Endometrial biopsy analysis

Machine learning models analyze endometrial biopsy images to determine properties like implantation readiness and pregnancy risk, overcoming the limitations of existing methods by providing accurate and efficient assessments without gene expression data.

WO2026099417A1PCT designated stage Publication Date: 2026-05-15UNIVERSITY OF WARWICK
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
UNIVERSITY OF WARWICK
Filing Date
2025-11-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for determining endometrial receptivity and timing of implantation are limited by their categorical approach, lack of temporal information, and require expensive and time-consuming gene expression data, making them unsuitable for routine clinical use.

Method used

A method using machine learning models to analyze image data from endometrial biopsies, enabling accurate determination of endometrial properties such as implantation readiness, decidualization, and risk of pregnancy loss, without the need for gene expression data.

Benefits of technology

Provides efficient, cost-effective, and accurate determination of endometrial properties, allowing for targeted interventions and personalized reproductive health strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025082255_15052026_PF_FP_ABST
    Figure EP2025082255_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A method of determining at least one property of the endometrium, the method comprising: receiving image data associated with an endometrial biopsy specimen from a subject; inputting the data into a machine learning model; and using the machine learning model to determine the at least one property of the endometrium of the subject based on the image data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] LUTEAL PHASE

[0002] FIELD OF THE INVENTION

[0003] The present disclosure relates to determining properties of the endometrium. In particular, but not exclusively, the present disclosure relates to methods and apparatuses for determining properties of the endometrium based on image data.

[0004] BACKGROUND

[0005] The inner lining of the uterus, called the endometrium, undergoes dynamic changes throughout the menstrual cycle and plays a critical role in fertility and reproductive health. The precise molecular and cellular changes during each menstrual cycle consists of proliferation, differentiation, and menstruation if pregnancy does not occur. This cyclic remodelling and regeneration under the influence of hormonal fluctuations, such as progesterone and oestrogen, produces changes in gene expression levels to modify the tissue, allowing for a functional implantation window to occur during the mid-luteal phase. Upon implantation of a developmentally competent embryo, the endometrium transforms from a cycling tissue into a robust, immune-privileged matrix that accommodates the semi-allogenic placenta throughout pregnancy. This process is known as decidualisation. Deviations in this finely orchestrated process can lead to disturbances in the window of implantation resulting in a range of gynaecological infertility disorders including recurrent implantation failure (RIF) and recurrent pregnancy loss (RPL). Understanding the complex biology and regulatory mechanisms governing the endometrium is crucial for understanding the physiological and pathological processes occurring within the uterus. Accurate timing of endometrial biopsies is therefore essential in endometrial pathology as it allows pathologists to evaluate tissue samples within their relative context of the stage of the menstrual cycle.

[0006] Commonly, it is practical to schedule biopsies relative to the preovulatory luteinising hormone surge, with the number of days following a positive urinary ovulation test known as LH+. However, this does not allow for direct comparison between samples due to individual differences in the timing of the luteinizing hormone surge in comparison with the day of ovulation as well as individual variability in patient progesterone levels after the luteinizing hormone surge which will affect the endometrial phenotype. There is also the risk of inaccuracies of self-reported LH+ days which can greatly impact the potential interpretation of images. To date, numerous techniques have been developed to detect the implantation window in the luteal phase, with a particular focus on utilizing transcriptomic and proteomic biomarkers. Furthermore, several computational techniques have been established to infer timing of the implantation window through the examination of gene expression patterns. Among these techniques, Win-Test employs a set of 11 genes, ER Map / ER Grade utilizes a more extensive panel of 184 genes, the TED model uses 73 genes, and the Endometrial Receptive Array (ERA) incorporates 238 genes. Although these methods are valuable in assessing endometrial receptivity, they are limited by their categorical approach in predicting receptivity status, such as pre-receptive, receptive, or post-receptive, which lacks specific temporal information and does not provide a continuous estimate. EndoTime is another computational method which has been developed to accurately estimate the exact timing of endometrial biopsies on a continuous basis using a panel of six genes (JL2RB, IGFBP1, CXCL14, DPP4, GPX3, and SLC15A2). This combines models for each gene to estimate a timing which generates sharper temporal expression profiles compared to patient-reported LH+ times, showing that the gene inferred timing represents a more accurate ordering for pathological comparison. A significant weakness of existing methods is that they require gene expression data, which is expensive and timeconsuming to generate, especially for large-scale studies. This limits their utilization and effectively rules them out for routine clinical use.

[0007] SUMMARY OF THE INVENTION

[0008] It is an object of the disclosure to at least partly address one or more of the shortcomings in the prior art.

[0009] According to an aspect of the invention, there is provided: a method of determining at least one property of the endometrium, the method comprising: receiving image data associated with an endometrial biopsy specimen from a subject; inputting the data into a machine learning model; and using the machine learning model to determine the at least one property of the endometrium of the subject based on the image data. Machine learning models are used to provide improved determination of properties of the endometrium of a subject based on image data, thereby providing an efficient and cost-effective way to handle unseen images. Accurate predictions, such as the ability to predict the timing of endometrial biopsies directly from histology images provides an objective tool that may be used in both a clinical setting, as well as for further investigations to advance understanding of endometrial biology and its implications for various pathological conditions.

[0010] Optionally, the at least one property comprises at least one of: a value associated with readiness to implantation and / or decidualisation and / or healthiness. Optionally, the at least one property is associated with a time within a luteal phase of the subject. Optionally, the at least one property comprises a relative time in the menstrual cycle of the endometrial biopsy specimen and / or a gene expression level and / or a protein expression level. The machine learning model may output determined properties with a high accuracy based on image data, thereby avoiding the need for additional physical processing of endometrial biopsy specimens. In turn, this enables faster, more efficient widespread sample analysis with reduced materials and costs in both clinical settings and beyond, in a way that was not previously possible.

[0011] Optionally, the at least one property comprises a risk of pregnancy loss, thereby enabling targeted interventions and personalized reproductive health strategies at an early stage due to improved accuracy without complex sample analysis.

[0012] There is also provided: a method of training a machine learning model for determining at least one property of the endometrium, the method comprising: receiving image data associated with an endometrial biopsy specimen of a subject, the received image data corresponding to at least one property of the endometrium; and training the machine learning model to determine at least one property of the endometrium based on the received image data.

[0013] Appropriate processing of image data and training of a machine learning model enables images to be used to determine properties of the endometrium in a way that avoids conventional sample processing.

[0014] Optionally, the at least one property of the endometrium is associated with a time within a luteal phase of a subject. Optionally, the method of training a machine learning model comprises processing an endometrial biopsy specimen of the subject, determining whole slide image data of the processed endometrial biopsy specimen of the subject; and determining one or more gene expression levels associated with processed endometrial biopsy specimen of the subject. Training a machine learning model in this way enables images to be used to provide accurate determinations, such as accurate time-based determinations of properties of the endometrium, even when images are not expected to reveal such information.

[0015] There is also provided a method of determining a risk of pregnancy loss. The method comprises: receiving image data associated with an endometrial biopsy specimen from a subject; measuring one or more attributes of histological constructs in the image data; and determining the risk of pregnancy loss based on the one or more attributes.

[0016] The use of image analysis allows for the identification of novel histological biomarkers for pregnancy risk. Such marker have not previously been identified due to the complex and highly variable nature of the histological composition of the endometrium.

[0017] There is also provided: a computing device configured to perform the methods and a computer program or computer readable medium comprising instructions that, when executed by a computer, cause the computer to perform the methods.

[0018] There is also provided: a system for performing the methods, the system comprising: an imaging apparatus for imaging an endometrial biopsy specimen of a subject; and the computer device.

[0019] There is also provided: a machine learning model for determining at least one property of the endometrium, wherein the machine learning model is configured to: receive image data associated with an endometrial biopsy specimen from a subject; and determine the at least one property of the endometrium based on the received image data.

[0020] DETAILED DESCRIPTION

[0021] Embodiments of the disclosure will be further described by way of example only with reference to the accompanying drawings in which:

[0022] Figure 1 shows a system for determining properties of the endometrium;

[0023] Figure 2 shows a process flow for determining properties of the endometrium;

[0024] Figure 3 shows a process flow for training a machine learning model to determine properties of the endometrium;

[0025] Figure 4 shows a process flow for processing and evaluating data;

[0026] Figure 5 shows a process flow for generating a graph representation;

[0027] Figure 6 shows a process flow using a graph neural network;

[0028] Figure 7 shows a visualisation of gene expression and LH+ prediction versus ground truth to the nearest day;

[0029] Figure 8 shows a visualisation of image-based prediction versus ground truth to the nearest day;

[0030] Figure 9 shows a visualisation of WSI node level predictions versus ground truth to the nearest day;

[0031] Figure 10 shows a scatter plot of predicted day versus ground truth;

[0032] Figure 11 shows representative patches for different days within the luteal phase;

[0033] Figure 12 shows a heatmap visualisation of each whole slide image in the training set for patch level scores during cross fold validation for different days within the luteal phase;

[0034] Figure 13 A shows the proportion of patients falling above the 25thpercentile and below the 75thpercentile for patients with a different number of pregnancy losses for predicted gene expression ratio values;

[0035] Figure 13B shows the proportion of patients falling above the 25thpercentile and below the 75thpercentile for patients with a different number of pregnancy losses for measured gene expression ratio values;

[0036] Figure 14 shows boxplots of day-normalized values of machine learning model features on days 8 to 11 comparing recurrent missed miscarriage (RMM) and control patients, along with illustrative images from whole slides; Figure 15 shows a heatmap depicting the correlation between the glandular characteristics discussed above and predicted gene expression variables;

[0037] Figure 16 shows visualisations of the correlation of various gene expression levels and histological attributes with pregnancy loss;

[0038] Figure 17 shows how different machine learning model features vary between RMM and control patients and the relative importance of the different features to predicting RMM;

[0039] Figure 18 shows the correlations of RMM prediction scores with self-reported LH+ day for genes with the largest group differences between RMM and control groups;

[0040] Figure 19 shows the correlations of prediction scores with LH+ day for all other genes;

[0041] Figure 20 shows representative images of exemplary cellular attributes; and

[0042] Figure 21 shows representative images of exemplary glandular attributes.

[0043] In the present disclosure, methods and apparatuses are described which use machine learning models in order to determine one or more properties of the endometrium based on image data. It has been realised that the integration of machine learning with histological image analysis provides a powerful approach to determining properties of the endometrium. In particular, surprisingly, it has been realised by the inventors that histology whole slide images may be used in order to infer the exact day in the luteal phase, along with properties such as readiness to implantation, decidualisation, healthiness, a relative time in the menstrual cycle, gene expression level, protein expression level, a prediction of the number of days following a luteinizing hormone surge or ovulation and / or risk of pregnancy loss.

[0044] In order to determine one or more properties of the endometrium, image data associated with an endometrial biopsy specimen of a subject is obtained through appropriate processing steps described herein and subsequently input into a machine learning model. The machine learning model, which is a model trained to determine at least one property of the endometrium based on the image data, is then used in order to output at least one determined property of the endometrium based on the input.

[0045] Figure 1 shows a system 100 for determining at least one property of the endometrium. The system 100 comprises a computing device 102 and an imaging apparatus 114. Optionally, the computing device 102 is in communication with one or more further computing devices forming part of a network 114. The computing device 102 comprises an input 104, an output 106, a memory 108 and a processor 110. In further implementations, the computing device 102 comprises additional or alternative components enabling the functionality described herein.

[0046] The computing device 102 is configured to receive image data associated with an endometrial biopsy specimen from a subject at the input 104 of the computing device 102. The endometrial biopsy specimen may be obtained and processed in a variety of different ways, as described in further detail below, in order to provide image data. The processed specimen may be imaged by an imaging apparatus 116 that is in communication with the computing device 102 via a communication pathway 101. Alternatively, or additionally, the processed specimen may be imaged by an imaging apparatus 116 that is in communication with the network 114 of one or more further computing devices via a communication path 105 between the imaging apparatus 116 and the network 114. The computing device 102 may then receive the image data at the input 104 of the computing device 102 via the communication path 103 between the network 114 and the computing device 102. Alternatively, or additionally, the image data may be acquired via an external imaging apparatus and input directly into the computing device 102 via the input 104.

[0047] The computing device 102 uses a machine learning model 112 to determine at least one property of the endometrium of a subject based on image data associated with an endometrial biopsy specimen from a subject. The machine learning model 112 is a model that may have been trained at the computing device 102 or at one or more other computing devices, such as the one or more further computing devices of the network 114. The machine learning model 112 may be stored in the memory 108 of the computing device 102. Alternatively, the machine learning model 112 may be stored at a different computing device, such as a computing device of the network 114 of one or more further computing devices. Accordingly, the computing device 102 is able to access and / or use the machine learning model 112 directly by accessing the memory 108 of the computing device 102, or indirectly by accessing the machine learning model 112 stored in the network 114 of one or more further computing devices. In an implementation, the machine learning model 112 is stored on a server in the network 114 of one or more further computing devices, thereby providing a distributed computing network. The use of a distributed computing network means that image data may be uploaded at one location, processed remotely and sent to a computing device at the original location and / or other locations. This means that improved determination of properties of the endometrium can be made without the need for specialised equipment, such as equipment required for measuring gene expression levels and without sample being physically present at the location at which they are analysed.

[0048] The computing device 102 is configured to output the determined one or more properties of the endometrium of the subject based on the image data at the output 106 of the computing device. The output 106 may be a display device. Alternatively, or additionally, the output 106 communicates the determined one or more properties of the endometrium to one or more further computing devices, such as those of the network 114 via the communication path 103 between the computing device 102 and the network 114. The communication paths 101, 103, 105 between different entities of the system 100 each comprise any appropriate combination of wired and / or wireless communication links. The communication paths 101, 103, 105 may be formed directly between the different entities of the system 100, or may include one or more intermediate devices.

[0049] Whilst the system 100 described with reference to Figure 1 is shown to include particular devices, it will be appreciated that in further implementations, additional and / or alternative devices are included whilst providing the functionality described herein. For example, the system 100 may comprise devices enabling at least partial automation of the methods described herein, such that steps relating to obtaining, processing and / or analysing samples are coordinated by the computing device 102 and / or one or more further computing devices of the network 114.

[0050] The system 100 described with reference to Figure 1 is used in order to implement at least part of a method of determining at least one property of the endometrium. Figure 2 shows a process flow S200 for determining at least one property of the endometrium. The process flow S200 is started at step S202 in response to one or more instructions. In an implementation, the one or more instructions are computer implemented instructions executed by the computing device 102 of Figure 1 in order to instigate the process flow S200. In further implementations, additionally, or alternatively, the one or more instructions comprise manual user instructions. As described above, the computing device 102 is configured to receive image data associated with an endometrial biopsy specimen from a subject. In some cases, the image data may have already been prepared externally based on an endometrial biopsy specimen that has already been obtained, in which case the process flow S200 may directly move to step S210. In other cases, initiation of the process flow S200 requires an endometrial biopsy specimen to be obtained, in which case the process moves to step S204, where an endometrial biopsy specimen may be obtained.

[0051] The endometrial biopsy specimen typically comprises endometrial cells or tissue. The endometrial cells or tissue may comprise, by way of example, luminal, glandular, and ciliated epithelial cells, arterial, venous, and lymphatic endothelial cells, pericytes / perivascular cells, lymphoid and myeloid cells, progesterone-resistant stromal cells, progesterone-responsive stromal cells and / or uterine natural killer (NK) cells. The endometrial biopsy specimen may be obtained using any appropriate technique, which may be manual or at least partly automated. In one implementation, the specimen is obtained by an endometrial biopsy catheter. In further implementations, the specimen is obtained by dilation and curettage, by biopsy forceps or curette during hysteroscopy, following hysterectomy or simultaneously with a procedure involving instrumentation of the uterus. In further implementations, the endometrial biopsy specimen is obtained using any appropriate additional or alternative technique. The endometrial biopsy specimen may be prepared for subsequent processing by any suitable technique, for example fixation, such as chemical fixation. Preferably a formalin-fixed paraffin embedded (FFPE) endometrial biopsy specimen is prepared. In one implementation, the endometrial biopsy specimen is obtained within 5 to 12 days after a positive urinary ovulation test, designated as LH+. In further implementations, the endometrial biopsy specimen is alternatively or additionally obtained at any appropriate time range.

[0052] Once the endometrial biopsy specimen has been obtained at step S204, the process moves to step S206, where the endometrial biopsy specimen is processed. The endometrial biopsy specimen is processed such that an endometrial whole slide image can be obtained that is suitable for further processing and inputting into the machine learning model. In one implementation, the endometrial biopsy specimen is a FFPE biopsy sample for which the tissue is subjected to staining for CD56, which serves as a marker for uterine natural killer (uNK) cells, and hematoxylin to visualise cell nuclei. However, in further implementations, the endometrial biopsy specimen is processed using any suitable additional or alternative technique, such as a different staining technique. Different staining techniques include, but are not limited to, Masson’s tri chrome staining for collagen deposits, pl6-INK4 staining for senescent cells, CD68 staining for macrophages, PAEP staining for differentiated glandular epithelial cells, acetylated a-tubulin staining for ciliated epithelial cells, Von Willebrand factor (vWF) staining for endothelial cells, routine Haematoxylin and Eosin (H&E) staining, PLA2G2A staining for (pre-)decidual cells, DIO2 staining for subluminal stromal cells and WNT7A staining for luminal epithelial cells.

[0053] Once the endometrial biopsy specimen has been processed at step S206, or when then sample has been prepared externally, the process moves to step S210, where image data is obtained. Image data may be obtained at the imaging apparatus 116 described with reference to Figure 1, or may be input into the computing device 102 by any appropriate mechanism. In one implementation, a whole slide image is obtained that is a bright-field image, for example, a bright-field image captured at the imaging apparatus 116. In further implementations, the whole slide image is in any appropriate form such that the image data comprises histologic whole slide image data that can be used to provide the functionality described herein. The whole slide image data may be formed of pixels at a given image resolution. In one implementation, the image resolution of the whole slide image data is 0.5 MPP. In further implementations, the image resolution of the whole slide image data is any appropriate image resolution to enable the functionality described herein.

[0054] Once the whole slide image data has been obtained at step S210, the process moves to step S212, where the whole slide image data is processed in order to provide data that is inputted into a machine learning model. The whole slide image data is processed in any appropriate manner in order to enable the machine learning model 112 to determine at least one property of the endometrium of a subject based on the image data. An exemplary implementation using a graph neural network is described in further detail below, with reference to Figures 4 to 13. In such cases, the whole slide image may be processed by segmenting the received whole slide image data into a plurality of regions and identifying a subset of one or more regions in order to generate graph representations for input into a graph neural network machine learning model. In further implementations, one or more alternative or additional machine learning models may be used and the whole slide image data is processed in any appropriate manner for input into the one or more alternative or additional machine learning models. For example, the whole slide image data may be processed by segmenting the whole slide image data into a plurality of regions, thereby facilitating the selective extraction of areas of interest from the whole slide image data when the processed image data is subsequently input into the machine learning model. In one implementation, processing of image data at step S212 may include analogous segmentation as described herein with reference to step S310 of Figure 3, in order to provide a suitable input for a machine learning model at step S214.

[0055] Once the image data has been processed at step S212, the process flow S200 moves to step S214, where the processed image data is input into a machine learning model, such as the machine learning model 112 described with reference to Figure 1.

[0056] The machine learning model 112 is a model that has been trained in order to receive image data associated with an endometrial biopsy specimen of a subject and to determine at least one property of the endometrium based on the received image data. The machine model 112 is trained on received image data corresponding to a least one property of the endometrium, as described in further detail herein with reference to Figure 3. For example, as described herein, the machine learning model 112 has been trained using at least one of: one or more gene expression levels corresponding to a biopsy specimen of a subject; one or more histological constructs; one or more cellular descriptors; one or more textural descriptors and one or more image features. In further implementations, the machine model 112 is additionally trained using any appropriate clinical outcome or label. For example, labels such as “fit for implantation” and “unfit for implantation may be used. Training the machine learning model 112 on such features provides improved determination of properties of the endometrium by learning relationships between image features and target variables (where target variables may be non-visual features, sub-visual features or other visual features). For example, the machine learning model 112 trained on gene expression levels may provide endometrium timing determinations with an accuracy that is similar to methods such as EndoTime solely based on unseen image data.

[0057] The process flow S200 then moves to step S216, where the machine learning model 112 determines at least one property of the endometrium based on the image data. In one implementation, the at least one property comprises at least one of a value associated with readiness to implantation and / or decidualisation and / or healthiness. In a further implementation, additionally or alternatively, the at least one property is associated with a time within a luteal phase of the subject. For example, the at least one property may be a gene expression level on a day within the luteal phase. In a further implementation, additionally or alternatively, the at least one property comprises a relative time in the menstrual cycle of the endometrial biopsy specimen. While the average length of menstrual cycle is 28 days, there is considerable intra- and inter-individual variation and the use of the machine learning model 112 enables determination of the relative time in the menstrual cycle irrespective of external factors.

[0058] In a further implementation, additionally or alternatively, the at least one property comprises a gene expression level and / or a protein expression level. The expression level may relate to a particular time within the menstrual cycle, or an evolution of expression level over a particular time period. In an example, the gene expression level allows identification of the day in the menstrual cycle. The gene expression level may be of at least one of IL2RB, IGFBP1, CXCL14, DPP4, GPX3 and SLC15A2. In further implementations, the gene expression level may be of any appropriate gene.

[0059] The gene expression level may comprise at least one marker gene for progesteronedependent decidual cells. Optionally, the at least one marker gene is selected from SLC15A2, DPP4, GPX3, IGFBP1, IL2RB, ITGAD, CXCL14, DIO2, PLA2G2A, TIMP3, IL15, TAGAP, SCARA5, IGF2 and WNT4. In further implementations, other additional or alternative marker genes may be used.

[0060] Additionally, or alternatively, the gene expression level comprises at least one marker gene for progesterone-resistant stromal cells, which may be selected from DIO2, HOXAIO, HOXA11, WNT5A and IGF1.

[0061] In an implementation, the at least one property comprises a ratio of gene expression levels. For example, the ratio of gene expression levels comprises a ratio between GPX3 and SLC15A2 gene expression levels and / or a ratio between PLA2G2A and DIO2 gene expression levels. In further implementations, additional or alternative ratios of gene expression levels are provided.

[0062] In an implementation, the at least one property comprises a prediction of the number of days following a luteinizing hormone surge or ovulation.

[0063] In an implementation the at least one property comprises a risk of pregnancy loss, optionally wherein the risk comprises a risk of pregnancy loss on a particular day in the endometrial cycle. The risk of pregnancy loss may be based on a ratio of gene expression levels associated with the biopsy specimen of the subject. For example, the ratio of gene expression levels comprises a ratio between PLA2G2A and DIO2 gene expression levels. In further implementations, the risk of pregnancy loss may be based on any suitable alternative or additional gene expression levels and / or ratios.

[0064] As mentioned above and described in more detail in relation to step 310 below, the processing of the image data may comprise segmentation of the image data. In some implementations, the segmentation may comprise segmentation into histological constructs, such as cells, glands, lumen, luminal epithelium, subluminal stroma and spiral arteries. Correspondingly, the at least one property of the endometrium may comprise features associated with these histological constructs. In particular, the at least one property may comprise characteristics of colour and morphology of cells, and / or characteristics (also referred to as features) of colour and morphology of glands, and / or attributes capturing gland-cell relationships and epithelial organization.

[0065] Colour attributes of cells and glands include the mean, maximum, median, and standard deviation of the Hematoxylin and CD56 channels within the bounding area of the cell, quantifying staining intensity and variability. Morphological characteristics of cells and glands comprised area, solidity, eccentricity, circumference, equivalent diameter, major and minor axis lengths, and convex area of the cell boundary contour. These characteristics collectively characterize cell size and shape and provide an informative summary of cellular phenotypes within the tissue.

[0066] Attributes of the gland-cell relationships and epithelial organization include the number of constituent cells, cell density (total cell area divided by gland area), and statistics (mean, median, maximum, standard deviation) of distances from epithelial nuclei centroids to their nearest gland boundary. To further quantify epithelial pseudostratification, attributes may further include the number of neighbouring epithelial cells within a predetermined radius (e.g. 50-pixels) and the distances to a predetermined number (e.g. eight) of nearest neighbours.

[0067] Figure 20 shows representative images illustrating examples of cellular attributes that may be used by the model and subsequent analysis. Figure 21 shows corresponding representative images for glandular attributes and attributes capturing epithelial organization. Tables 1 and 2 show feature descriptions and corresponding histological interpretations for the cellular and glandular characteristics.

[0068] Table 1

[0069]

[0070] Table 2

[0071] The property of the endometrium determined at step S216 may be provided at a whole slide image level, or a node level, where the whole slide image data can be considered to be formed from contributions of spatially distributed nodes. For example, in one implementation the machine learning model 112 comprises a graph neural network that is able to generate nodelevel determinations before aggregation into a whole slide image data prediction. Node-level determinations, which may be provided in the form of scores, enable identification of specific areas within a whole slide image that contribute positively or negatively to the whole slide image determination of at least one property of the endometrium of a subject. This provides a further layer of interoperability in both training and using a machine learning model in order to determine at least one property of the endometrium of a subject. Whilst this functionality may be provided through the use of a graph neural network, in further implementations, the functionality is provided by any suitable machine learning model.

[0072] When at least one property of the endometrium has been determined, the process flow S200 then moves to step S218, where one or more determined properties of the endometrium are output. The determined properties are output at an output display 106 of the computing device 102. The determined properties may be output at a whole slide image data level and / or at a node level for spatially distributed nodes associated with the whole slide image data. Additionally, or alternatively, the determined properties are output at any alternative or additional device, such as one or more computing devices in the network 114. The one or more determined properties of the endometrium may be output by the machine learning model 112 in any suitable format. In one implementation, the one or more determined properties of the endometrium are output by the machine learning model 112 in the form of a predicted value and associated confidence value. The process flow S200 ends at step S220.

[0073] Whilst the steps of process flow S200 are described in a particular order, it will be appreciated that in further implementations, the steps may be performed in a different order and / or concurrently with additional, fewer and / or alternative steps in accordance with the functionality described herein, in order to determine at least one property of the endometrium of a subject based on received image data associated with an endometrial biopsy specimen of the subject being input into a machine learning model.

[0074] The machine learning model 112 used in the process flow S200 of Figure 2 is a machine learning model that has been trained in order to determine at least one property of the endometrium based on received image data. Figure 3 shows a process flow S300 of a method of training a machine learning model for determining at least one property of the endometrium. The method may be used to train the model 112 described with reference to Figures 1 and 2. The process flow S300 may be at least partially implemented at the system 100 described with reference to Figure 1 and is initiated at step S302 in response to one or more computer implemented and / or user instructions.

[0075] In order to provide an effective machine learning model, appropriate data is used to form a dataset for training and testing the machine learning model. At step S304, an endometrial biopsy specimen of a subject is obtained. As described above with reference to Figure 2, an endometrial biopsy specimen of a subject may be obtained by any appropriate technique such that it can be processed to form part of the training, validation and / or testing data set for the machine learning model. For example, the endometrial biopsy specimen may be obtained: by an endometrial biopsy catheter; by dilation and curettage; by biopsy forceps or curette during hysteroscopy; following hysterectomy or simultaneously with a procedure involving instrumentation of the uterus. In one implementation, the endometrial biopsy specimen is obtained within 5 to 12 days after a positive urinary ovulation test, this self-reported time being designated as LH+. In further implementations, the endometrial biopsy specimen is alternatively or additionally obtained from any appropriate range of times. The endometrial biopsy specimen may be prepared for subsequent processing by any suitable technique, for example fixation, such as chemical fixation. Preferably a FFPE endometrial biopsy specimen is prepared. In one implementation, the endometrial biopsy specimen is a FFPE biopsy sample for which the tissue is subjected to staining for CD56, which serves as a marker for uterine natural killer (uNK) cells, and hematoxylin to visualise cell nuclei. However, in further implementations, the endometrial biopsy specimen is processed using any suitable additional or alternative technique, such as a different staining technique. Different staining techniques include, but are not limited to, Masson’s tri chrome staining for collagen deposits, pl6-INK4 staining for senescent cells, CD68 staining for macrophages, PAEP staining for differentiated glandular epithelial cells, acetylated a-tubulin staining for ciliated epithelial cells, Von Willebrand factor (vWF) staining for endothelial cells, routine Haematoxylin and Eosin (H&E) staining, PLA2G2A staining for (pre-)decidual cells, DIO2 staining for subluminal stromal cells and WNT7A staining for luminal epithelial cells.

[0076] Once the specimen has been obtained, it is processed at step S306 in order to enable the eventual training and / or testing of the machine learning model. In an implementation, processing the endometrial biopsy specimen includes associating the endometrial biopsy specimen of the subject with a time within a luteal phase of a subject and / or a relative time in the menstrual cycle of the endometrial biopsy specimen. The association of samples with further information enables effective training of the machine learning model 112. In one implementation, RTq-PCR data is associated with each endometrial biopsy specimen, along with expressions of individual genes along with EndoTime as calculated using the method described in Lipecki, J., Mitchell, A.E., Muter, J., Lucas, E.S., Makwana, K., Fishwick, K., Odendaal, J., Hawkes, A., Vrljicak, P., Brosens, J. J., Ott, 8.: EndoTime: non-categorical timing estimates for luteal endometrium. Human Reproduction 37(4), 747-761 (2022).

[0077] In one implementation, alternatively or additionally, the endometrial biopsy specimen is associated with a gene expression level and / or a protein expression level. Preferably, the gene expression level may comprise at least one gene that allows identification of the day in the menstrual cycle. Alternatively, or additionally, the gene expression level is of at least one of IL2RB, IGFBP1, CXCL14, DPP4, GPX3 and SLC15A2. However, in further examples of the implementation, the endometrial biopsy specimen is associated with any appropriate gene expression level. In one implementation, additionally or alternatively, the gene expression level comprises at least one marker gene for progesterone-dependent decidual cells, optionally selected from SI.C15A2, DPP4, GPX3, IGFBP1, II.2RB, ITGAI), CXCI.14, DIO2, PLA2G2A, TIMP3, IL15, TAGAP, SCARA5, IGF2 and WNT4 and / or at least one marker gene for progesterone-resistant stromal cells, optionally selected from DIO2, HOXA Kk H0XA 1 L WNT5A and IGF1. However, in further examples of the implementation, alternatively or additionally, the endometrial biopsy specimen is associated with any appropriate marker gene for progesteronedependent decidual cells and / or for progesterone-resistant stromal cells.

[0078] In one implementation, alternatively or additionally, the endometrial biopsy specimen is associated with a ratio of gene expression levels. In an example the ratio of gene expression levels comprises a ratio between GPX3 and SLC15A2 gene expression levels and / or a ratio between PLA2G2A and DIO2 gene expression levels. However, in further examples, any appropriate ratio of gene expression levels may be alternatively or additionally associated with the endometrial biopsy specimen.

[0079] In one implementation, alternatively or additionally, the endometrial biopsy specimen is associated with a number of days following a luteinizing hormone surge or ovulation.

[0080] In one implementation, alternatively or additionally, the endometrial biopsy specimen is associated one or more clinical variables, such as previous pregnancy losses associated with the subject from whom the endometrial biopsy specimen is obtained.

[0081] In further implementations, additionally or alternatively the endometrial biopsy specimen of the subject is associated with further parameters and / or values that enable determination of at least one property of the endometrium by the trained machine learning model 112. For example, in further implementations, additionally or alternatively, the endometrial biopsy specimen is processed to determine at least one of a cell count; a histological construct; one or more cellular descriptors and one or more textural descriptors. Cellular descriptors can include morphological descriptors of the nuclear or cell shape (area, irregularity, etc), colour features within the cell, as well as other properties of the areas around the cell. Textural descriptors are features of image texture that can be obtained through conventional texture characterization approaches or deep learning models. The histological construct may be selected from glands, lumen, luminal epithelium, subluminal stroma and / or spiral arteries. Such determinations may be made by analysing a whole slide image associated with the endometrial biopsy specimen and data may be labelled accordingly.

[0082] Once the specimen has been processed at step S306, whole slide image data is determined at step S308. Determining whole slide image data may comprise imaging a processed endometrial biopsy specimen. Additionally, determining whole slide image data may comprise associating one or more parameters and / or values with the processed endometrial biopsy specimen, such as those described with reference to processing image data at step S306. For example, additionally or alternatively, determining whole slide image data may comprise determining one or more gene expression levels associated with the processed endometrial biopsy specimen of the subject. Whilst gene expression is known to provide accurate timing within the luteal phase, for example using the particular genes identified in EndoTime, such gene expression is not a visible effect. Nevertheless, the use of gene expression may be used indirectly in the machine learning model in order to provide accurate determinations based solely on images, in a way that would not have been expected.

[0083] Once whole slide image data has been determined at step S308, the whole slide image data is subsequently processed sequentially and / or simultaneously for each sample that is being used in the dataset. Suitably processed whole slide image data may then be input into a machine learning model in order to train the machine learning model. The processing of the whole slide image data that is required may depend on the type of machine learning model that is being used. Additional processing steps may therefore be implemented at step S308.

[0084] Whilst steps S304 to S308 describe processes for determining whole slide image data for one endometrial biopsy specimen from a subject, it will be appreciated that in order to train a machine learning model, a dataset based on determining whole slide image data for multiple subjects enables the effective training and testing of a machine learning model. Therefore, steps S304 to S308 are performed for each subject sample in order to build an appropriate dataset. The dataset may then be refined in order to improve the training, validation and / or testing of a machine learning model. For example, the dataset may be refined by excluding specimens for which there are quality issues, missing data and / or labels. The dataset may comprise any appropriate number of specimens, which may be obtained and processed sequentially and / or simultaneously.

[0085] Once the whole slide image data has been determined at step S308, the process advances. Optionally, at step S310, the determined whole slide image data may be processed in order to provide segmented image data. Segmenting received image data into a plurality of regions enables the selective extraction of areas of interest from the whole slide image data. In one implementation, segmenting whole slide image data comprises user-defined identification of boundaries of endometrial tissue for processing. In further implementations, additional or alternative techniques may be used. In one implementation, the whole slide image data is additionally or alternatively segmented into a plurality of regions, also referred to as patches, each measuring 224 x 224 pixels at an image resolution of 0.50 MPP. In further implementations, whole slide image data may be segmented into any suitable number of patches of any suitable shape and size, including different pixel distributions and / or image resolutions. For example, whole slide image data may additionally or alternatively be segmented into histological constructs, such as glands, lumen, luminal epithelium, subluminal stroma and spiral arteries. These constructs may then be used as an input to a machine learning model together with cellular and textural descriptors, or deep features of various image regions to provide improved predictions of properties of the endometrium. By segmenting whole slide image data and extracting regions of interest, a machine learning model may be trained to identify and / or exclude artefacts that may otherwise negatively influence determinations by the machine learning model. For example, artefacts arising through the use of different imaging apparatuses may be eliminated. Further, unexpected links between features can be exploited by training a machine learning model on combinations of processed image data.

[0086] In further example implementations, additionally or alternatively, the whole slide image data may be segmented into a plurality of regions at least partially based on pixels of the whole slide image data and / or the resolution of the whole slide image data. For example, the whole slide image data may be segmented into a plurality of regions formed from multiple pixels of the whole slide image data. The number of pixels of each region of the plurality of regions may depend on the image resolution of the whole slide image data. Segmentation of the whole slide image data into a plurality of regions, each region comprising multiple pixels, provides a practical way of processing whole slide image data whilst enabling the training of a machine learning model in order to determine one or more properties of the endometrium. For example, where multi-gigapixel histopathology images are used to provide the whole slide image data, the images may be segmented into regions (also referred to as patches) measuring 224 x 224 pixels at an image resolution of 0.50 MPP. However, other combinations of region size and image resolution may be used when segmenting the whole slide image data at step S310. It will be appreciated that whilst a plurality of regions in arrays of 224 x 224 pixels are described, alternatively or additionally, the pixels of the plurality of regions may be provided in different sized arrays with any appropriate array and pixel dimensions so as to form regions corresponding to any regular or irregular polygon, or other appropriately shaped patch that may be used to segment the whole slide image data. For example, irregular polygons may be correlated with features such as glands, cells, or other appropriate feature. Such segmentation may be performed in combination with segmentation based on histological constructs. The combination of segmentation may be implemented in any appropriate manner and order, such as by segmenting based on histological constructs in a first step, then segmenting based on pixel distributions and / or image resolution in a second step, or vice versa. Segmentation of the whole slide image data based on pixel distribution and / or image resolution provides an effective route to selectively train a machine learning model on regions of interest in the whole slide image data. For example, appropriate segmentation enables region-based parameters, such as mean pixel intensity, to be used in order to determine extraction of a subset of the plurality of regions without the need to evaluate data on a per-pixel basis.

[0087] Once the image data has been optionally segmented at step S310, the process flow S300 then moves to step S312, where segmented patches are processed and extracted from the segmented image data in order to provide patch image data. The processing and extraction of patches from the segmented image data provides one or more regions of the plurality of regions that can be used selectively to train the machine learning model 112. Patch processing and extraction may be based on identifying a subset of one or more regions of the plurality of regions corresponding to a tissue, so that the machine learning model can be trained on the subset. In an implementation, where the whole slide image data has been segmented into a plurality of regions each comprising multiple pixels, subsets of the plurality of regions may be identified for extraction based on a metric encompassing contributions from each of the pixels within the patches, rather than analysing each individual pixel separately. For example, the subset may be identified by determining one or more regions of the plurality of regions having a mean pixel intensity of at least a predetermined threshold. In further implementations, any appropriate additional or alternative techniques are used in order to identify a subset of the plurality of regions.

[0088] In an implementation, segmented patches are processed and extracted from the segmented image data in order to provide patch image data in order to develop a graph representation for input into a graph neural network, as described with reference to Figures 4 to 13. In further implementations, additionally or alternatively, the segmented patches are processed and extracted from the segmented image data in order to provide patch image data that may be processed for input into any appropriate machine learning model. Patches of the patch image data may be associated with spatially distributed nodes in order to facilitate machine learning model predictions associated with contributions from individual nodes in addition to predictions associated with whole slide image data.

[0089] Determining the whole slide image data may further comprise applying a normalization procedure. As discussed above, there is considerable intra- and inter-individual variation in the level of gene expression over the menstrual cycle. This can make comparisons between different patients or between the same patient at different times very difficult and confuse the determination of properties such as a risk of pregnancy loss.

[0090] To account for natural variation across days, a normalization procedure may be applied to the data, which may be referred to as day-normalization. For each day within the menstrual cycle, the method may comprise grouping the image data (or the attributes derived therefrom such as gene expression levels and colour or morphological characteristics discussed above) and fitting a normal probability distribution that described the mean and variability specific to that day. Each measurement can then be converted into a percentile indicating its relative position within that day’s distribution. This transformation allows values collected on different days to be compared on a standardized scale that accounts for normal day-to-day fluctuations in measurement patterns.

[0091] Once the whole slide image data has been determined at step S308, or where the determined whole slide image data has been processed in order to determine segmented image data and once patch image data has been processed and extracted at step S312, the process flow then moves to step S314, where the processed image data is input into a machine learning model in order to train the machine learning model. The machine learning model may comprise a neural network. In further implementations, the machine learning model may comprise a graph neural network, as described further herein with reference to Figures 4 to 13. In further implementations, the machine learning model comprises any suitable additional or alternative machine learning model.

[0092] Once suitable data has been provided, the process progresses to step S314, where a machine learning model may be developed by inputting the processed image data into a machine learning model.

[0093] At step S314, the machine learning model processes the image data input at step S314 and outputs a prediction at step S316. The processed image data that is input into the machine learning model may comprise data associated with one or more images and the prediction may comprise corresponding output predictions associated with each of the one or more images. In such cases, the processed image data comprises a batch of grouped inputs and the machine learning model outputs a batch of grouped output predictions, each prediction corresponding to an input of the batch of grouped inputs. The machine learning model may be any appropriate type of machine learning model for receiving image data associated with an endometrial biopsy specimen from a subject and for determining at least one property of the endometrium of the subject based on the image data. The machine learning model may be trained by iteratively adjusting parameters of the machine learning model in response to a comparison between a prediction of the machine learning model and ground truth. Accordingly, at step S318 it is determined whether the prediction error associated with the predicted output at step S318 is below a predetermined threshold value. Where the predicted output corresponds to multiple inputs of processed image data as part of a batch of grouped inputs, the prediction error associated with the prediction output at step S316 may comprise an average of prediction errors associated with the output for each corresponding processed image data input. Prediction errors may be determined by based on a comparison of a predicted output with a target output in order to determine a difference between the predicted output and the target output. In further implementations, the prediction error may be determined using any appropriate additional or alternative method. If the prediction error is below the predetermined threshold value, the machine learning model is trained to a sufficient accuracy and the machine learning model is output at step S322.

[0094] In some implementations, additionally, at step S318, validation of the output prediction from step S316 may be performed. Validation of the output prediction may be performed by determining the error or loss associated with the machine learning model over a separate held- out validation set of data. Such validation may be performed by inputting data, such as processed image data corresponding to one or more images or other appropriate inputs, from the validation set of data into the machine learning model used at step S314 and outputting a prediction in the same manner as described with reference to step S316. The validation data may comprise data obtained and processed in the same manner as described with reference to the data obtained and processed at steps S304 to S312 of Figure 3. At step S318, the error or loss associated with the predicted output based on the validation data set may be compared with a predetermined threshold value in order to ensure that the machine learning model produces a correct output over unseen data.

[0095] Where validation of the machine learning model is performed at step S318, the determination of the prediction error threshold may comprise an additional determination of whether the error or loss associated with the predicted output based on the validation data set is below a predetermined threshold. In such a case, if both the prediction error associated with the predicted output at step S318 and the error or loss associated with a predicted output based on the validation data set are both below respective predetermined thresholds, the machine learning model may be considered to have been trained to a sufficient accuracy and the machine learning model is output at step S322. If one or both of the prediction error associated with the predicted output at step S318 and the error or loss associated with a predicted output based on the validation data set is / are not below the respective predetermined threshold, the process moves to step S320 for further adjustment of the machine learning model.

[0096] Each iteration of predicting an output at step S316 and adjusting the parameters of the machine learning model at step S320 is known as an epoch. The number of iterations, or epochs may be monitored. In some implementations, validation may be performed at each epoch when the machine learning model outputs a prediction at step S316. In further implementations, validation may be performed at one or more selected epochs, or in response to one or more events, such as one or more events based on the epoch number, time, error associated with the prediction, etc. Whilst the validation may be performed at step S318, in further implementations, validation may be performed at any suitable time in the process flow S300.

[0097] Alternatively, or additionally, the number of epochs may be monitored and the process flow may move from step S318 to S322 after a predetermined number of epochs has passed. This may be in combination with error determination associated the output prediction and / or a prediction based on validation data, or independent thereof.

[0098] In an example implementation, the machine learning model is a graph neural network that receives graph representations through the extraction and processing of patch image data in order to output whole slide image level predictions and node level predictions associated with spatially distributed nodes of the whole slide image data. In further implementations, additionally or alternatively, the machine learning model is any appropriate machine learning model that receives whole slide image data and outputs whole slide image level predictions and / or node level predictions, wherein the node level predictions are predictions associated with spatially distributed nodes, each node being associated with a different region of the whole slide image data. The predictions may be node level or whole slide image data level predictions for any appropriate property of the endometrium described herein. The process flow S300 then ends at step S324.

[0099] If the error prediction output at step S318 is not below the corresponding predetermined threshold value and / or the error or loss associated with a prediction based on the validation data set is not below its corresponding predetermined threshold and / or if the number of epochs has not reached a predetermined number, the process flow S300 moves to step S320, where one or more parameters of the machine learning model are adjusted.

[0100] Adjustment of the parameter of the machine learning model may be made using any appropriate method. In an example implementation, the machine learning model is initially configured with one or more random and / or selected parameters that are used to make predictions and processed image data is put into the machined learning model as described with reference to step S314. Based on the prediction error determined at step S318, which may be an average error for all inputs of a batch of grouped inputs of processed image data, the machine learning model may be adjusted at step S320 by changing one or more of the initially configured parameters. The process flow S300 then moves back to step S314, where the image data is input into the adjusted machine learning model, which then outputs an updated prediction at step S316. Subsequently, the error associated with the updated prediction at step S316 is compared to the predetermined threshold at step S318. If the prediction error is below the predetermined threshold and / or the error or loss associated with a prediction based on the validation data set is below its corresponding predetermined threshold and / or if the number of epochs has reached a predetermined number, the machine learning model is output at step S322. Otherwise, the process moves back to step S320 for further adjustment of the model parameters. This process is repeated until a sufficiently trained machine learning model is provided and the process ends at step S324. The iterative training of the machine learning model uses any appropriate number of epochs, learning rate and / or weight decay in order to optimise the machine learning model. In some implementations, additionally or alternatively, training of the machine learning model is terminated if no improvement in the validation set performance is achieved over any appropriate number of epochs.

[0101] In an implementation, the iterative training of the machine learning model may include cross-validation with a subset of the dataset. Alternatively, or additionally, any suitable process for refining the machine learning model may be implemented.

[0102] Once the machine learning model has been trained using an appropriate method, the trained machine learning model is output at step S322 and the process ends at step S324. Whilst the steps of process flow S300 are described in a particular order, it will be appreciated that in further implementations, the steps may be performed in a different order and / or concurrently, with additional, fewer and / or alternative steps in order to train a machine learning model to determine at least one property of the endometrium of a subject based on received image data associated with an endometrial biopsy specimen of the subject.

[0103] In an implementation, where the image data input into the machine learning model has been segmented to provide a plurality of regions, the machine learning model may be selectively trained on one or more regions of the plurality of regions. In such cases, selectively training the machine learning model may comprise determining a spatial distribution of values associated with the at least one determined property of each of the plurality of regions and selectively training the machine learning model on a subset of one or more of the plurality of regions based on the values. Each spatially distributed value may be associated with a node, such as a node of a graph neural network. Such training enables an improved model through focused model adjustment based on targeted subsets of one or more of the plurality of regions. In further implementations, additional or alternative selective training techniques are used in order to provide a machine learning model enabling the functionality described herein. The machine learning model that is trained in accordance with Figure 3 can subsequently be used in order to determine at least one property of the endometrium based on received image data associated with an endometrial biopsy specimen of a subject. As described with reference to Figure 2, processed image data may be input into a machine learning model at step S214 of process flow S200 in order to determine at least one property of the endometrium at step S216 subsequently provide an output at step S218.

[0104] Examples

[0105] An example of an implementation of the system of Figure 1 and methods described with reference to Figures 2 and 3 is now described in further detail with reference to Figures 4 to 13.

[0106] An example of datasets for construction and testing of a machine learning model is described with reference to Figure 4, which shows an example of a consort diagram S400 for refining such datasets. The datasets described with reference to Figure 4 may be used to train, validate and test a machine learning model, such as the machine learning model 112 of Figure 1, that is used to determine at least one property of the endometrium of a subject based on image data associated with an endometrial biopsy specimen from the subject. The datasets may be stored in the memory 108 of the computing device 102 described with reference to Figure 1, or at another suitable device, such as at one or more of the computing devices of the network 114.

[0107] In the example of Figure 4, in accordance with step S202 of Figure 2 and step S304 of Figure 3, endometrial biopsy specimens were obtained for a test set of 641 patients in a first step S402 within 5 to 12 days after a positive urinary ovulation test, this self-reported timing being designated as LH+. Whilst the use of patient data for samples for LH+ between 5 to 12 days was used, in further implementations, samples for any appropriate range of times may be used.

[0108] Once the endometrial biopsy specimens were obtained, each was processed in accordance with step S206 of Figure 2 and step S306 of Figure 3. The endometrial biopsy specimens may be obtained and processed sequentially and / or simultaneously. In the example of Figure 4, the endometrial biopsy specimens are each formalin-fixed paraffin-embedded (FFPE) biopsy tissue samples. The 641 tissue samples were subjected to staining for CD56, which serves as a marker for uterine natural killer (uNK) cells, and hematoxylin to visualise cell nuclei. Furthermore, RTq-PCR data was provided for the expressions of 13 individual genes along with corresponding EndoTime as calculated using the method described in Lipecki, J., Mitchell, A.E., Muter, J., Lucas, E.S., Makwana, K., Fishwick, K., Odendaal, J., Hawkes, A., Vrljicak, P., Brosens, J.J., Ott, S. : EndoTime: non-categorical timing estimates for luteal endometrium. Human Reproduction 37(4), 747-761 (2022). It will be appreciated that in further implementations, alternative or additional techniques may be used to prepare any suitable number of specimens. In order to determine whole slide image data in accordance with step S210 of Figure 2 and step S308 of Figure 3, for each of the processed samples of the 641 patients, bright-field images were obtained by scanning the samples at a spatial resolution of 0.50 microns per pixel (MPP). The whole slide image data from the 641 patients was then screened in order to exclude data associated with missing image or gene expression data, for example. With reference to the example of Figure 4, at step S404, 139 samples were excluded due to lack of a usable whole slide image. A further 8 samples were removed at step S406 due to sample quality issues and a further 38 were removed at step S408 due to missing labels. The resultant 456 samples were output at step S410 of Figure 4 for use as training data.

[0109] Similarly, patient data for 167 patients were gathered for independent testing at step S416. This initial test set of patient data was reduced following steps S418, S420, S422 to exclude samples for which there was no whole slide image, or where the sample had quality issues, and / or was missing a label in an analogous process as the one described for refining the 641 samples of patient data for the training data. Following the exclusion of 9 samples, 159 independent test samples were provided in a test set at step S424.

[0110] Whilst Figure 4 describes an implementation including a particular number of patient samples and techniques for sample preparation, processing and imaging, it will be appreciated that in further examples of the implementation of Figure 3, any appropriate number of patient samples and techniques for sample preparation, processing and imaging may be used in order to provide an appropriate dataset comprising training and test data for implementing the functionality described herein.

[0111] In the example of the 641 training and validations samples of the dataset described with reference to Figure 4, the majority of whole slide images contain control tissue from the small bowel. The whole slide image data was processed in accordance with steps S310 and 312 described with reference to Figure 3 in order to provide patch image data for input into a machine learning model. To accurately process the correct tissue and concurrently eliminate any artefacts such as staining errors, a user-defined area selection approach was used to identify the boundaries of the desired endometrial tissue for processing. Within the defined area, a threshold of 250 was applied on the average pixel intensity to effectively segregate the tissue region from the background, though other thresholds may be used. Due to the excessively large size of multi-gigapixel histopathology images, direct utilization with standard image analysis methods is impractical. However, the segmentation of the whole slide images into patches measuring 224 x 224 pixels at an image resolution of 0.50 MPP provides a practical way of processing the data. In further implementations, different patch sizes and image resolutions may be used. In the example of the implementation at Figure 4, in order to ensure the inclusion of informative tissue areas and discard non-relevant regions, patches with less than 5% of informative tissue area, determined by pixels with a mean intensity greater than 220, were excluded. In further implementations, alternative or different thresholds may be used in order to determine relevant regions. The remaining patches were then used as the basis for further analysis.

[0112] In the example implementation of Figure 4, the refined 456 processed samples output for training and validation data at step S410 were used in the development of a graph neural network machine learning model as described with reference to Figure 5. For each whole slide image associated with the refined dataset, a graph was created from its respective patches that are extracted as described with reference to step S312 of Figure 3. A process flow S502 for forming graph representations is shown at Figure 5. At Figure 5, the process commences at step S502, where the extracted patches are obtained for an endometrial biopsy specimen. The process moves to step S504, where feature extraction is performed. A graph can be written as G=(V, E), where V = {vj\i=l..n} is the set of vertices, which in the case of the example of Figure 4, are the patches of a whole slide image, and E is the edge set defining the connectivity between the nodes. All nodes can be expressed as v, = (gi, hi) where gtrepresents the spatial coordinates and hi is the feature representation of a patch in the whole slide image. Accordingly, at step S504, for each 224 x 224 patch the features are taken from the penultimate fully connected output layer of ShuffleNet V.2, trained on ImageNet, giving a 1024-dimensional feature vector hi. The set of edges E is computed by connecting adjacent patches, with a distance of less than 4000 pixels, using Delaunay triangulation. The process them moves to step S506, where a graph representation is provided in the form of a planar graph such that if is connected to Vj then there will be an edge in the edge set G E. The graph representation may serve as an input into a graph neural network.

[0113] Whilst the formation of graph representations from patches is described with reference to particular processes, data and values, in further implementations, alternative or additional processes, data and values may be used.

[0114] In the example of the dataset described with reference to Figure 4, the data was used to provide graph representations that are input into a graph neural network in accordance with step S314 of Figure 3. In the example of Figure 4, graph representations were obtained through the extraction and processing of patch image data, as described with reference to Figure 5. However, graph representations obtained through other methods may be used. The graph representations of the whole slide images were then passed into a graph neural network (GNN) for node and whole slide image level predictions of gene expression values, EndoTime, and LH+ simultaneously as part of the model development at step S412 of Figure 4. In further implementations, alternative or additional predictions may be provided. Model development at step S412 may be evaluated during or after training a machine learning model, for example by evaluating Spearman correlations at step S414 of Figure 4.

[0115] In order to train a machine learning model based on the dataset described with reference to the example of Figure 4, a graph neural network architecture was used such that node feature representations were passed through EdgeConv layers L={1,2,3}. Each EdgeConv layer updates all graph nodes feature representations by aggregating adjoining nodes to create a new node level feature representation for each layer. For node vmwith neighbours S(m) and a feature representation hmat layer I the new embedding of the EdgeConv operation can be mathematically expressed as follows: where Hlis a neural network. Using GNNs and EdgeConv layers allows patches to share information with neighbouring patches, letting more spatial context be considered in the prediction. By repeatedly applying EdgeConv layers in the GNN architecture, the model can capture structure and connectivity patterns of nodes in the graph structure from a wide neighbourhood, enabling effective information propagation and feature learning in the graphs. To get a node level prediction score for a node vm, the feature representation hmlis passed into a multi-layer perceptron which is then aggregated over all layers L to get a final node prediction score for all target variables.

[0116] The whole slide image level score is then obtained by pooling and aggregating all node level prediction scores in a whole slide image, thereby providing an output prediction in accordance with step S316 of Figure 3.

[0117] Both EdgeConv layers and node-level multi-layer perceptrons trainable parameters are learned using backpropagation. For batch size N, the prediction score for k={l..K} target variables are compared with the ground truth using the pairwise ranking loss:

[0118] Where Pk= {(a, 6) \y > ya> a, b = 1 .... N] is the set of all patients in the batch where the value of the target score k is larger for a than b. Minimisation of the loss function L will rank each patient for each variable. For example, a patient with EndoTime of day 10, should always be ranked higher than a patient with EndoTime of day 6.

[0119] In the example of the dataset obtained through refinement with reference to Figure 4, in order to provide a comprehensive evaluation of the model’s performance, 5-fold cross-validation was performed on the images. This uses 5 mutually exclusive sets of 20% of the data for testing, with the other 80% of data in each fold used for training.

[0120] In the training data, 10% was used for validation and parameter optimisation. The model was trained over 300 epochs using a learning rate of 0.001 and weight decay of 0.0001. During each epoch, the training set was divided into batches of size 16, and the learnable parameters were updated using the Adam optimizer. To prevent overfitting, early stopping was implemented by monitoring the validation set performance. If there was no improvement for 20 consecutive epochs, training was stopped. By following these steps, the graph neural network model was improved through the comparison with a threshold and adjustment of model parameters as described with reference to steps S318 and S320 of Figure 3.

[0121] Whilst machine learning model trained using the example dataset of Figure 4 was adjusted using particular epochs, learning rate, weight decay and other parameters, it will be appreciated that in further implementations, the machine learning model may be trained by adjusting any suitable parameter by any suitable amount and / or frequency.

[0122] In the example of the dataset of Figure 4, from each epoch trained model, the 10 bestperforming models of the validation set are selected to ensemble. The prediction scores from these models are then averaged for test set inference. At step S426 of Figure 4, for final independent test set inference, the same method is performed with all training data to generate final prediction scores for each instance.

[0123] Once a machine learning model has been trained, it may be used to provide predictions and the performance may be evaluated. In the example of Figure 4, since a ranking loss was used, the prediction scores were not in the same range as the EndoTime ground truth. To convert these into interpretable predictions the values were scaled to put them in the same range as original EndoTime scores. At step S428 of Figure 4, the predictive performance of this model is reported across each fold, the statistical summary of the evaluation metrics for all prediction targets is presented, specifically the mean and standard deviation of the Spearman’s correlation coefficient which quantifies the relationship between the predicted scores and the ground truth. To assess the statistical significance of the findings from the different folds, the mean of the p- values associated with these obtained Spearman’s correlations are also obtained. Furthermore, the R2score for the EndoTime prediction is included as an additional performance to measure to provide further insight into the accuracy of the predictions. By employing rigorous performance evaluation techniques, a robust and comprehensive assessment of the quality of the model is provided.

[0124] In the example of the data obtained with reference to Figure 4, the independent test data may be tested at step S426 by following the process S600 described with reference to Figure 6.

[0125] Figure 6 illustrates a process flow S600 of using a graph neural network to determine one or more properties of the endometrium of a subject based on image data. At step S602, appropriate data is obtained, for example in accordance with step S424 of Figure 4. Image data may be obtained and processed in the manner described above with reference to steps S304 to 312 of Figure 3. For each sample for which the machine learning model is being tested, a graph representation may be constructed at step S604 in the manner described with reference to Figure 5. The graph representation is then input into the machine learning model at step S606. The trained machine learning model then outputs prediction scores at step S608. Graph explanations are determined at step S610 and node level scores are output at step S612.

[0126] The process flow S600 described with reference to Figure 6 may be used to output data that can be analysed in order to evaluate the performance of a machine learning model that is used to determine at least one property of the endometrium based on received image data associated with an endometrial biopsy specimen of a subject. For example, the training and validation dataset consisting of 456 patients and the independent test data set consisting of 159 patients, described with reference to Figure 4 may be used to analyse the performance of a graph neural network machine learning model that has been trained as described herein with reference to Figures 3. In such cases, the trained machine learning model is trained and used to predict timing estimates of the luteal endometrium. In an example of the implementation described herein, the machine learning model determines gene expression for 13 genes, including the 6 genes used in EndoTime, as well as a ratio of two genes and LH+. The determinations are made through the prediction of one or more properties with an associated confidence. Once these determinations have been made, the results may be analysed. In further implementations, a machine learning model may determine any appropriate property of the endometrium of a subject, as described with reference to Figure 2.

[0127] Examples of results obtained using the data described with reference to Figure 4 are shown at Figures 7 to 13. Evaluation of the results is provided in Table 1, which shows Spearman Correlation results for determined properties of the endometrium of a subject based on a machine learning model comprising a graph neural network that has been trained using the data referred to at Figure 4.

[0128] Figure 7 shows a visualisation of gene expression and LH+ prediction versus ground truth to the nearest day. In Figure 7, a heatmap 700 of log-transformed average gene expression for individual genes, ratio of genes and average LH+ across the luteal phase is presented.

[0129] Along the x-axis 702 columns 704 to 718 are denoted, which are associated with ground truth, represented to the nearest day based on EndoTime data, from day 5 in the first column 704 on the left-hand side to day 12 in the last column 718 on the right hand-side.

[0130] Along the y-axis, a first row 722 shows a prediction by the machine learning model of average LH+ based on received image data associated with endometrial biopsy specimen from a subject. The result shows that the average LH+ increases from day 5 in the first column 704 to day 12 in the last column 718.

[0131] A second row 724 illustrates a prediction of the GPX3 / SCL15A2 ratio based on received image data associated with endometrial biopsy specimen from a subject. The ratio increases from day 5 at the first column 704 on the left to the day 12 at the last column 718 on the righthand side. This particular ratio has been employed as an inference tool for gene-derived timing in previous studies. As shown in Table 1, there is a strong correlation with ground truth for the prediction of the GPX3 / SCL15A2 ratio. This highlights the significant influence of the ratio of these two genes on the endometrial phenotype and the substantial changes that occur during the luteal phase, thereby emphasizing their potential as important determinant in the biological processes underlying endometrial development and function.

[0132] Log transformed gene expressions for six different genes are shown from the third row 726 along the y-axis 720 to the eighth row 736 along the y-axis. For each of the six genes, the log transformed gene expression value increases from day 5 in the first column 704 on the left to day 12 in the last column 718 on the right 718. The third row 726 shows the log transformed gene expression value for SLC15A2, the fourth row 728 shows the log transformed gene expression value for II ARP.jthe fifth row 730 shows the log transformed gene expression value for IGFRP1, the sixth row 732 shows the log transformed gene expression value for GPX3, the seventh row 734 shows the log transformed gene expression value for CXCL14 and the eighth row 736 shows the log transformed gene expression value for DPP4. As shown at Table 1, there is a strong correlation between the predicted gene expression level and ground truth for each of the genes.

[0133] Training a machine learning model in accordance with the methods described herein enables the effective prediction of endometrial timing estimates directly from histology images. In the example of the graph neural network described based on the datasets obtained with reference to Figure 4, deep features may be extracted from patches using transfer learning with a pretrained network and graphs constructed to represent whole slide images. The use of a graph neural network machine learning model enables accurate predictions directly from the whole slide images. In particular, the machine learning model is able to accurately predict gene-based timing estimates and associated gene expressions from histology images.

[0134] Figure 8 shows a visualisation of image-based predictions of EndoTime versus ground truth to the nearest day. That is, Figure 8 shows a plot 800 of EndoTime prediction based on received images associated with an endometrial biopsy specimen from a subject along the y-axis 816 versus ground truth in respect of EndoTime rounded to the nearest day. Box plots 802, 804, 806, 808, 810, 812, 814 are shown for each day from day 5 at the first box plot 802 on the lefthand side, to day 12 in the last box plot 816 on the right-hand side. It can be seen that the correlation between the image-based prediction and the ground truth is strong until approximately day 9. The plot 800 illustrates that the performance of a graph neural network machine learning model in accurately predicting EndoTime, which represents the gene expression-based timing of the luteal endometrium. The results of 5-fold cross-validation on the training data set generated at step S410 of Figure 4 is shown at Table 3.

[0135] Table 3

[0136] There, a strong and statistically significant association with a mean Spearman’s correlation coefficient of 0.778 (standard deviation = 0.047). This finding indicates robust and consistent relationship between the predictions and the actual timing estimates. Statistical analysis further confirms the significance of these predictions, with a mean p-value of less than IO’15. This shows that there is a significant relationship between tissue morphology and genebased timing estimates.

[0137] The ability of the machine learning model to accurately predict gene-based timing estimates and associated gene expressions from histology image is shown through rigorous cross-validation, thereby ensuring the robustness of the results. The consistently high performance achieved by the predictive models across different folds of the cross-validation as well as independent testing supports the generalizability of the model and indicates that the whole slide images associated with endometrial biopsy specimens contain valuable information that reflects gene expression patterns in the luteal endometrium. Analysis of individual genes shows that the top performing genes (DPP4, GPX3, and SLC15A2') are all associated with epithelial glandular expression. The ratio between GPX3 and SLC15A2 is shown to be the highest performing predicted variable, underscoring the substantial influence of those genes on endometrial tissue.

[0138] As shown at Table 3, the superior performance of the machine learning model in predicting EndoTime, based on incorporating all genes, compared with the performance of the individual component genes shows that information from multiple genes is integrated to provide enhanced predictive power and providing a comprehensive assessment of all gene expression patterns rather than relying on a single gene for prediction. Additionally, uNK cells are known to increase over time, which is clearly reflected in a significant increase above 10 in the representative patches described below, thereby further validating the machine learning model.

[0139] The approach described herein provides an efficient way to handle large datasets efficiently, contrasting with traditional methods for analysing histology images which often involve manual inspection and subjective scoring by pathologists, which can be time consuming and prone to variability. Surprisingly, it has been found that accurate timings can be provided based on gene expressions without the use of expensive gene expression kits. Accordingly, quick objective predictions are provided.

[0140] One advantage of using graph neural networks is their ability to generate node-level prediction scores before aggregation into a whole slide image prediction. This characterisation adds a layer of interpretability to the predictions, as it allows the identification of specific areas within the whole slide image that contribute positively or negatively to the final prediction. To visualise this contribution, an overlay representation is used by colorising patches according to their node-level scores, indicating high and low predictions.

[0141] In order to find representative patches for analysis, the image-based prediction of EndoTime may be compared with the ground truth and for each rounded EndoTime day, the top 10 most accurate whole slide images are chosen to select patches from. In cases where fewer than 10 whole slide images are available for a particular day, the node level scores are compared to the top scoring whole slide images. The top 10 patches that displayed the highest similarity to the whole slide image level score are then chosen. The features for the 10 patches for all 10 top scoring whole slide images are then clustered into 5 groups using k-medoids with the centroid of each cluster being chosen as a representative patch for visualisation. This aimed to visualize the most relevant and representative patches for each day to visualise how these patterns change over the luteal phase. Histological features of the images can then be assessed to interpret any patterns in the model’s predictions.

[0142] In order to visualise the spatial arrangements of node contributions, overlays may be created from all whole slide images. In order to create overlays from all whole slide images, firstly, all node-level scores are subjected to normalization to a range of 0 to 1. Any outside of this range are set to the nearest integer in the range. This normalization process uses a min-max scaler with the lowest and highest patch scores derived from the highest and lowest wholes slide image level score for the most accurately predicted whole slide image in days 5 and 12 in the dataset. Subsequently, the resulting scores are visualised through the TIAtoolbox Bokeh application. Leveraging this application, visual representations are created that allowed the effective analysis and interpretation of normalized patch scores of a whole slide image.

[0143] Figure 9 shows a visualisation of whole slide image node level predictions versus ground truth to the nearest day. Representative whole slide images with overlays are shown in three rows along the y-axis 904 for each ground truth day from day 5 to day 12 along the x-axis. The three rows illustrate the top scoring whole slide images for each day, with an increase from lower scoring regions on the left-hand side to higher scoring regions on the right-hand side, along the direction of the x-axis 902. For visualisation, the scores were normalised across all patch-level scores and scaled within a range of 0 to 1, with 0 and 1 being the wholes slide image level score for the most accurately predicted day 5 and day 12 whole slide image respectively. Consequently, the heatmap scale represents the sign and magnitude of the node-level scores. The observations shown at Figure 9 show that as the timing estimates increase along the x-axis 902, the patches generally exhibit a shift from lower scoring to higher scoring. This aligns with the fact that slide-level scores are aggregations of the patch-level scores. There are also differences in the heterogeneity of the patches within the different days, with those in days 7 to 10 being more heterogeneous compared with those at the extreme ends of the data. This suggest that the effect of the changes in gene expression is not uniform across the tissue and may affect different areas at different rates.

[0144] Figure 10 shows a scatter plot 1000 of predicted EndoTime timing estimates along the y- axis 1004 versus ground truth along the x-axis 1002. As shown by the mean average trace 1006, there is a strong correlation (R2= 0.692). This highlights the strength and accuracy of the machine learning model-based approach in predicting EndoTime, indicating the potential for use as a reliable tool for timing estimation in luteal endometrium histology images. After day 9, there is some decrease in performance, which can be attributed to the stabilisation of gene expression values at this time, leading to loss of temporal information within the tissue structure.

[0145] Figure 11 shows representative patches extracted from whole slide image data, collected as described with reference to Figure 4, for each different EndoTime day. The plot 1100 of Figure 11 shows the five most representative patches for each day, from day five 1102 to day twelve 1116. By selecting a subset of patches that are most representative of each day, notable trends and patterns can be observed. For example, for whole slide images, there is usually an increase in uterine natural killer (uNK) cell activity over time. This correlation underscores the importance of uNK cells in the physiological changes that occur during the luteal phase in establishing a receptive environment for implantation. The increase in abundance of uNK cells within the tissues is reflected by an increment in the proportion of CD56 staining, as seen illustrated in Figure 11.

[0146] Figure 12 shows a heatmap visualisation 1200 of each whole slide image in the training set, collected as described with reference to Figure 4, for patch level scores during cross fold validation for different days, from day 5 1202 to day 12 1216. The heatmap visualisation 1200 shows that the predicted scores for patches closely matches that of the actual day for the patient.

[0147] The use of a machine learning model described herein in order to determine at least one property of the endometrium of a subject based on image data associated with an endometrial biopsy specimen from the subject may be used to determine the risk of pregnancy loss on a particular day in the endometrial cycle. Recurrent pregnancy loss poses a significant reproductive health challenge, with a documented stepwise increase in risk following each successive pregnancy loss, a phenomenon independent of age and other clinical variables. It has also been shown that there is a similar stepwise increase in the ratio of PLA2G2A and DIO2 gene expressions following each pregnancy loss which may provide an explanation for the increase in risk. This ratio remains stable across menstrual cycles, indicating its reliability as a marker of underlying endometrial dysfunction and a predictor of future inadequate decidual responses. The PLA2G2A / DIO2 gene expression ratio is included in an implementation of the machine learning model and has been shown to demonstrate a strong Spearman’s correlation of 0.62 (SD = 0.035, P < io9).

[0148] In order to assess the feasibility of inferring pregnancy loss risk directly from endometrial histology images, the analysis was extended to an additional 1935 whole slide images. These images did not include gene expression values, rendering them unsuitable for training or testing. However, clinical variable such as the number of previous pregnancy losses were available for evaluation. The graph neural network machine learning model developed based on the data collected with reference to Figure 4 was applied, thereby enabling gene ratio values to be predicted. Accordingly, the association between the predicted PLA2G2A / DIO2 gene expression ratio and the risk of pregnancy loss may be obtained.

[0149] Day-normalized predicted PLA2G2A / DIO2 ratios demonstrated a strong association with miscarriage risk. Each one percentage point increase in the ratio corresponded to a 0.48% reduction in expected miscarriage counts (IRR = 0.9952, p = 0.002), replicating the ground-truth ratio effect (IRR= 0.9947, p < 0.001). Boxplot analyses revealed a stepwise decline in the scaled PLA2G2A / DIO2 ratio across groups stratified by increasing previous pregnancy loss counts. Nonparametric Mann-Whitney U tests confirmed that the group with more than five miscarriages exhibited significantly lower percentile predictor values compared to the zeromiscarriage group (p < 0.001). Bayesian estimation also produced a consistent median effect (IRR = 0.9979, 90% CrI = 0.997-0.999) after adjusting for age and BMI, supporting the predicted PLA2G2A / DIO2 ratio as an independent predictor of miscarriage risk.

[0150] Figure 13A shows the proportion (along the y-axis 1304A) of patients falling above the 25thpercentile 1306 A and below the 75thpercentile 1308 A, for patients with a different number of pregnancy losses (from 0 to 5+, along the x axis 1302A) for predicted gene expression ratio values. Figure 13B shows the proportion 1304B of patients falling above the 25thpercentile 1306B and below the 75thpercentile 1308B for patients with a different number of pregnancy losses (from 0 to 5+ along the x-axis 1302B) for measured gene expression ratio values.

[0151] As shown at Figure 13 A, there is a consistent stepwise change in the gene expression ratio following each pregnancy loss, aligning with the observed recurrence risk patterns as shown at Figure 13B. Patients with a low number of pregnancy losses are more likely to have a ratio value fall above the 75thpercentile of ratio scores and less likely to fall in the 25thpercentile. Inversely, those with 5+ pregnancy losses show the opposite, with a stepwise decrease and increase for those bins respectively for intermediate values. This outcome not only underscores the impact of pregnancy loss on the endometrium, specifically on the PLA2G2A / DIO2 gene expression ratio, but also highlights the utility of the machine learning model in predicting the PLA2G2A / DIO2 gene expression ratio. Accordingly, the predictive capability of the machine learning model provides a tool to offer a convenient proxy for evaluating pregnancy loss risk, eliminating the necessity for direct gene expression testing. This establishes the feasibility of quantifying miscarriage risk directly through image-based analysis, offering a practical and scalable alternative to conventional molecular assays. As a consequence, targeted interventions and personalized reproductive health strategies may be implemented.

[0152] As demonstrated in connection with Figure 13, the machine learning model can determine properties of the endometrium of the subject such as gene expression levels. At least some of the properties determined by the model, such as the PLA2G2A / DIO2 gene expression ratio discussed above, may be known biomarkers associated with clinical conditions such as risk of pregnancy loss. This allows the model to also determine further properties of the endometrium such as a risk of pregnancy loss based on the known association between the properties and the clinical condition. However, the machine learning model may also be used as a discovery tool to identify novel biomarkers, such as histological biomarkers, associated with clinical conditions.

[0153] For example, identifying patients at risk of recurrent pregnancy loss (RPL) or recurrent missed miscarriage (RMM) remains a critical clinical challenge. In particular, many RMM events occur asymptomatically and are often detected only through routine ultrasound. RPL is defined as two or more lost pregnancies, while RMM is defined as three or more consecutive pregnancy losses. Despite differing thresholds (2 vs 3), both definitions describe the same underlying disorder (repeated spontaneous loss of pregnancy prior to viability).

[0154] To explore potential histological markers associated with RMM, we performed Bayesian logistic regression analyses focusing on glandular and cellular features determined from endometrial biopsy specimens using the machine learning model.

[0155] Quantitative assessment of tissue-level features in whole-slide images (WSIs) revealed a significant reduction in relative glandular area in 46 patients with RMM compared to 1,398 controls during days 9 to 11 post-LH surge (t-test, p < 0.001). Figure 14 shows boxplots of day- normalized feature values on days 8 to 11 comparing recurrent missed miscarriage (RMM) and control patients, showing significant differences between groups. Three representative histological features are displayed with corresponding p-values. Adjacent to the boxplots are example slides illustrating typical characteristics for each patient group. On average, day normalised glandular proportion was reduced by 0.23 in the RMM group, indicating considerable depletion of glandular structures in affected individuals. Additionally, we observed significant differences in day normalised uNK cell proportion and gland solidity between cases and controls, suggestive of impaired endometrial receptivity associated with RMM.

[0156] These results highlight clear distinctions in histological features between RMM and control patients, demonstrating that measurable tissue-based differences exist across groups and supporting the clinical relevance of these biomarkers for further investigation. Notably, while the endometrial glands play a central role in uterine function, there are no previously described clinical disorders that are specific to endometrial glands themselves. Thus, our findings identify distinctive tissue features associated with both miscarriage and missed miscarriage, supporting their pathophysiological relevance and potential for clinical application.

[0157] Figure 15 shows on the left a heatmap depicting the correlation between the glandular characteristics discussed above and predicted gene expression variables. Rows represent model features, grouped using k-means clustering, while columns correspond to gene expression variables, clustered by functional similarity. Each square indicates the strength and direction of correlation according to the scale at the bottom of the figure, facilitating identification of feature groups most strongly associated with each prediction. Representative examples illustrate temporal changes in gland solidity for different predicted EndoTime days. Given that levels of expression of some genes are known to correlate with risk of pregnancy loss, the associations suggest that the histological characteristics can also be predictive of pregnancy loss.

[0158] Extending to predicted gene expression, univariate Bayesian modelling identified all genes as showing an effect SLC15A2 expression increased risk (median IRR = 1.002, 90% CrI = 1.001-1.003), whereas SCARA5 was strongly protective (median IRR= 0.997, 90%CrI = 0.996- 0.998). Further, multivariate analysis of averaged expression profiles stratified by gene function (glandular, stromal, and immune) consistently demonstrated that elevated immune gene expression correlates with decreased miscarriage burden across these key endometrial compartments. Notably, ground truth gene expression measurements did not achieve statistical significance in this context. By contrast, predictions generated by our model using histological image features yielded significant associations, with narrower credible intervals. This disparity highlights the advantage of model-based prediction, particularly when powered by large-scale data, in enhancing signal detection emphasizing the strength of image-based prediction for mapping molecular drivers of miscarriage risk.

[0159] These molecular signatures corresponded to and were reinforced by morphological predictors. Glands with greater solidity, eccentricity, tissue gland proportion displayed reductions in miscarriage risk suggesting a transcriptional programme which creates a tissue environment where protective glandular architecture is maintained. For example, high eccentricity is correlated with higher predicted GPX3 expression (Figure 15). At the cellular level, uterine NK cell proportion and mean cell size were also protective factors in risk of miscarriage. Identifying these characteristics enhances understanding of the biological mechanisms underlying pregnancy loss risk and may support the development of targeted interventions to improve reproductive outcomes. Detailed visualizations of these effects are presented in Figure 16.

[0160] As further illustrated in Figure 1 (c), higher gland solidity was found to be significantly protective against RMM with an odds ratio (OR) of 0.645 (90% CrI 0.519-0.902). Similarly, a higher tissue gland proportion was associated with decreased odds of RMM (OR=0.780; 90% CrI 0.636-0.944). Both predictors demonstrated strong evidence for real association, supporting their potential as histological biomarkers protective against recurrent missed miscarriage.

[0161] A supervised Random Forest classifier was trained on endometrial histological profiles incorporating gland and cell features demonstrated reasonable discrimination for the identification of RMM patients. During cross-validation, the model achieved an AUROC of 0.708 ± 0.041, effectively distinguishing 105 RMM patients from 3,238 non-RMM controls. Figure 17(a) shows a boxplot of Random Forest prediction scores between RMM and control patients, showing significant differences.

[0162] Figure 1 (b) shows the average feature importance measured by SHAP values across five stratified cross validation folds using the Random Forest classifier with optimal hyperparameters. Feature importance represents the mean absolute SHAP value aggregated over all test samples from each fold, reflecting the average contribution of each feature to model predictions. The features are ordered by descending importance, highlighting the most influential variables for the classification task. The SHAP analysis highlighted the relative importance of features contributing to model predictions, showing uNK cell proportion, gland solidity, and tissue gland proportion as important predictors, reinforcing their strong role in miscarriage risk.

[0163] Taken together, these results show the potential of interpretable machine learning approaches to identify RMM at-risk patients directly from histological images. They also demonstrate the utility of the machine learning model and associated methods disclosed herein for identifying novel histological and other biomarkers that may be associated with clinical conditions such as RMM.

[0164] In connection with this, the method of determining at least one property of the endometrium may comprise receiving image data associated with an endometrial biopsy specimen from a plurality of subjects, the plurality of subjects associated with two or more clinical classes. The method may comprise using the machine learning model to determine the at least one property of the endometrium for each of the plurality of subjects based on the image data. The method may further comprise identifying a difference in the at least one property between subjects associated with different clinical classes. The clinical classes may include attributes such as age or ethnicity, or may include pathologies such as recurrent missed miscarriage or recurrent pregnancy loss.

[0165] Further, these results support the use of histological and morphological features for the assessment of risk of pregnancy loss. Consequently, the present disclosure also provides a method of determining a risk of pregnancy loss. The risk of pregnancy loss may comprise one or more of i) a risk of pregnancy loss on a particular day in the endometrial cycle, ii) a risk of recurrent pregnancy loss, or iii) a risk of recurrent missed miscarriage.

[0166] The method comprises receiving image data associated with an endometrial biopsy specimen from a subject. The image data and the endometrial biopsy specimen may be substantially similar or the same as described for the methods described above.

[0167] The method further comprises measuring one or more attributes of histological constructs in the image data. The histological constructs may comprise any of those discussed above, such as one or more of cells, glands, lumen, luminal epithelium, subluminal stroma and spiral arteries.

[0168] The one or more attributes may comprise any of those discussed above. These may comprise colour or morphological attributes, for example one or more of: a mean, maximum, median, or standard deviation of a predetermined colour channel within the histological construct; an area, solidity, eccentricity, circumference, equivalent diameter, major or minor axis length, and convex area of the histological construct.

[0169] The predetermined colour channel may be a colour channel corresponding to a stain used to stain the endometrial biopsy specimen, such as one of the stains discussed above. In particular, the stain may be CD56 or hematoxylin.

[0170] The one or more attributes may also include attributes derived from the colour or morphological attributes. For example, in an exemplary implementation, the histological constructs may comprise cells and the one or more attributes may comprise one or more of a mean, maximum, median, or standard deviation of a colour channel corresponding to a CD56 stain within the histological construct. As discussed above, CD56 serves as a marker for uterine natural killed cells. Consequently, the method may further comprise determining a proportion of uterine natural killer cells in the biopsy specimen based on the one or more attributes, i.e. based on the mean, maximum, median, or standard deviation of the colour channel corresponding to CD56. In this case, determining the risk of pregnancy loss based on the one or more attributes may comprise determining the risk of pregnancy loss based on the proportion of uterine natural killer cells.

[0171] Measuring the one or more attributes of the histological constructs in the image data may comprise segmenting the image data into a plurality of regions, each region corresponding to a histological construct, and measuring the one or more attributes from the plurality of regions. The attribute may be that of individual regions corresponding to particular types of histological construct. Alternatively or additionally, the attribute may be a collective attribute, for example representing an average (such as mean or median) or the properties of multiple regions corresponding to particular types of histological construct.

[0172] The method further comprises determining the risk of pregnancy loss based on the one or more attributes. This is achieved based on the known correlations between the attributes and risk of pregnancy loss, for example as illustrated in Figure 16.

[0173] We also investigated disrupted stromal synchrony in recurrent miscarriage. To evaluate differences in the temporal regulation of gene expression in the endometrium, we compared Spearman correlations between the timing of the luteinizing hormone surge (LH+) and model- predicted gene expression profiles for patients with recurrent missed miscarriage and non-RMM controls. We performed 1,000 bootstrap resampling iterations to generate distributions of correlation coefficients for each group providing robust group-level estimates.

[0174] Differences between groups were assessed by comparing the means of these distributions and using a two-sample t-test, which identified statistically significant divergence for all genes.

[0175] Figure 18 shows the correlations of prediction scores with self-reported LH+ day for genes with the largest group differences, bootstrapped to generate distributions for recurrent missed miscarriage (RMM) and control patients. The difference in means between groups is shown with a two tailed t-test p-value. Displayed are gene expression profiles for the three genes with the largest positive and largest negative differences. Glandular genes exhibit increased coordination in RMM patients, while stromal genes show disrupted regulation.

[0176] Figure 19 shows the correlations of prediction scores with LH+ day for all other genes predicted by the machine learning model, bootstrapped to give distributions for RMM patients vs control. Difference in means between the groups are shown along with t-test p-value. 7 predicted gene expressions with smallest absolute difference in means are shown.

[0177] Examination of gene-level results revealed a pronounced compartment-specific pattern. Genes showing positive shifts in absolute correlation, indicating stronger temporal regulation with respect to LH+, were predominantly glandular, including GPX3 and CXCL14. This suggests that phase-dependent coordination is preserved or even accentuated in glandular tissue. In contrast, genes with reduced or negative correlation differences reflecting diminished regulatory synchrony, corresponded primarily to stromal and uNK cell markers, such as IL15, SCARA5, and TIMP3. These findings imply that recurrent missed miscarriage is associated with a distinct molecular profile a highly co-ordinated glandular gene expression programme across the cycle, coupled with disrupted stromal regulation.

Claims

CLAIMS:

1. A method of determining at least one property of the endometrium, the method comprising: receiving image data associated with an endometrial biopsy specimen from a subject; inputting the data into a machine learning model; and using the machine learning model to determine the at least one property of the endometrium of the subject based on the image data.

2. The method according to claim 1, wherein the at least one property comprises at least one of: a value associated with readiness to implantation and / or decidualisation and / or healthiness.

3. The method according to any preceding claim, wherein the at least one property is associated with a time within a luteal phase of the subject.

4. The method according any preceding claim, wherein the at least one property comprises a relative time in the menstrual cycle of the endometrial biopsy specimen.

5. The method according to any preceding claim, wherein the at least one property comprises a gene expression level and / or protein expression level.

6. The method according to claim 5, wherein the gene expression level comprises at least one gene that allows identification of the day in the menstrual cycle.

7. The method according to claim 5, wherein the gene expression level is of at least one of IL2RB, IGFBP1, CXCL14, DPP4, GPX3 and SLC15A2.

8. The method according to any of claims 5 to 7, wherein the gene expression level comprises at least one marker gene for progesterone-dependent decidual cells, optionally selected from SLC15A2, DPP4, GPX3, IGFBP1, IL2RB, ITGAD, CXCL14, DIO2, PLA2G2A, TIMP3, IL15, TAGAP, SCARA5, IGF2 and WNT4 and / or at least one marker gene for progesterone-resistant stromal cells, optionally selected from 7)7(92, HOXAIO, H0XA11, WNT5A and IGF 7.

9. The method according to any preceding claim, wherein the at least one property comprises a ratio of gene expression levels.

10. The method according to claim 9, wherein the ratio of gene expression levels comprises a ratio between GPX3 and SLC15A2 gene expression levels and / or a ratio between PLA2G2A and DIO2 gene expression levels.

11. The method according to any preceding claim, wherein the at least one property comprises a prediction of the number of days following a luteinizing hormone surge or ovulation.

12. The method according to any preceding claim, wherein the at least one propertycomprises a risk of pregnancy loss, optionally wherein the risk comprises one or more of i) a risk of pregnancy loss on a particular day in the endometrial cycle, ii) a risk of recurrent pregnancy loss, or iii) a risk of recurrent missed miscarriage.

13. The method according to claim 12, wherein the risk of pregnancy loss is based on a ratio of gene expression levels associated with the biopsy specimen of the subject.

14. The method according to claim 13, wherein the ratio of gene expression levels comprises a ratio between PLA2G2A and DIO2 gene expression levels.

15. The method according to any preceding claim, wherein the image data comprises histologic whole slide image data.

16. The method according to any preceding claim, wherein the machine learning model has been trained using at least one of: one or more gene expression levels corresponding to a biopsy specimen of a subject; one or more histological constructs; one or more cellular descriptors; one or more textural descriptors and one or more image features.

17. A method of training a machine learning model for determining at least one property of the endometrium, the method comprising: receiving image data associated with an endometrial biopsy specimen of a subject, the received image data corresponding to at least one property of the endometrium; and training the machine learning model to determine at least one property of the endometrium based on the received image data.

18. The method according to claim 17, wherein the at least one property of the endometrium is associated with a time within a luteal phase of a subject.

19. The method according to claim 17 or 18, further comprising: processing the endometrial biopsy specimen of the subject; determining whole slide image data of the processed endometrial biopsy specimen of the subject; and determining one or more gene expression levels associated with the processed endometrial biopsy specimen of the subject.

20. The method according to any of claims 17 to 19, further comprising: processing the endometrial biopsy specimen of the subject to determine at least one of a cell count; a histological construct; one or more cellular descriptors and one or more textural descriptors.

21. The method according to any of claims 17 to 20, further comprising: segmenting the received image data into a plurality of regions; selectively training the machine learning model on one or more regions of the plurality of regions.

22. The method according to any of claims 17 to 21, wherein the endometrial biopsy specimen is obtained: by an endometrial biopsy catheter; by dilation and curettage; by biopsy forceps or curette during hysteroscopy; following hysterectomy or simultaneously with a procedure involving instrumentation of the uterus.

23. The method according to any preceding claim, wherein the machine learning model comprises a neural network, optionally wherein the neural network is a graph neural network.

24. The method according to any of claims 21 to 23, comprising: identifying a subset of one or more regions of the plurality of regions corresponding to a tissue; and selectively training the machine learning model on the subset of one or more regions of the plurality of regions, optionally wherein identifying the subset comprises determining one or more regions of the plurality of regions having a mean pixel intensity of at least a predetermined threshold.

25. The method according to any of claims 22 to 24, comprising: determining a spatial distribution of values associated with the at least one determined property of each of the plurality of regions; and selectively training the machine learning model on a subset of one or more regions of the plurality of regions based on the values.

26. The method of any of claims 17 to 25, wherein processing the endometrial biopsy specimen comprises identifying histological constructs, and the histological constructs comprise one or more of cells, glands, lumen, luminal epithelium, subluminal stroma and spiral arteries.

27. The method according to any of claims 1 to 16, wherein the machine model is trained in accordance with any of claims 17 to 26.

28. A machine learning model for determining at least one property of the endometrium, wherein the machine learning model is configured to: receive image data associated with an endometrial biopsy specimen from a subject; and determine the at least one property of the endometrium based on the received image data.

29. The machine model according to claim 28, wherein the machine model is trained in accordance with any of claims 17 to 26.

30. A method of determining a risk of pregnancy loss, the method comprising: receiving image data associated with an endometrial biopsy specimen from a subject; measuring one or more attributes of histological constructs in the image data; and determining the risk of pregnancy loss based on the one or more attributes.

31. The method of claim 30, wherein the one or more attributes comprise colour or morphological attributes, for example one or more of: a mean, maximum, median, or standard deviation of a predetermined colour channel within the histological construct; an area, solidity, eccentricity, circumference, equivalent diameter, major or minor axis length, and convex area of the histological constructs.

32. The method of claim 31, wherein the predetermined colour channel is a colour channel corresponding to a stain used to stain the endometrial biopsy specimen, optionally wherein the stain is CD56 or hematoxylin.

33. The method of any of claims 30 to 32, wherein the histological constructs comprise one or more of cells, glands, lumen, luminal epithelium, subluminal stroma and spiral arteries.

34. The method of any of claims 30 to 33, wherein: the histological constructs comprise cells; the one or more attributes comprise one or more of a mean, maximum, median, or standard deviation of a colour channel corresponding to a CD56 stain within the histological construct; the method further comprises determining a proportion of uterine natural killer cells in the biopsy specimen based on the one or more attributes; and determining the risk of pregnancy loss based on the one or more attributes comprises determining the risk of pregnancy loss based on the proportion of uterine natural killer cells.

35. The method of claim 34, wherein a higher proportion of uterine natural killer cells is associated with a lower risk of pregnancy loss.

36. The method of any of claims 30 to 35, wherein the histological constructs comprise glands and one or more of: a) the one or more attributes comprise solidity or eccentricity of the glands, optionally wherein a higher gland solidity or eccentricity is associated with a lower risk of pregnancy loss; b) the one or more attributes comprise a mean of a predetermined colour channel within the gland, optionally wherein i) a higher mean of the predetermined colour channel is associated with a higher risk of pregnancy loss; and / or ii) the predetermined colour channel corresponds to a CD56 or hematoxylin stain; and c) the one or more attributes comprise a mean area of the glands, optionally wherein a higher mean gland area is associated with a higher risk of pregnancy loss.

37. The method of any of claims 30 to 36, wherein the histological constructs comprise cells and one or more of: a) the one or more attributes comprise solidity or eccentricity of the cells, optionally wherein a higher cell solidity or eccentricity is associated with a lower risk of pregnancy loss;b) the one or more attributes comprise a mean of a predetermined colour channel within the cells, optionally wherein i) a higher mean of the predetermined colour channel is associated with a lower risk of pregnancy loss; and / or ii) the predetermined colour channel corresponds to a hematoxylin stain; and c) the one or more attributes comprise a mean area of the cells, optionally wherein a higher mean cell area is associated with a lower risk of pregnancy loss.

38. The method of any of claims 30 to 37, wherein the one or more attributes comprise a proportion of glands within tissue of the endometrial biopsy specimen, optionally wherein a higher proportion of glands is associated with a lower risk of pregnancy loss.

39. The method of any of claims 30 to 38, wherein measuring one or more attributes of the histological constructs in the image data comprises: segmenting the image data into a plurality of regions, each region corresponding to a histological construct; and measuring the one or more attributes from the plurality of regions.

40. The method of any of claims 30 to 39, wherein the risk of pregnancy loss comprises one or more of i) a risk of pregnancy loss on a particular day in the endometrial cycle, ii) a risk of recurrent pregnancy loss, or iii) a risk of recurrent missed miscarriage.

41. A computing device configured to perform the method of any of claims 1 to 27 or 30 to 40.

42. A computer program or computer readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method of any of claims 1 to 27 or 30 to 40.

43. A system for performing the method of any of claims 1 to 27 or 30 to 40, the system comprising: an imaging apparatus for imaging an endometrial biopsy specimen of a subject to obtain image data associated with the endometrial biopsy specimen; and a computing device configured to perform the method of any of claims 1 to 27 or 30 to40.