Reproductive health assessment
A computer-implemented method using supervised learning to generate targeted analyte panels and integrate machine learning models addresses inefficiencies in reproductive health assessments, reducing false positives and clinician time, enabling timely and cost-effective diagnoses.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-09
AI Technical Summary
Current reproductive health assessments are inefficient and costly, often leading to long waiting times and delayed diagnoses, which can result in infertility and the need for assisted reproductive technology, due to reliance on general practitioners and broad spectrum analyte testing with high false positive rates.
A computer-implemented method using supervised learning models to generate targeted analyte panels based on user health data, reducing false positives and clinician time by selecting relevant analytes from predefined sets, and integrating machine learning models to determine reproductive health conditions.
This approach reduces false positive rates, minimizes blood sample volume, and decreases clinician time, enabling earlier interventions and improved health outcomes by providing targeted and efficient reproductive health assessments.
Smart Images

Figure GB2025052146_09042026_PF_FP_ABST
Abstract
Description
[0001] REPRODUCTIVE HEALTH ASSESSMENT
[0002] Field of Invention
[0003] The invention relates to a computer implemented method of generating an analyte panel for a reproductive health assessment, a method of training a machine learning model for determining a probability or likelihood of a reproductive health condition, and computer implemented methods of reproductive health assessment, preferably based at least in part on analyte data according to the generated analyte panel.
[0004] Background
[0005] Reproductive health conditions can affect one in six females. Despite this high rate of occurrence, there is an increasing number of female reproductive health conditions not being addressed promptly, and some that in some cases go undiagnosed completely. Currently, most patients are reliant upon getting appointments with general practitioners in a primary care setting, who apply clinical guidelines and their own assessment based on the patient’s medical history, symptoms and blood test results in order to give a potential diagnosis. However, often specialist blood tests or investigations are required which leads to referral to secondary care, such as gynaecology clinics. This naturally creates a bottleneck in the healthcare process, meaning patients typically must wait long periods of time for an appointment, potential diagnosis and access to relevant information or care pathways. For example, as of 2024, there are approximately 592,000 people on the UK National Health Service gynaecology waiting list, marking a 171 % increase over the past 5 years. Meanwhile, late or delayed diagnosis can in some cases result in infertility and / or the need for assisted reproductive technology. The escalating burden on healthcare services underscores a critical need for more efficient reproductive healthcare interventions. In addition to the potentially long waiting times, obtaining a fertility health assessment can be costly both in terms of clinician time and tests. There are a number of different reproductive health conditions that can affect fertility, and the cost and time required to test for all of these can be prohibitive.
[0006] Summary
[0007] Aspects and examples of the invention are set out in the claims and aim to address at least a part of the above-described technical problem, and other problems.
[0008] In one aspect of the invention there is provided a computer implemented method of generating an analyte panel for a reproductive health assessment comprising: receiving user health data preferably including a plurality of medical history and symptom variables; determining, preferably using a supervised learning model, at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables; and generating an analyte panel based on the determined at least one target analyte.
[0009] By determining, for example using a supervised learning model, at least one target analyte associated with at least one reproductive health condition based on user health data; and generating an analyte panel based on the determined at least one target analyte, the advantage of providing a targeted approach for analyte panels that can reduce false positive rates that occur with broad spectrum analyte testing is afforded. The method can thus generate an improved analyte panel that is bespoke to user health data, and furthermore the method can reduce the clinician time needed to produce an analyte panel. The targeted approach may also permit more efficient use of user blood samples.
[0010] The step of determining may comprise: selecting and / or excluding, using for example the supervised learning model, at least one target analyte from a predefined set of target analytes based on the one or more medical history and symptom variables.
[0011] Where a supervised learning model is used, the supervised learning model may comprise a plurality of decision nodes to determine the at least one target analyte, each decision node configured to select and / or exclude at least one target analyte from a predefined set of target analytes based on a subset of the user health data and one or more clinical criteria associated with the respective decision node. The one or more clinical criteria that are associated with at least some of the decision nodes may comprise a minimum set of clinical diagnostic criteria for a given reproductive health condition compiled from one or more clinical data sources.
[0012] In one implementation, the supervised learning model may comprise a decision tree with a plurality of decision nodes, and wherein the step of determining comprises: traversing the decision tree based on the user health data; and selecting and / or excluding, at each decision node of the decision tree encountered, at least one target analyte from the predefined set of target analytes based on a subset of the user health data and one or more clinical criteria associated with the respective decision node.
[0013] The analytes in the predefined set of target analytes may be grouped into a plurality of groups of one or more complementary target analytes, and wherein selecting and / or excluding at least one target analyte from a predefined set of target analytes comprises selecting and / or excluding at least one group of target analytes from the predefined set of target analytes.
[0014] The plurality of groups may comprise one or more, or all, of: a) Anti-Mullerian hormone (AMH); b) Cycling hormones (oestradiol (E2), luteinising hormone (LH) and follicle-stimulating hormone (FSH); c) Thyroid hormones (FT4 and TSH); d) Androgens (testosterone (T), dehydroepiandrosterone sulphate (DHEAS) and sex hormone-binding globulin (SHBG); and e) Prolactin (PRL).
[0015] The supervised learning model may be configured to generate an analyte panel that includes up to and including 10 analytes through selection and / or exclusion of the plurality of groups of one or more complementary target analytes. For example, the analyte panel may include a minimum of three analytes, for example Anti-Mullerian hormone (AMH) and Thyroid hormones (FT4 and TSH).
[0016] In the above computer implemented method the step of determining may comprise: determining, using the supervised learning model, a likelihood, probability or risk of at least one of a plurality of reproductive health conditions based on the medical history and symptom variables and one or more clinical diagnostic criteria, and selecting and / or excluding at least one target analyte from a predefined set of target analytes based on the determined likelihood, probability or risk of the at least one of a plurality of reproductive health conditions.
[0017] In addition, the computer implemented method may comprise obtaining the user health data using a dynamic decision tree, comprising: sending, to a user device, a series of queries relating to medical history and symptom variables; and receiving a series of responses for each query, wherein at least some of the queries are determined, using the dynamic decision tree, based on the responses to one or more previous queries; and wherein the step of determining is performed concurrently with the obtaining of user health data.
[0018] In a further aspect of the invention, there is provided a method of training a machine learning model for determining a likelihood (for example a probability, or other value indicative thereof) of a reproductive health condition based on health data and / or analyte data, method comprising: receiving health data, the health data comprising: a plurality of user health data and / or analyte data (for example previous user health data and / or the user analyte data); selecting features in the health data and / or analyte data for training the machine learning model; generating a training data set comprising the selected features; and training the machine learning model using the generated training data set to determine a likelihood (for example a probability, or other value indicative thereof) of said reproductive health condition based on the user health data and / orthe user analyte data (for example current user health data and / orthe user analyte data). Health data or user health data may further comprise features determined or obtained from a scan performed on the user (i.e. scan data), e.g. ultrasound or other suitable scan techniques.
[0019] By providing a machine learning model for determining a likelihood or probability based on the user health data and / orthe user analyte data, the advantage of reducing clinician time needed to diagnosis reproductive health conditions and also the time needed to provide a diagnosis, is afforded. This allows earlier interventions to be made and improves health outcomes and reduces long term costs associated with undiagnosed or late diagnosed reproductive health conditions. The trained models also provide an improvement over diagnosis based only on existing clinical guidelines.
[0020] In this training method the training data may comprise at least one prediction signal for each of the plurality of user health data and / or analyte data, for example wherein the prediction signal is based on a diagnosis of the user health data and / or analyte data.
[0021] In some embodiments, the user health data may exclude health data and / analyte data in which concurrent reproductive health conditions have been diagnosed.
[0022] Selecting features from the user health data and / or analyte data may comprise selecting all features or a subset of features from the user health data and / or analyte data.
[0023] The machine learning model may be or comprise an ensemble machine learning model, preferably an ensemble classification model such as a decision tree ensemble model. Alternatively, the machine learning model may be or comprise a Bayesian Network model.
[0024] Selecting features from the user health data and / or analyte data step may comprise: selecting features from the user health data and / or analyte data that meet a threshold for statistical significance, for example a threshold p-value, for example selecting features with a p-value less than or equal to 0.05, for example using a chi-squared test for discrete variables and an F-test for continuous variables, wherein the target variable is either a positive or a negative diagnosis of the reproductive health condition.
[0025] Additionally or alternatively, selecting features from the health data and / or analyte data may comprise: calculating forthe selected features the mutual information gain between each selected feature and a target variable, wherein the target variable is either a positive or a negative diagnosis of the reproductive health condition; comparing the calculated mutual information gain to a baseline mutual information gain; and deselecting features which do not meet a threshold when compared to the baseline mutual information gain.
[0026] Training the machine learning model using the generated training data set may comprise: adding a selected feature from the training data set into the machine learning model; testing, after each addition of a selected feature, the machine learning model against the target variable; and repeating the method steps until a threshold condition between iterations of the machine learning model has been met, for example a threshold condition based on a convergence of a performance metric between iterations of the model against the target variable, for example an f1 score against a positive diagnosis.
[0027] The selected features may be ranked in order of mutual information gain against the target variable, and wherein the selected features are added into the machine learning model in order of ranking.
[0028] The above-described training method may also comprise: tuning, after each addition of a selected feature, a set of parameters of the machine learning model so as to maximise a performance metric of the machine learning model against a target variable, for example to maximise an f1 score against a positive diagnosis.
[0029] In a further aspect of the invention, there is provided a machine learning model for determining a likelihood of a reproductive health condition based on user health data and / or analyte data, wherein the machine learning model is trained using the above-described training method for example wherein the machine learning model comprises a decision tree ensemble machine learning model, such as a gradient boosted decision tree ensemble.
[0030] In another aspect of the invention, there is provided a computer implemented method of reproductive health assessment, the method comprising: receiving user health data comprising medical history and symptom variables of a user and analyte data obtained from results of a user analyte panel test; providing the user health data and analyte data to a model (for example a supervised learning model) comprising a logic algorithm encoded with clinical diagnostic criteria for reproductive health conditions; and determining, using the model, a likelihood or probability of at least one reproductive health condition based on the user health data and the user analyte data.
[0031] By providing a model (for example a supervised learning model) comprising a logic algorithm encoded with clinical diagnostic criteria for reproductive health conditions; and determining, using the model, a likelihood of at least one reproductive health condition based on the user health data and the user analyte data, the advantage of reducing clinician time needed to diagnosis reproductive health conditions and also the time needed to provide a diagnosis, is afforded. This allows earlier interventions to be made and improves health outcomes and reduces long term costs associated with undiagnosed or late diagnosed reproductive health conditions. The models (for example the trained models) can also provide an improvement over diagnosis based only on existing clinical guidelines.
[0032] The computer implemented method may also comprise encoding the user health data and analyte data, wherein encoding comprises extracting from the user health data and analyte data a set of predefined variables to define a plurality of logic blocks, and wherein the providing step comprises providing the encoded user health data and analyte data to the supervised learning model, further wherein the logic algorithm is configured to assemble the logic blocks based on clinical guidelines for a plurality of health conditions to determine a likelihood of at least one reproductive health condition based on the user health data and the user analyte data.
[0033] In yet another aspect of the invention, there is provided a computer implemented method of reproductive health assessment, comprising: providing user health data including a plurality of medical history and symptom variables and / or analyte data to a machine learning framework comprising at least one machine learning model, each machine learning model trained to determine a probability or likelihood of a separate reproductive health condition based on the user health data and / or the user analyte data; and determining, using the machine learning framework, a probability likelihood of at least one reproductive health condition based on the user health data and / or the user analyte data.
[0034] The at least one machine learning model may be trained using the training method described above, for example wherein the at least one machine learning model comprises a decision tree ensemble model machine learning model trained using the above-described training method.
[0035] Additionally or alternatively, the at least one machine learning model may be or comprise a Bayesian Network model.
[0036] The method may further comprise generating a user health assessment report based on the determined likelihood of the at least one reproductive health condition; and preferably, wherein the method comprises encoding the user health data and / or analyte data and the likelihood of at least one reproductive health condition, and wherein generating the user health assessment report comprises selecting predefined text for inclusion in the report based on the encoded health data, analyte data and likelihoods. Optionally or preferably, the report includes a relative uncertainty of the determined probabilities of the at least one reproductive health condition, e.g. derived from a Bayesian Network model. The computer implemented method or machine learning models described herein may be implemented for various reproductive health condition for example at least one of: Polycystic ovary syndrome, Overt / subclinical hypothyroidism, Overt / subclinical hyperthyroidism, Premature Ovarian Insufficiency, Peri / Menopause, Hyperandrogenism, Hyperprolactinaemia, Hypothalamic Amenorrhea, ovulatory disorder, Endometriosis, pelvic issues (such as fibroids or adenomyosis), and Iron deficiency anaemia.
[0037] The analyte data in the above-described methods may be obtained using an analyte panel generated by the above-described computer implemented method of generating an analyte panel for a reproductive health assessment.
[0038] In a further aspect of the invention, there is provided a computer readable non-transitory storage medium comprising program instructions that, when executed by one or more processing devices, cause the one or more processing devices to perform the methods described herein.
[0039] In a further aspect of the invention, there is provided a computer readable non-transitory storage medium, comprising the trained machine learning model detailed above. The computer readable non-transitory storage medium may contain program instructions that, when executed on processing circuitry, cause the processing circuitry to operate a machine learning model trained as detailed above.
[0040] The machine learning model may be or comprise an ensemble machine learning model, preferably an ensemble classification model and / or a decision tree ensemble machine learning model. Alternatively, the machine learning model may be or comprise a Bayesian Network model.
[0041] In a further aspect of the invention, there is provided a system comprising one or more processing devices configured with instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform the method of any previous aspect.
[0042] Any feature of the methods as described herein may also be provided as a system feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure. Any, some and / or all features in one aspect of the invention may be applied to other aspects of the invention, in any appropriate combination or sub-combination. In particular, method aspects may be applied to corresponding system aspects, and vice versa.
[0043] It should also be appreciated that particular combinations of the various features described and defined in any aspect of the invention can be implemented and / or supplied and / or used independently. The invention extends to the methods and systems substantially as herein described and / or as illustrated with reference to the accompanying figures. The invention also extends to any novel aspects or features described and / or illustrated herein. In this specification the word 'or' can be interpreted in the exclusive or inclusive sense unless stated otherwise. Brief Description of Drawings
[0044] Some practical implementations will now be described, by way of example only, with reference to the accompanying drawings in which:
[0045] Figure 1 shows a computer implemented method of generating a bespoke analyte panel for testing for reproductive conditions;
[0046] Figure 2 shows an illustrative example of part of a supervised learning model for determining at least one target analyte associated with at least one reproductive health condition;
[0047] Figures 3a to 3b shows a schematic example of the method of Figure 1 ;
[0048] Figure 4 shows a computer implemented method of reproductive health assessment, according to an embodiment of the present disclosure;
[0049] Figure 5 shows a schematic illustration of the method of Figure 4;
[0050] Figure 6 shows a further schematic illustration of the method of Figure 4 involving encoding user health data and analyte data into logic blocks;
[0051] Figure 7 shows a computer implemented method of reproductive health assessment and in particular details a method of obtaining user health data from a user that can be used in the method of Figure 1 ;
[0052] Figure 8 shows a computer implemented method of reproductive health assessment;
[0053] Figure 9 shows a method of training a machine learning model to determine a likelihood of a separate reproductive health condition based on the user health data and / or the user analyte data;
[0054] Figure 10 shows an example of the feature selection process of the method Figure 8;
[0055] Figure 11 shows an example method of training a machine learning model to determine the likelihood of a reproductive health condition;
[0056] Figure 12 shows an example flow diagram for training machine learning models to predict the presence of\determine the likelihood of a diagnosis of reproductive health condition;
[0057] Figure 13 shows an example flow diagram of training machine learning models to predict the presence of\determine the likelihood of a diagnosis of reproductive health condition;
[0058] Figure 14 shows an example flow diagram of using a machine learning model to provide a probabilityMikelihood of diagnosis of a reproductive health condition;
[0059] Figure 15 shows an example flow diagram of using a machine learning model to provide a probabilityMikelihood of diagnosis of a reproductive health condition;
[0060] Figure 16 shows a schematic diagram of an example system for implementing the disclosed methods;
[0061] Figure 17 shows a simplified example of Bayesian Network model configured to determine a likelihood of a reproductive health condition based on user health data and / or the user analyte data;
[0062] Figure 18 shows a simplified example of a machine learning framework comprising machine learning models, each machine learning model trained to determine a probability or likelihood of a separate reproductive health condition based on the user health data and / or the user analyte data; and
[0063] Figure 19 shows the probability or likelihood of a number of reproductive health conditions calculated by a machine learning framework based on a set of user health data and / or analyte data.
[0064] In the drawings like reference numerals are used to indicate like elements. Specific Description
[0065] Testing of analytes in blood samples is important for diagnosing reproductive health conditions. An important consideration in such testing is the determination of which analytes to test for. One reason that this is important is because broad spectrum analysis leads to unnecessary over testing for analytes, which gives rise to high false positive diagnosis rates. A more targeted approach is therefore needed to reduce false positive rates. In addition, testing for a large number of analytes requires larger blood samples to be provided by a patient or user. Reducing the number of blood samples and the volume of these samples that are needed is desirable. Furthermore, reducing the number of analytes to be tested reduces testing time and cost.
[0066] Figure 1 shows a schematic diagram of a computer implemented method 100 of generating a bespoke analyte panel for testing for reproductive conditions that seeks to address these problems.
[0067] In step 110, method 100 comprises receiving user health data including a plurality of medical history and symptom variables. As an example, a user / patient may be provided with an online or virtual health assessment to complete. The results from the completed health assessment form the user health data that is received in method 100. The virtual health assessment includes a predefined set of queries or questions to target and flag reproductive health conditions the user may be at risk of. Specifically, the queries are configured to target symptoms associated with reproductive health conditions that the user / patient may have been experiencing along with details of the user / patient’s medical history including any current and past medications. The virtual health assessment may also include queries targeting patient / user risk factors associated with reproductive conditions such as smoking and diet. The medical history variables of the user health data may include, but are not limited to: BMI, menstrual cycle characteristics, previous medical diagnosis for example a diagnosis of Polycystic ovary syndrome (PCOS); medication history, contraception used. The symptom variables may include, but are not limited to: headaches; acne; hirsutism; double vision; partial loss of vision; nausea; vomiting; fatigue and nipple discharge. The user health data may also include data or features from medical scans or other diagnostics procedures. For example it may include data or features from ultrasound scans, for example ultrasound scans of the user’s uterus and / or ovaries.
[0068] In the targeted approach described herein, the target medical history and symptom variables (user health data points) obtained by the virtual health assessment are selected by first identifying a set of common reproductive health conditions to screen for, and then determining a minimum set of clinical criteria for diagnosis and risk factors commonly associated with the conditions based on evidence published via clinical guidelines by governing bodies such as the National Institute of Clinical Excellence (NICE), the Royal College of Obstetricians and Gynaecologists (RCOG), the European Society of Human Reproduction and Biology (ESHRE), International Guidelines for PCOS, and the Endocrine Society. These are the same guidelines used by clinicians and specialists to investigate symptoms and diagnose conditions related to reproductive and thyroid health. The virtual health assessment is thus constructed and configured to target and collect a plurality of user health data points required to assess the relevant diagnostic criteria.
[0069] In preferred examples, the user health data is obtained using a dynamic decision tree model with a plurality of nodes corresponding to all possible queries in the virtual health assessment. In this way, obtaining the health data involves traversing the decision tree by sending, to a user device such as a computing device running a mobile or web application, a series of queries relating to target medical history and symptom variables; and receiving a series of responses for each respective query, wherein at least some of the queries are determined, using the dynamic decision tree, based on the responses to one or more previous queries. The connections or links between each node are determined by logic and / or decision rules based, at least in part by, the compiled clinical criteria.
[0070] In step 120, the method 100 comprises determining, using a supervised learning model, at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables of the user health data. An example supervised learning model is shown in Figure 2 and is described in more detail below. The user health data received in step 110 is used by the supervised learning model to determine at least one target analyte for inclusion in the user’s analyte panel.
[0071] In step 130, the method 100 comprises generating a user analyte panel based on the determined at least one target analyte. The resulting user analyte panel is thus bespoke to the user health data. Once generated, blood samples collected from the user can be tested according to the bespoke analyte panel. The results of blood tests provide user analyte and / or hormone data that can be used, together with the user health data in the diagnosis of reproductive health conditions, as will be described in more detail below.
[0072] In preferred embodiments, determining the at least one target analyte comprises selecting and / or excluding, using the supervised learning model, at least one target analyte from a predefined set of target analytes based on the one or more medical history and symptom variables.
[0073] In general, the analytes that are determined, selected or excluded in step 120 of the method 100 are taken from a predefined set of target analytes. These analytes include but are not limited to: Anti-Mullerian hormone (AMH); Cycling hormones (oestradiol (E2); luteinising hormone (LH); follicle-stimulating hormone (FSH); Thyroid hormones (FT4 and TSH); Androgens (testosterone (T), dehydroepiandrosterone sulphate (DHEAS); sex hormone-binding globulin (SHBG); and Prolactin (PRL).
[0074] It is advantageous to group these predefined analytes into a plurality of groups of one or more complementary analytes. In particular, certain hormones / analytes should be tested in tandem with others in order to interpret their results correctly. For example, the thyroid hormone free thyroxine (FT4) can be tested alongside thyroid-stimulating hormone (TSH) to improve understanding of the result.
[0075] In an embodiment, the plurality of groups of complementary analytes comprise: a) Anti-Mullerian hormone (AMH); b) Cycling hormones (oestradiol (E2), luteinising hormone (LH) and follicle-stimulating hormone (FSH); c) Thyroid hormones (FT4 and TSH); d) Androgens (testosterone (T), dehydroepiandrosterone sulphate (DHEAS) and sex hormone- binding globulin (SHBG); and e) Prolactin (PRL).
[0076] Each of these particular groupings correspond to a minimum analyte testing required to identify at least one reproductive condition or coincident reproductive conditions. In restricting the selected analytes to these groupings, over testing of analytes can be reduced whilst maintaining the minimum required to test for a reproductive condition.
[0077] The supervised learning model may restrict the analyte panel to a minimum of two or three analytes and a maximum of 10 analytes. In some examples, a default set of three analytes groups a) Anti-Mullerian hormone (AMH) and c) Thyroid hormones (FT4 and TSH) are always included in the analyte panel to begin with, which may be removed / excluded by the supervised learning model based on the user health data, along with selecting / adding further groups. In restricting an analyte panel in this way, over testing, which can give rise to false positive rates is further reduced. It also permits more efficient use of a blood sample. For example, a blood sample with only a limited volume for example a sample of 1 millilitre (ml) may be the only blood sample available for a given user and so wide spectrum testing for analytes associated with a plurality of reproductive conditions may not be possible using the limited sample. In such cases, the selection and reduction of the analyte panel to between 3 and 10 analytes by the supervised learning model based on one or more of the medical history and symptom variables is particularly advantageous.
[0078] In preferred examples, up to 10 analytes are included in the user analyte panel which can screen for up to ten reproductive health conditions which rely on at least one abnormal analyte / hormone result for diagnosis, i.e. the respective clinical diagnostic criteria include at least one abnormal analyte / hormone result. The reproductive health conditions of method 100 for which analyte panels may be generated for may comprise at least one of: Polycystic ovary syndrome, Overt / subclinical hypothyroidism, Overt / subclinical hyperthyroidism, Premature Ovarian Insufficiency, Peri / Menopause (e.g. if aged 45 or under), Hyperandrogenism, Hyperprolactinaemia and Hypothalamic Amenorrhea . The supervised learning model may also be configured to generate analyte panels for other reproductive health conditions. Other conditions can be identified based on the user health data alone, such as ovulatory disorder, Endometriosis, pelvic issues (such as fibroids or adenomyosis), and Iron deficiency anaemia.
[0079] Figure 2 shows an illustrative example of part of a supervised learning model 200 for determining at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables.
[0080] In this example, the supervised learning model is a dynamic decision tree 200 having a plurality of decision nodes 210 arranged sequentially in the decision tree 200. Based on the supervised learning, the nodes 210 of the dynamic decision tree 200 determine at least one target analyte for an analyte panel for a given set of user health data. Specifically, each decision node 210 is configured to select and / or exclude at least one target analyte (or group of target analytes) from the predefined set of target analytes based on a subset of the user health data and one or more clinical criteria associated with the respective decision node 210.
[0081] At least some of the nodes 210 of the decision tree 200 may have a weighting which associates a likelihood of a given reproductive health condition with the user health data associated with the respective node 210, for example with one or more symptom and / or medical history variables. The weighting of the decision nodes 210 is encoded based on the determined clinical criteria for identifying the given reproductive health conditions (as discussed above). The decision nodes 210 of the decision tree 200 are thus used to determine if an analyte should be added to or removed from the analyte panel.
[0082] For example, there may be a decision node 210 associated with the symptom of headaches. Based on the user health data associated with this symptom and one or more clinical criteria the decision tree 200 and in particular the decision node 200 may determine that an analyte should be added to the panel. The user health data associated with the symptom may also cause the dynamic decision tree 200, for example a further decision node, to consider the further details of the symptom provided in the user health data for example the frequency of the headaches / migraines. Again, this further decision node may have a weighting which associates a likelihood of a reproductive health condition with at least a part of the user health data.
[0083] For example, in the case of the user health data indicating a symptom of frequent headaches / migraines the associated decision nodes of the dynamic decision tree 200 may determine that an analyte (prolactin (e)) should be added to the analyte panel, because this symptom is, according to clinical guidelines, associated with hyperprolactinaemia.
[0084] The weighting associated with decision nodes of the dynamic decision trees may also be interdependent or combined so that the outcomes of several decision nodes are determinative of whether a target analyte is selected for or excluded from the bespoke analyte panel. That is to say, the weighting of a decision node 210 may also be dependent on determinations made by other / previous decision nodes in the dynamic decision tree 200. For example, a medication indicated in the user health data may affect the testing of a particular analyte (for example it might be detrimental to the testing of that analyte) and a decision node 210 associated with a symptom may take this into account as to whether that analyte should be included in the bespoke analyte panel.
[0085] Method 100 therefore may select and / or exclude, at each decision node 210 of the decision tree encountered, at least one target analyte from a predefined set of target analytes based on one or more medical history and symptom variables input to the respective decision node and one or more clinical criteria associated with the respective node. The decision tree 200 and method 100 may run concurrently with collection of the user health data. For example, it may run concurrently with a virtual health assessment collecting user health data from the user.
[0086] Once the supervised learning model 200 has determined one or more analytes based on the user health data, the method 100 generates in step 130 a bespoke user analyte panel. The user analyte panel can be used in the testing of a user blood sample so that potential reproductive health conditions can be identified. The bespoke generation of the analyte panel based on user health data reduces over testing and the associated false positive rates, ensures the necessary analytes are tested on a small blood sample, and is cost-efficient. Figures 3a and 3b show a schematic example of method 100. In Figure 3a, a user undertakes a virtual health assessment 1610 thereby providing the information for the health data used in method 100. In the example shown, the user is provided with queries relating to their periods and indicates that their periods are irregular and that they experience acne, excessive body hair and pelvic pain. This information is encoded into the user health data and is received by a supervised learning model 200, in this example a dynamic decision tree such as the dynamic decision tree 200 described above.
[0087] In Figure 3b, the dynamic decision tree 200 determines 1620 at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables and one or more clinical criteria (e.g. see dashed boxes in Figure 3a). In particular, the decision nodes 210 of the dynamic decision tree 200 associated with the symptoms of irregular periods, acne, excessive body hair and pelvic pain determines, based on the user health data and the weighting assigned to the decision nodes 210, which analytes to add or remove.
[0088] In this particular case, as indicated in Figure 3b the dynamic decision tree 200 determines that there is a likelihood that the user might have polycystic ovary syndrome (PCOS) and the analyte grouping d), containing Androgens (testosterone (T), dehydroepiandrosterone sulphate (DHEAS) and sex hormone- binding globulin (SHBG)), is selected from the groupings described above and added to the user’s analyte panel. A default set of analytes groups a) Anti-Mullerian hormone (AMH) and c) Thyroid hormones (FT4 and TSH) are also included in the analyte panel as shown in Figure 3b. These default analytes may be used in the determination of the likelihood that a user has a reproductive health condition. For example, AMH levels in the resulting analyte data may also be used to determine a likelihood that a user might have PCOS.
[0089] Figure 4 shows a computer implemented method 300 of reproductive health assessment, according to an embodiment of the present disclosure. In more detail, the method 300 comprises in step 310 receiving user health data 1510 comprising medical history and symptom variables, and analyte data 1520 obtained from results of a user analyte panel test (preferably generated as described above in method 100). In step 320, the method comprises providing the user health data and analyte data to a supervised learning model encoded with a minimum set of clinical diagnostic criteria or logic blocks for a plurality of reproductive health conditions, as compiled from various clinical guidelines and data described above. In one example, the supervised learning model comprises a logic algorithm encoded with clinical diagnostic criteria for the plurality of reproductive health conditions. In step 330, the method comprises determining, using the supervised learning model, a likelihood or probability of at least one reproductive health condition based on the user health data and the user analyte data. The method 300 preferably further comprises a step 340 of generating a user health assessment report based, at least on part, on the determined likelihood or probability of the at least one reproductive health condition.
[0090] Figure 5 shows a schematic illustration 1500 of method 300. A blood sample of the user is tested for the selected analytes according to the user’s bespoke analyte panel and the results of the analyte panel in combination with the user health data are used to identify a potential reproductive health condition. The results of an example analyte panel are shown on the lefthand side of Figure 5, whereby AMH, T and DHEAS levels are high, TSH and FT4 levels are normal and SHBG level is low. This analyte data combined with the user health data of irregular periods, acne and excessive body hair (see Figure 3a) can be used to determine or identity, by the supervised learning model or the machine learning model frameworks of method 300, 500 described in more detail below, a likelihood of at least one potential reproductive health condition based on clinical diagnostic criteria, in this example suspected PCOS. A clinician may also use this data to arrive at a diagnosis or check the potential diagnosis produced by the supervised learning model or the machine learning model frameworks.
[0091] Preferably, the method 300 also includes encoding the user health data and analyte data, wherein encoding comprises extracting or selecting from the user health data and analyte data a set of predefined variables to define a plurality of logic blocks 1530, 1540 or input parameters / variables for the logic algorithm, as illustrated in Figure 6. In this example, the providing step 320 then provides the encoded user health data and analyte data to the supervised learning model. The logic algorithm is configured to assemble the logic blocks 1530, 1540 based on clinical guidelines. This is shown in Figure 6 in which the logic blocks have been assembled 1550, 1560, 1570, according to the logic algorithm to identify three reproductive health conditions X,Y,Z. For example, a logic block from the analyte data may be associated with an identified increased or high level of a hormone (for example AMH), and a logic block from the user health data may be associated with the presence of a target symptom (for example “irregular periods”). The presence of these two logic blocks / input variables may be sufficient for the logic algorithm to assemble the logic blocks based on clinical guidelines to identify a reproductive health condition.
[0092] As noted above, step 320 preferably includes encoding the user health data and analyte data into logic blocks or input parameters, and providing the encoded user health data and analyte data as inputs to the supervised learning model. The supervised learning model includes, e.g. in the logic algorithm, at least one set of clinical diagnostic criteria for each reproductive health condition being screened for. For example, for a given reproductive health condition, there will be a number of clinical scenarios that define the minimum blood result(s) and / or OHA answer(s) required to flag the condition(s) in the user’s results, i.e. the minimum set of clinical diagnostic criteria. For example, there may be multiple different assemblies of logic blocks identified from the clinical guidelines for a single reproductive health condition and the supervised learning model can be encoded to build the logic blocks into any one of these assemblies in order to identify the reproductive health condition. The use of a number of different assemblies / diagnostic criteria for each reproductive condition provides for a differential diagnosis approach - whereby all the possible causes of a symptom / abnormal blood result are considered by the method 300. To diagnose a condition, a bare minimum set of criteria are required to be met to confirm the presence of a particular reproductive health condition and not one of the other conditions considered in the differential diagnosis. However, a reproductive health condition often comes with a number of "extra" symptoms and blood results on top of the minimum criteria. For some conditions, this is very simple, for example hyperthyroidism requires minimum of a low TSH and high FT4 for diagnosis but "extra" symptoms like weight loss can assist in the diagnosis. For other conditions, like PCOS, it is more complex as there are multiple sets of criteria which serve as the minimum for diagnosis, leading to multiple clinical scenarios where someone can be diagnosed with PCOS according to clinical guidelines. Method 300 takes this into account by utilising a number of different assemblies for each reproductive condition and by having assemblies for a number of reproductive conditions.
[0093] In order to flag or generate a correct likelihood or diagnosis of a new conditions, for a user who has indicated in their health data one or more certain pre-existing medical conditions the supervised learning model first checks against the encoded clinical diagnostic criteria for the pre-existing condition(s) to determine whether the user has any expected abnormal analyte data (blood results) and / or health data (symptom variables) for that condition. If it is determined that the user has abnormal analyte data (blood results) and / or health data which do not satisfy the clinical diagnostic criteria for the pre-existing condition(s), the supervised learning model then proceeds to check if any of the remaining encoded clinical diagnostic criteria are satisfied to flag or determine a likelihood of a new condition.
[0094] In an example, step 340 of generating the report involves encoding the flagged or determined likelihood of at least one reproductive health condition, and generating the user health assessment report by selecting, or retrieving from a database, associated text for the report based on the encoded health data, analyte data and likelihoods.
[0095] The generated report can also include any “linked” abnormal analyte data (blood results) and health data which are commonly observed with the flagged condition but are technically not required for diagnosis (according to the encoded minimum set of diagnostic criteria). By way of example, a clinical scenario for a diagnosis of PCOS may require an abnormal blood result of high testosterone (T) and user health data indicating a menstrual cycle length of >35 days (which is the minimum required for diagnosis). However, if the user’s health data also indicates a symptom of dark skin patches (which is linked to PCOS), then this is also included in the generated report as an associated symptom commonly observed in the condition. Using this approach, all of the relevant analyte data (blood results) and health data for each reproductive health condition can be grouped together in the report with associated text so the user is provided with a holistic understanding of any abnormal results / symptoms.
[0096] The report may “triage” the determined likelihood in the report so that a clinician can quickly identify determinations that need a more thorough review. For example, based on the user health data and / or analyte data used in the identification of the reproductive health condition, the determined likelihood may be higher or lower. The identified reproductive health conditions with a lower determined likelihood may be flagged to the clinician. For example, a traffic light system (red, amber, green) can be assigned in the report to identify where detailed checks - red, checks - amber , and review - green may be needed depending on the determined likelihood.
[0097] Figure 7 shows a computer implemented method 400 of reproductive health assessment, and in particular details a method of obtaining user health data from a user that can be used in method 100. The method comprises sending 410, to a user device, a series of queries relating to medical history and symptom variables. The user responds to each query and a series of responses for each query are received 420. At least some of the queries sent to the user and / or presented to the user are determined, using the dynamic decision tree, based on the responses to one or more previous queries wherein the step of determining is performed concurrently with the obtaining of user health data. In such a way, a user is presented with a dynamic set of queries that are dependent on previous queries, which tailors the obtained medical history and symptom variables in the user data to each user. The obtained user health data can then be used in, for example, method 100 or the other methods 300 detailed herein.
[0098] Machine learning frameworks to identify reproductive health conditions
[0099] The identification of a reproductive health condition from either the user health data and / or blood analyte panel test results can take up clinician time and may require input from specialist clinicians. There is therefore a need to reduce the clinician time taken to identify reproductive health reproductive health conditions. This has many benefits, for example by reducing clinician time users can be provided with results at an earlier stage and take actions to mitigate identified reproductive health conditions sooner.
[0100] Furthermore, when deciding on a final diagnosis, clinicians will use the appropriate clinical guidelines for diagnosis. However, depending on the diagnosis, some clinicians may take further steps, such as utilising their own clinical experience or recent peer-reviewed literature, to further aid the diagnosis. Where a clinician utilises their experience in this way to improve on the clinical guidelines then this indicates that something fundamental is added to the diagnosis process which is beyond the information available in the clinical guidelines. In this instance, a clinician may be using intuition or other experience to make a decision. A machine learning model that captures this approach would provide additional value over the existing clinical guidelines. As used herein, a ‘machine learning model’ refers to a model which has undergone a training or learning stage, in which certain parameters of the model are refined using training data. A machine learning framework refers to a collection of at least one machine learning model trained to determine the probability or likelihood of at least one reproductive health condition.
[0101] The hypothesis that the clinician provides better decisions than the clinical guidelines was evaluated by: looking at cases where the clinician agreed with the rules-based clinical guidelines diagnoses, and where they didn’t; comparing true / false positives, and true / false negatives for each feature from the user health data and analyte data; and determining if the clinician is adding knowledge by identifying the differences in distribution of features (there will be clear differences in the distributions of some features if the clinician is adding knowledge and there will be no obvious difference if the clinician is only adding human error). It was found that several features for each reproductive health condition showed a clear difference in the distributions, which indicated the clinician is consistently adding their own interpretation of the guidelines and providing an improvement over the existing clinical guidelines. Replicating the clinician diagnosis is therefore desirable.
[0102] As an example for PCOS, one of the deciding factors of this condition is relatively high AMH for the user’s age. It was found that the true positive diagnoses (i.e. a clinician agrees with clinical guidelines) tends to have higher AMH than the false positive diagnoses (i.e. clinical guidelines suggest PCOS but a clinician disagrees). This suggests that a clinician will often not diagnose PCOS where the AMH level is only slightly high, thereby reducing false positives and providing an improvement over simply following the clinical guidelines. The machine learning framework described below is trained to replicate clinician decisions and provide a reproductive health assessment that determines the probability or likelihood of at least one reproductive health condition.
[0103] In more detail, the method 500 shown in Figure 8 seeks to provide a computer implemented method of reproductive health assessment. User health data including a plurality of medical history and symptoms and optionally scan features and / or user analyte data, for example user analyte data obtained from results of a user analyte panel test, is optionally received and is provided to a machine learning framework 510, 520. Wherein the machine learning framework comprises at least one machine learning model, each machine learning model is trained to determine a probability or likelihood of a separate reproductive health condition based on the user health data and / or the user analyte data. The machine learning framework determines 530 a likelihood or probability of at least one reproductive health condition based on the user health data and / or the user analyte data.
[0104] These frameworks and at least one machine learning model are described in detail below. It will be appreciated that the machine learning models employed by the framework may in general be unsupervised, semi-supervised or supervised. The examples described herein are examples of supervised machine learning models.
[0105] In an example, the supervised machine learning framework comprises at least one, and in most cases several, supervised machine learning models. Each of the supervised machine learning models are configured and trained to each identify a different reproductive health condition using user health data and / or user analyte data.
[0106] The user health data comprises information on a user’s medical history for example medication that has been used or symptoms that the user has experienced, as described above. The health data may also include data on other activities that modify the likelihood or probability of a reproductive health condition being present, for example user diet or whether a user smokes or has smoked. The user health data may also include data or features obtained or derived from scan, e.g. ultrasound features, such as thin endometrium, fibroids, polyps, abnormal lesion, hypervascularity, etc. as described above.
[0107] The analyte data comprises blood results from an analyte panel testing for reproductive health conditions, preferably a bespoke user analyte panel generated by method 100. In particular, the analyte data may contain concentrations of analytes associated with one or more potential reproductive health conditions.
[0108] Supervised machine learning model topography
[0109] It will be appreciated that in general, a number of different machine learning models may be suitable to be trained to identify a reproductive health condition using user health data and / or user analyte data.
[0110] One example of a suitable machine learning model is a decision tree ensemble model such as random forests and gradient-boosted trees ensembles, for example that are provided by XGBoost models (a distributed gradient-boosted decision tree, GBDT, machine learning library). This model structure is beneficial because such ensemble models can learn non-linear signals from the selected training features. These models are also adept at handling missing data, which in this specific application is particularly useful because the user health data may not be complete for each user and each user is likely not to have been tested for every possible analyte. Furthermore, such GBDT models are relatively quick to train, compared to some deep learning architectures, and are very customisable and known to give good results on a wide variety of data, making it relatively easy to efficiently loop through several model variations for each reproductive health condition and get good outcomes.
[0111] Another example of a suitable machine learning model is a Bayesian network model, as described in more detail below. Bayesian network models may offer advantages over classification type models in terms of explainability of results, since Bayesian Network models are constructed entirely around the relationships between variables / features and diagnosis, such that the impact of each feature on the diagnosis is inherent in the design of the model; symptoms are explicitly causally linked to conditions. In addition, as a Bayesian Network model effectively encodes dependencies among all variables, it may readily handle situations where some data is missing and may potentially require less input data to reach a diagnosis.
[0112] Supervised machine learning model training
[0113] Each supervised machine learning model of the machine learning framework is preferably trained to identify a different reproductive condition, each using a different training data set comprising at least some features extracted from a user's health data and / or analyte data that have resulted in a particular diagnosis (positive or negative) for said reproductive health condition. Health data and / or analyte data in which multiple / concurrent reproductive health conditions were identified by a clinician are preferably removed from the training data sets or not included in the health / analyte data from which the training dataset was generated. Each training data set may contain an associated or paired training output\prediction signal for each of the plurality of the health data and / or analyte data. For example, the associated / paired training output\prediction signal may comprise the clinical diagnosis (referred to herein as the prediction signal) for the particular user’s health data for said reproductive health condition, for example a positive or negative diagnosis.
[0114] An outline of an example training process 600 is shown in Figure 9. The process 600 comprises two main steps, the first of which is the extraction or selection of features 610 from the health data, and generation of a training data set comprising the extracted / selected features. The training data set may also include the paired training output\prediction signal for each user's health data. The second step 620 is the training of the machine learning model using the generated training data. Selection of features may include selecting all features, or a subset of features e.g. relevant to the model. The selection of features may be dependent on the machine learning model implemented. For example, feature selection may be particularly beneficial for classification models, such as a decision tree ensemble model, as described in more detail below. However, where the machine learning model is a Bayesian network, the selection of features may be less important, e.g. it may include selecting all features, or simply features relevant for the nodes of the given Bayesian Network model, or the step of selection may be omitted. Ensemble machine learning models
[0115] The training data set may comprise a large number of features, for example hundreds of features, or even thousands of features. As such, in some embodiments an important step in training / building the machine learning model, particularly for classification type models, is the selection of those features that were most relevant to diagnosis of the reproductive health condition. An outline of the feature selection process is shown in the method 700 of Figure 10. Method 700 is preferably applied to classification type models, such as a decision tree ensemble models.
[0116] In step 710, the method 700 comprises determining a statistical dependence or statistical significance of a feature in the user health data and / or analyte data with respect to a target variable - the target variable in this example being a positive diagnosis of the reproductive health condition. The target variable could alternatively be a negative diagnosis of the reproductive health condition. In one example, the statistical dependence / significance of a feature is determined by testing and producing a p-value (probability value) as is known in the art, in this example using a chi-squared test for discrete variables such as presence of a symptom for example headaches and an F-test for continuous variables such as analyte concentration. Other statistical tests may also be suitable. The resultant p-value from the tested statistical dependence of a feature is compared to a threshold value, for example a p-value of greater than 0.05. A p-value greater than 0.05 indicates no statistical dependence between the feature and the target, whereas p-value equal to or less than 0.05 indicates a statistically significant dependence between the feature and the target. Thus, features having a p-value greater than 0.05 are not selected and those with p-value equal to or less than 0.05 are selected for the training of the machine learning model.
[0117] Optionally, step 710 may be preceded by an initial step of selecting, for a given reproductive health condition, the features from the user health data and / or from the user health data which are most relevant for that reproductive condition according to the clinical guidelines (i.e. without determining their statistical significance). This optional step can provide an appropriate starting point for feature selection and can reduce the time taken to select appropriate features. Accordingly, if certain features have been selected using the optional initial step the statistical dependence of these selected features does not need to be calculated in step 710.
[0118] In step 720, for the selected features, those with a p-value less than or equal to 0.05 or those selected in the initial optional step, the mutual information gain between the feature and the target variable is calculated. This score quantifies how much information about the likelihood of the condition in question can be gained from knowing the value of a particular feature. In an example implementation, the mutual information gain is calculated using an entropy estimation from k-nearest neighbours distances between the feature and the condition indicators.
[0119] In step 730, the calculated mutual information gain is compared to a baseline mutual information gain value. For example, the baseline value may be a calculated average mutual information gain between the target and several (for example 10) randomly generated features. If the mutual information gain (score) for a feature does not meet a threshold value when compared to the baseline mutual information gain, then the feature is discarded / deselected in step 740. For example, if a feature has a mutual information gain score that is not at least 10% higher than the baseline average mutual information score, then the feature is discarded and deselected.
[0120] In step 750, the selected features may optionally be ranked in descending order by mutual information score. The selected features using method 700 are then used to generate a training data set that can be used to train a supervised machine learning model. In the example of an ensemble machine learning model, the model is trained by an additive training approach, as described below.
[0121] An example method 800 of training a machine learning model to determine a probability or likelihood of a reproductive health condition is shown in Figure 11 . Method 800 may preferably be applied to classification type models, such as a decision tree ensemble models. To train a machine learning model, the method 800 begins in step 810 with a default set of features, based on the clinical guidelines for each condition. For example, for PCOS, irregular periods and elevated AMH are indicators specified in clinical guidelines, so the patient's AMH value and health assessment response in the user health data indicating regular or irregular periods are included. Following this, the process then continues by adding a selected feature from the training data set into the machine learning model 810, in order of highest mutual information score. The machine learning model is then tested against the target variable, for example a positive diagnosis of the reproductive health condition 820, to determine an outcome. The steps 810, 820 are repeated, by adding selected features one by one, until a threshold condition between iterations of the machine learning model has been met. As with the preparation of the training data set described above, the inclusion of the default set of features is optional in step 810. The inclusion of these features from the outset can make the training process more efficient as these default features are highly likely to be added by step 810.
[0122] In one example, features are additively incorporated 810 into the machine learning model starting with the highest ranked feature from method 700 to create an ensemble model. After each addition of a feature, the model is tested in step 820 against the target variable to determine a predictive accuracy / performance, for example an F-score (f1) against the positive class (i.e. occurrence of the reproductive health condition). The number of and / or the parameter values of the machine learning model (e.g. in the case of GBDT models, the number and / or value of estimators and model hyperparameters such as the learning rate, maximum tree depth, subsample ratio, subsample ratio for columns by tree) may be tuned in the training process 800 at each iteration and selected to maximise the f1 -score against the positive class (i.e. occurrence of the reproductive health condition). This can be done for example by using a grid search algorithm to select combinations of model parameters to use in the model training with the list of features, and iterating like this to find the combination with the highest f1 -score against the positive class. Similarly, the choice of whether to balance the classes in the training data (i.e. sampling the examples of nonoccurrence of the condition to get an even split of occurrence and non-occurrence of the condition, which can help improve model performance on predicting the minority class, in this case the occurrence of a condition), and whether to rescale the training data before training so all features have mean and variance on the same order of magnitude, can be made based on the model performance. In the example of a GBDT model, it was found that leaving the training data imbalanced and unsealed produced the best model accuracy / performance, e.g. as quantified by the f1 -score. An additional feature from the training dataset is added in each iteration at step 830 until a threshold condition, e.g. based on F-score (f1) convergence between successive iterations of the model, is met.
[0123] Preferably, the features to train the machine learning model on are selected by ranking them in order of mutual information gain against the target variable, and then determining model performance after each iteration. A list of features which maximised the model performance is then selected.
[0124] The supervised machine learning model is tested using a test data set specific to the reproductive health condition and that is different to the training data set. The test data set comprises health data and / or analyte data for the reproductive health condition. For example, the test data set might comprise 25% of the available data, stratified for the target variable. A Log loss function, which is particularly suited to binary classification tasks may be used to verify the output of the model.
[0125] The machine learning model training process 800 detailed above can be used to create multiple models for each reproductive health condition, with different features and different hyperparameters in use or based on a different ranking of features, for example the order of the top ranked features may be shuffled and used instead as a starting point. The final model for each reproductive health condition can then be selected from the multiple models that were created based on a target variable of the f1 -score for the positive class of having a diagnosis in the test set. The f1 score was chosen in this example as the selection criteria as it is more difficult to correctly find positive diagnoses than no diagnoses due to imbalances that often occur in user health data and analyte data, for example in some cases there may be many more examples in such data of users not being diagnosed with a given condition than getting a diagnosis. Furthermore, use of the f1 -score means that the models will be more useful if geared towards accurately suggesting a positive diagnoses. Other selection criteria may be used, for example the true positive and true negative rates, otherwise known as sensitivity and specificity.
[0126] Figure 12 and Figure 13 show example flow diagrams 1100, 1200 of training machine learning models to predict the presence of reproductive condition\determine the likelihood or probability of a diagnosis of reproductive health condition. The two process flows differ in that the machine learning model of Figure 12 is trained using only features from the user health data and the machine learning model of Figure 13 is trained using features from the user health data and / or features from analyte data.
[0127] In more detail, a customer flow 1110,1210 may be used to obtain user health data and / or analyte data. Alternatively, this data may simply be received. The customer flow comprises the user completing a virtual or online health assessment 1112,1212 such as that described above. A dynamic decision tree such as decision tree 200 may then generate a bespoke user analyte panel using method 100 described above which is used to test 1114, 1214 a user blood sample to produce user analyte data. Data from the user analyte panel along with user health data from the user health assessment is provided to a clinician who provides a diagnosis for a reproductive health condition 1116.
[0128] Features from the user health data and / or analyte data are selected 1120, 1220 using for example method 700 described above and used to train 1130, 1230 a machine learning model 1132, 1232 for example using method 800 to provide a probability of diagnosis of the reproductive health condition based on the data comprising user health data and / or analyte data. The Clinician diagnosis is used as the target variable (prediction signal) in the training of the machine learning model.
[0129] Once a machine learning model has been trained and verified (e.g. a classification type model or a Bayesian Network model), the machine learning model can be used to provide a probabilityMikelihood of diagnosis of a reproductive health condition based on user health data and / or analyte data. The flow diagrams shown in Figure 14 and Figure 15 provide examples of this. The two process flows 1300,1400 differ in that the machine learning model 1332 of Figure 14 was trained using only features from the user health data and the machine learning model 1432 of Figure 15 was trained using features from the user health data and / or features from analyte data. As such the trained machine learning model of Figure 14 can provide a probabilityMikelihood of diagnosis of a reproductive health condition based solely on user health data without the need for analyte data. It will be appreciated that the trained machine learning model may be a classification type model or a Bayesian Network model.
[0130] In more detail, a customer flow 1310,1410 may be used to obtain user health data and / or analyte data for a user requiring a diagnosis of a reproductive health condition. Alternatively, this data may simply be received. The customer flow comprises the user completing a user health assessment 1312, 1412 such as that described above. A dynamic decision tree such as decision tree 200 may then generate a user analyte panel using method 100 described above which is used to test 1314, 1414 a user blood sample to provide user analyte data. Features used to train the machine learning model in processes 1100, 1200 are extracted 1320, 1420 from the user analyte panel and / or user health data and are provided to the trained machine learning model 1332, 1432 (the trained machine learning model having been trained for example using method 600, 800). The trained machine learning model provides 1334, 1434 a probability of diagnosis of the reproductive health condition based on the extracted features from the user health data and / or the user analyte data. This probability of diagnosis can then be provided, along with the user health data and analyte data, to a clinician report for the clinician to make the final assessment and complete the report 1316,1416. The machine learning probability provides a more accurate (meaning fewer changes and less clinician time needed) assessment and one that is more interpretable because it provides a probability of diagnosis, not just a binary indicator of yes / no for a given condition that would otherwise be provided by a clinician.
[0131] A number of trained machine learning models can be used to form a machine learning framework to provide a reproductive health assessment that can be used to determine a likelihood or probability of a number of reproductive health conditions based on the user health data and / or the user analyte data.
[0132] Example supervised learning model for predicting the probability of diagnosis of PCOS
[0133] In this example, a supervised machine learning model (of the classification type, such as a decision tree ensemble model) was trained using methods 600, 700 and 800, to identify PCOS using health data comprising a plurality of user health data and results from an analyte panel in which a PCOS diagnosis had been confirmed by a clinician. Datasets in which concurrent reproductive health conditions were identified were removed. The user health data comprised the following features: customer goals and objectives; demographics (age, ethnicity); period and cycle information (duration, bleeding description, regularity, cycle length); symptoms experienced and how frequently; contraception currently or recently used; medical history (previous diagnosed conditions, cancer, medications taken, past surgeries and cosmetic procedures, STIs); pregnancy and miscarriage history; lifestyle (smoking, vaping, drinking, drugs, stress, exercise, diet); and body variations (bmi, body shape); and the analyte data comprised the following features: analytes values; and analyte range indicators (e.g. low, high, in normal range).
[0134] Features for training a supervised machine learning model to predict a probability of diagnosis of PCOS were selected from the user health data and / or analyte data using method 700 described above.
[0135] In the example of the PCOS supervised machine learning model the following features were selected:
[0136] - analyte values: AMH, LH, FSH, TEST, DHEAS, SHBG; period features: irregular or no periods, cycle length, bleeding description, no periods due to contraception or hrt; symptoms: number of symptoms, acne, hair loss, hair growth (and where on body), ovulation pain, pelvic pain, anxiety, increased hunger, depressive mood, feeling cold, hot flushes, fatigue, irritability, bleeding between periods; cosmetic surgery: face fillers, if answered in the virtual health assessment, (and confirmation of no surgery, if answered); whether there is cosmetic surgery or previous diagnosis information; demographics and lifestyle: age, bmi, some body shape indicators, stress, exercise; number of questions answered; if the motivation for the test is getting help with symptoms.
[0137] This effectively reduced the number of features from the user health data and analyte data from hundreds of features to tens of features.
[0138] The selected features were also ranked according to their mutual information gain using method 700 described above. A training dataset was generated from the selected features and also included a predictive signal for each feature. In this example, the prediction signal\target variable was the positive diagnosis of PCOS provided by a Clinician.
[0139] Multiple decision tree ensemble models were then trained using the ranked selected features. To train the machine learning model the selected features were additively incorporated as decision trees into a machine learning model in accordance with rank for example using method 800 described in detail above. The multiple models were constructed using at least some of these selected features and different hyperparameters in use or based on a different ranking of features. The final model was selected based on the f1 -score for the positive class of having a diagnosis in the test set. The test data set comprised of 25% of the available data, stratified for the target variable and a Log loss function was used for verification of the models.
[0140] The table below shows the f1 scores for the final selected model as well as the f1 scores for a number of other models trained to predict the probability of diagnosis of other reproductive health conditions in a similar manner. These machine learning models form a machine learning framework trained to determine a likelihood of separate reproductive health conditions based on the user health data and / or the user analyte data.
[0141] Using the methods detailed above for training supervised machine learning models, supervised machine learning models may be trained to identify reproductive health conditions based on health data alone, or health data and analyte data. In particular supervised machine learning models may be trained to identify Polycystic ovary syndrome, Overt / subclinical hypothyroidism, Overt / subclinical hyperthyroidism, Premature Ovarian Insufficiency, Peri / Menopause, Hyperandrogenism, Hyperprolactinaemia, Hypothalamic Amenorrhea, ovulatory disorder, Endometriosis, pelvic issues (such as fibroids or adenomyosis), and Iron deficiency anaemia. The trained models may be used in the machine learning framework to provide determination of the likelihood of different reproductive health conditions based on user health data and / or analyte data. Bayesian Network models
[0142] While the machine learning frameworks and machine learning models have been discussed above in part in the context of ensemble / classification type machine learning models, other machine learning models may also be used to for determining a likelihood (for example a probability) of a reproductive health condition based on user health data and / or analyte data. An example of another suitable machine learning model that may be used for determining a probability of a reproductive health condition is a Bayesian Network model.
[0143] An example Bayesian Network model 1700 is shown in Figure 17. The Bayesian Network model comprises a plurality of nodes 1710, wherein each node is representative of a variable that is at least hypothesised to be causally linked to the probability / likelihood of a reproductive health condition with which the model is concerned. In the example shown, the Bayesian Network model is configured to determine the probability of a PCOS diagnosis. The nodes in this example Bayesian Network model represent the variables: age; weight; irregularity of periods; FSH analyte data; LH analyte data; and ultrasound data. It will be appreciated that additional variables can be included as additional nodes, but which are not shown in this simplified diagram. In addition, it may be the case that some variables have a diminutive effect on the probability of PCOS, and the node or nodes associated with these variables may be removed, e.g. as part of the training process.
[0144] Other Bayesian Network models may be configured similarly to determine the probability of other reproductive health conditions for example Polycystic ovary syndrome, Overt / subclinical hypothyroidism, Overt / subclinical hyperthyroidism, Premature Ovarian Insufficiency, Peri / Menopause, Hyperandrogenism, Hyperprolactinaemia, Hypothalamic Amenorrhea, ovulatory disorder, Endometriosis, pelvic issues (such as fibroids or adenomyosis), and Iron deficiency anaemia. The number and nature of the variables which are selected to be represented as nodes will differ depending on the reproductive health condition of interest. For example, the analytes LH and FSH may be considered relevant for PCOS but not for determining the probability of another reproductive health condition, e.g. hyperthyroidism. The selection of variables that are implemented as nodes may be determined / established initially at least by previous experience, for example by medical literature.
[0145] For example, in determining the probability of a diagnosis of PCOS the following features may be selected from user health data and / or user analyte data as variables represented as nodes in the Bayesian Network model: analyte values: for example, one or more of: AMH, LH, FSH, FT3, FT4, TEST, DHEAS, SHBG, OEST, and PROL; period features: for example, one or more of: irregular or no periods, cycle length, bleeding description, and no periods due to contraception or hrt; symptoms: for example, one or more of: number of symptoms, acne, hair loss, hair growth (and where on body), ovulation pain, pelvic pain, anxiety, increased hunger, depressive mood, feeling cold, hot flushes, fatigue, irritability, and bleeding between periods; demographics and lifestyle: for example, one or more of: age, ethnicity, bmi, some body shape indicators, stress, exercise, smoking, vaping, drugs and alcohol usage, and cosmetic surgery treatments; number of questions answered;
[0146] Test motivations and objectives: for example, getting help with symptoms, understanding fertility.
[0147] A Bayesian Network model also comprises edges extending between and connecting nodes. The simplified model 1700 shown in Figure 17 depicts a number of edges 1720 that extend between nodes 1710. Edges depict or represent a conditional dependency of one node upon another node connected by that edge, in other words the causal link between variables. The directional arrow of the edge indicating the hierarchy of the dependency. In this example, the “age node” 1712 is connected to the “weight node” via an edge 1722 with an arrow directed from the “age node” to the “weight node”, since weight may be causally linked with age. The “weight node” 1714 is then connected via edge 1724 to the PCOS node 1716 indicating a dependency of PCOS on weight (and by extension age) in the probability of a PCOS diagnosis.
[0148] The structure of the Bayesian Network model, i.e. the arrangement nodes and edges, may be determined at least initially by experience, e.g. by an expert. For example, by reference to medical literature. The arrangement of nodes and edges may then be refined and / or updated using the training methods described below.
[0149] It should be understood that other network structures are possible, and may in practice be more complex e.g. involving a greater number of variables and dependencies; the example shown in Figure 17 is a simplified example of a Bayesian Network model for determining the probability of PCOS. It will further be appreciated that the arrangement of nodes and edges will be different depending on the reproductive health condition for which the Bayesian Network model has been configured.
[0150] Associated with each node in the Bayesian Network model is a probability distribution. For example, in the example Bayesian Network model of Figure 17 the probability distribution for the “age node” may be [AGE 12, 13, 14 35 PROBABILITY 0.001 , 0.004, 0.006 0.23]. Where a node has multiple edges connected to it, the node may have an associated joint probability distribution. For example, the “weight node” 1714 in the example shown in Figure 17 may have a probability distribution [WEIGHT underweight, normal weight, overweight; PROBABILITY 0.2, 0.4 , 0.4]. The PCOS node may then have a probability given an age and weight. For example, P(PCOS| known age, known weight) for a positive diagnosis of PCOS (PCOS : YES) the probability may be, in an underweight user of age 40, (Age 40, weight under weight: p(yes) = 0.001), or in an overweight user of age 40, (Age 40, weight over weight: p(yes) = 0.005). The resulting Bayesian network is thus a graphical model representing the joint probability distribution of the set of variables in a directed acyclic graph (DAG).
[0151] As discussed, the probability distributions may be initially determined by experience, for example from preexisting medical literature and then updated using the training methods described below. The initial probability distribution values are referred to as ‘prior’ distributions. The priors associated with each node may be updated through training of the model to provide updated probability distributions referred to as ‘posterior’ distributions. Through training of the Bayesian Network model the accuracy of the model may be improved by updating the structure and probability distributions of the model.
[0152] Bayesian Network model training
[0153] As noted above the structure of a Bayesian Network model such as that depicted in Figure 17 may initially be determined by a set of assumptions, for example the number of and type of variables represented as nodes and the positioning of the edges connecting those nodes may be initially based on prior experience, for example on pre-existing medical literature and / or by an expert. Several different models having different structures may be created in such a way and from these an optimal model may be selected using selection criteria, for example a Bayesian information criteria (BIC) or a ‘likelihood function’, as is known in the art.
[0154] The probability distribution of each of the nodes in the Bayesian Network model may be trained for example through the sampling of an appropriate dataset, for example sampling of a dataset of user health data. A common sampling technique uses a Markov chain Monte Carlo (MCMC) algorithm. In the example of a model for determining the probability of a reproductive health condition, the dataset may comprise previous user health data and / or analyte data associated with the reproductive health condition of interest to the model. Each previous user health data and / or analyte data may also include a prediction signal denoting for each whether there was a positive or negative diagnosis of the reproductive health condition of interest. The previous user health data and / or analyte data may comprise similar data to the user health data and / or analyte data as described above.
[0155] The dataset is sampled, and based on the sampling, the probability distributions of the nodes of the Bayesian Network model are updated. For example, the nodes of the Bayesian network may be updated to incorporate a posterior probability determined using Bayesian inference. To this end, the Bayesian Network comprising initial edges and probabilities can be thought of as an initial hypothesis P(H). The hypothesis can be updated using observed data (or evidence P(E)) and determining a probability of the observed data given the initial hypothesis (a likelihood P(E|H)). A relationship between the prior probability (or hypothesis, P(H)), new observed data P(E), the likelihood of the observed data given the hypothesis PE|H) and the updated hypothesis (the posterior probability, P(H|E) ), can be given by Bayes theorem: P(H|E) = P(E|H) P(H) / P(E), as is known in the art.
[0156] For example, the sampling of the dataset and associated convergence of the Bayesian Network model (i.e. the shift from a prior to posterior distribution for a given variable) may indicate a stronger or weaker dependency of a reproductive health condition on a particular variable than assumed or adopted in the priors of the Bayesian Network model, indicating that an updated probability distribution would provide an improved model based on the sampled dataset. Implementation of the Bayesian Network model
[0157] Once trained the Bayesian Network model will comprise a plurality of nodes 1710, edges 1720 and probability distributions. The trained Bayesian Network model may then be used to determine the probability of a diagnosis of a reproductive health condition. For example, a user may provide current user health data and / or analyte data for example using one of the methods described above. The current user health data and / or analyte data may then be entered into the trained Bayesian Network model to establish the probability of a diagnosis of a reproductive health condition. In the example shown in Figure 17, the current user data and / or analyte data may comprise user data associated with the variables represented by the nodes of the model, which in this example includes information on irregularity of periods, age, weight, FSH analyte data, LH analyte data, and ultrasound scan data. The data may then be entered into the trained Bayesian Network model and based on the probability distributions and edge connections a probability of a reproductive health condition (for example PCOS) calculated.
[0158] A Bayesian Network may be constructed and trained for each reproductive health condition of interest. For example, a number of Bayesian Network models each constructed and trained to determine the probability of a different reproductive health condition may be used to form a machine learning framework. The machine learning framework may comprise ensemble machine learning models as described above, Bayesian network models, or may comprise both ensemble machine learning models as described above and Bayesian network models. The framework may comprise an ensemble machine learning model and a Bayesian network model for the same reproductive health condition so that a comparison between the two may be performed. A simplified machine learning framework 1800 comprising a plurality of Bayesian Network models for determining a probability of different respective reproductive health conditions is shown in Figure 18.
[0159] Example results from a machine learning framework
[0160] Figure 19 shows example results from a machine learning framework that comprises several trained Bayesian Network models each configured to determine the probability of a diagnosis of reproductive health condition (such as the framework 1800 of Figure 18); the Bayesian Network models having been trained using previous user health and / or analyte data. The graph 1900 in Figure 19 shows the probability of each of the reproductive health conditions calculated by the machine learning framework based on a set of current user health data and / or analyte data, including irregular periods, high LH analyte, and high FSH analyte, and scan features. In this example the machine learning framework comprises models configured to determine the probability of a diagnosis of PCOS, Hypothyroidism, Hyperthyroidism, Adenomyosis, OvarianCysts, Fibroids and Malignancy. Current user health data and / or analyte data is entered into the machine learning framework and the probability of the reproductive health conditions is evaluated using the trained Bayesian Network models. In the example shown, a posterior probability (P) of 0.475 of PCOS was calculated from the trained Bayesian Network model. The framework also provided the probability of Hypothyroidism (P=0.079), Hyperthyroidism (P=0.075), Adenomyosis (P=0.057), OvarianCysts (P=0.041), Fibroids (P=0.037) and Malignancy (P=0.027).
[0161] An advantage of the Bayesian network models described above is that they provide not only a probability of a user having a particular reproductive health condition but also an indication of how much each node (variable) influences that probability. For example, a Bayesian Network trained to determine the probability that a user has, for example PCOS, may indicate that irregular periods, high LH, and high FSH are strong indicators of that particular condition being present.
[0162] Figure 16 shows a schematic diagram of a system 1000 for implementing the above-described methods. The system 1000 comprises one or more processing devices 1100 configured with instructions that, when executed by the one or more processing devices 1100, cause the one or more processing devices 1100 to perform any of the above-described methods. The system 1000 may comprise a computer-readable medium 1200 in communication with the one or more processing devices 1100 storing the instructions. The processing devices 1100 may include a user computing device and / or one or more servers.
[0163] Accordingly, aspects of the present disclosure may be implemented entirely in hardware, entirely in software (including firmware, resident software, micro-code, etc.) or combining software and hardware implementations that may all generally be referred to herein as a “unit,” “module,” or “system”. Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having instructions or computer readable program code embodied thereon. Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including wireless, wireline, optical fibre cable, RF, or the like, or any suitable combination of the foregoing.
[0164] The computer readable medium 1200 may include a mass storage, a removable storage, a volatile read- and write memory, a read-only memory (ROM), or the like, or any combination thereof. Exemplary mass storage may include a magnetic disk, an optical disk, a solid-state drive, etc. Computer program code or instructions for carrying out disclosed methods may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB. NET, Python or the like, conventional procedural programming languages, such as the "C" programming language, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages such as Python, Ruby and Groovy, or other programming languages.
[0165] The disclosed methods and / or program code may execute entirely on a user's computing device, partly on a user's computing device, as a stand-alone software package, partly on a user's computing device and partly on a remote computer, or entirely on a remote computer or server. In the latter scenarios, the remote computer / server may be connected to a user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external server / computer (for example, through the Internet using an Internet Service Provider) or in a cloud computing environment or offered as a service such as a Software as a Service (SaaS).
[0166] It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention. Each feature disclosed in the description, and (where appropriate) the claims and drawings may be provided independently or in any appropriate combination. Any feature described in relation to any one example may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or any combination of any other of the examples. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.
[0167] Although the appended claims are directed to particular combinations of features, it should be understood that the scope of the disclosure of the present invention also includes any novel feature or any novel combination of features disclosed herein either explicitly or implicitly or any generalisation thereof, whether or not it relates to the same invention as presently claimed in any claim and whether or not it mitigates any or all of the same technical problems as does the present invention. Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.
Claims
CLAIMS1 . A method of training a machine learning model for determining a likelihood of a reproductive health condition based on user health data and / or analyte data, the method comprising: receiving user health data and / or analyte data; selecting features in the user health data and / or analyte data for training the machine learning model; generating a training data set comprising the selected features; and training the machine learning model using the generated training data set to determine a likelihood of said reproductive health condition based on the user health data and / or the user analyte data.
2. The method of claim 1 , wherein the training data comprises at least one prediction signal for each of the user health data and / or analyte data.
3. The method of claim 2, wherein the prediction signal is based on a prior diagnosis based on the user health data and / or analyte data.
4. The method of any preceding claim, wherein the user health data excludes health data and / analyte data in which concurrent reproductive health conditions have been previously diagnosed.
5. The method of any preceding claim, wherein selecting features from the user health data and / or analyte data comprises selecting all features or a subset of features from the user health data and / or analyte data.
6. The method of any preceding claim, wherein selecting features from the user health data and / or analyte data comprises: selecting features from the user health data and / or analyte data that meet a threshold for statistical significance.
7. The method of claim 6, wherein the threshold of statistical significance is a threshold p-value.
8. The method of claim 6 or claim 7, wherein selecting features from the user health data and / or analyte data comprises: calculating for the selected features the mutual information gain between each selected feature and a target variable, wherein the target variable is either a positive or a negative diagnosis of the reproductive health condition; comparing the calculated mutual information gain to a baseline mutual information gain; and deselecting features which do not meet a threshold when compared to the baseline mutual information gain.
9. The method of any preceding claim, wherein training the machine learning model using the generated training data set comprises: adding a selected feature from the training data set into the machine learning model; testing, after each addition of a selected feature, the machine learning model against the target variable; and repeating the method steps until a threshold condition between iterations of the machine learning model has been met.
10. The method of claim 9, wherein the threshold condition is based on a convergence of a performance metric between iterations of the model against the target variable.11 . The method of claims any preceding claim, wherein the selected features are ranked in order of mutual information gain against the target variable, and wherein the selected features are added into the machine learning model in order of ranking.
12. The method of any preceding claim, comprising: tuning, after each addition of a selected feature, a set of parameters of the machine learning model to increase or maximise a performance metric of the machine learning model against a target variable.
13. The method any preceding claim, wherein the machine learning model is or comprises an ensemble machine learning model, preferably an ensemble classification model or a decision tree ensemble model.
14. The method any one of claims 1 to 5, wherein the machine learning model is or comprises a Bayesian Network model.
15. A computer implemented method of reproductive health assessment, the method comprising: receiving user health data comprising medical history and symptom variables of a user and analyte data obtained from results of a user analyte panel test; providing the user health data and analyte data to a model comprising a logic algorithm encoded with clinical diagnostic criteria for a plurality of reproductive health conditions; and determining, using the model, a likelihood or probability of at least one reproductive health condition based on the user health data and the user analyte data.
16. The computer implemented method of claim 15, further comprising encoding the user health data and analyte data, and providing the encoded user health data and analyte data to the model; and / or wherein the model comprises a supervised learning model.
17. The computer implemented method of claim 16, wherein encoding the user health data and analyte data comprises: extracting from the user health data and analyte data a set of predefined variables to define a plurality of logic blocks, and wherein the logic algorithm is configured to assemble the logic blocks basedon clinical guidelines for a plurality of reproductive health conditions to determine a likelihood or probability of at least one reproductive health condition based on the user health data and the user analyte data.
18. A computer implemented method of reproductive health assessment, the method comprising: providing user health data including a plurality of medical history and symptom variables and / or analyte data to a machine learning framework comprising at least one machine learning model, each machine learning model trained to determine a probability of a separate reproductive health condition based on the user health data and / or the user analyte data; and determining, using the machine learning framework, a probability of at least one reproductive health condition based on the user health data and / or the user analyte data.
19. The method of claim 18, wherein the at least one machine learning model is trained using the method of any one of claims 1 to 14.
20. The method of claim 18 or 19, wherein the at least one machine learning model comprises an ensemble machine learning model, preferably an ensemble classification model and / or a decision tree ensemble model.
21. The method of claim 18 or 19, wherein the at least one machine learning model comprises a Bayesian Network model.
22. The method of any of claims 15 to 21 comprising: generating a user health assessment report based on the determined probability of at least one reproductive health condition.
23. The method of claim 22, wherein the method comprises encoding the user health data and / or analyte data and the probability of at least one reproductive health condition, and wherein generating the user health assessment report comprises selecting predefined text for inclusion in the report based on the encoded health data, analyte data and probabilities; and optionally or preferably, wherein the report includes a relative uncertainty of the determined probabilities of at least one reproductive health condition.
24. The method of any preceding claim, wherein the reproductive health condition is at least one of: Polycystic ovary syndrome, Overt / subclinical hypothyroidism, Overt / subclinical hyperthyroidism, Premature Ovarian Insufficiency, Peri / Menopause, Hyperandrogenism, Hyperprolactinaemia, Hypothalamic Amenorrhea, ovulatory disorder, Endometriosis, pelvic issues (such as fibroids or adenomyosis), and Iron deficiency anaemia.
25. The methods of claims 1 to 24, comprising obtaining the analyte data, wherein obtaining the analyte data comprises: receiving user health data including a plurality of medical history and symptom variables; determining at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables; andgenerating an analyte panel based on the determined at least one target analyte.
26. A computer implemented method of generating an analyte panel for a reproductive health assessment, the method comprising: receiving user health data including a plurality of medical history and symptom variables; determining at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables; and generating an analyte panel based on the determined at least one target analyte.
27. The method of claim 25 or 26, wherein the step of determining at least one target analyte comprises: determining, using a supervised learning model, at least one target analyte associated with at least one reproductive health condition based on one or more of the medical history and symptom variables.
28. The method of claim 27, further comprising: selecting and / or excluding, using the supervised learning model, at least one target analyte from a predefined set of target analytes based on the one or more medical history and / or symptom variables.
29. The method of claim 27 or 28, wherein the supervised learning model comprises a plurality of decision nodes to determine the at least one target analyte, each decision node configured to select and / or exclude at least one target analyte from a predefined set of target analytes based on a subset of the user health data and one or more clinical criteria associated with the respective decision node.
30. The method of claim 29, wherein the one or more clinical criteria associated with at least some of the decision nodes comprise a minimum set of clinical diagnostic criteria for a given reproductive health condition compiled from one or more clinical data sources.31 . The method of claim 29 or 30, wherein the supervised learning model comprises a decision tree with a plurality of decision nodes, and wherein the step of determining comprises: traversing the decision tree based on the user health data; and selecting and / or excluding, at each decision node of the decision tree encountered, at least one target analyte from the predefined set of target analytes based on a subset of the user health data and one or more clinical criteria associated with the respective decision node.
32. The method of any of claims 28 to 31 , wherein the analytes in the predefined set of target analytes are grouped into a plurality of groups of one or more complementary target analytes, and wherein selecting and / or excluding at least one target analyte from a predefined set of target analytes comprises selecting and / or excluding at least one group of target analytes from the predefined set of target analytes.
33. The method of claim 32, wherein the plurality of groups comprise: a) Anti-Mullerian hormone (AMH);b) Cycling hormones (oestradiol (E2), luteinising hormone (LH) and follicle-stimulating hormone (FSH); c) Thyroid hormones (FT4 and TSH); d) Androgens (testosterone (T), dehydroepiandrosterone sulphate (DHEAS) and sex hormone- binding globulin (SHBG); and e) Prolactin (PRL).
34. The method of any of claims 29 to 33, wherein the supervised learning model is configured to generate an analyte panel that includes up to and including 10 analytes through selection and / or exclusion of at least one of a plurality of groups of one or more complementary target analytes.
35. The method of any of claims 25 to 34, wherein the analyte panel includes a minimum of three analytes, preferably Anti-Mullerian hormone (AMH) and Thyroid hormones (FT4 and TSH).
36. The computer implemented method of any of claims 25 to 35, wherein the step of determining comprises: determining, using a supervised learning model, a likelihood or risk of at least one of a plurality of reproductive health conditions based on the one or more medical history and symptom variables and one or more clinical diagnostic criteria, and selecting and / or excluding at least one target analyte from a predefined set of target analytes based on the determined likelihood or risk of the at least one of a plurality of reproductive health conditions.
37. The method of any of claims 25 to 36, further comprising obtaining the user health data using a dynamic decision tree, comprising: sending, to a user device, a series of queries relating to medical history and symptom variables; and receiving a series of responses for each query, wherein at least some of the queries are determined, using the dynamic decision tree, based on the responses to one or more previous queries; and wherein the step of determining is performed concurrently with the obtaining of user health data.
38. A system comprising one or more processing devices configured with instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform the method of any preceding method claim.
39. A computer readable non-transitory storage medium comprising program instructions that, when executed by one or more processing devices, cause the one or more processing devices to perform the methods of any preceding method claim.
40. A computer readable non-transitory storage medium comprising program instructions that, when executed by one or more processing devices, cause the one or more processing devices to operate a machine learning model for determining a probability of a reproductive health condition based on userhealth data and / or analyte data, wherein the machine learning model is trained using the method of any one of claims 1 to 14.41 . The method of claim 40, wherein the machine learning model comprises an ensemble machine learning model, preferably an ensemble classification model and / or a decision tree ensemble machine learning model.
42. The method of claim 40, wherein the machine learning model comprises a Bayesian Network model.
Citation Information
Patent Citations
Methods for assessing the probability of achieving ongoing pregnancy and informing treatment therefrom
US20200011883A1