Method, system and computer program product for improved analysis of image data

The method addresses the challenge of diagnosing ultra-rare genetic disorders by employing advanced neural networks for facial feature recognition and disease classification, enhancing accuracy and explainability in disease identification.

WO2025191122A1PCT designated stage Publication Date: 2025-09-18HUSTINX ALEXANDER
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/057011
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2025-03-14
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing methods for diagnosing ultra-rare genetic disorders based on facial features are challenging due to the scarcity of data and reliance on medical practitioners' experience, with current models lacking optimal explainability and accuracy for ultra-rare diseases.

Method used

A method using a machine learning model, such as a convolutional neural network or residual neural network, trained with additive angular margin loss, for improved facial feature recognition and disease classification, incorporating facial landmark alignment and fine-tuning with diverse datasets, and generating class activation maps for enhanced explainability.

Benefits of technology

The method provides improved accuracy and explainability in diagnosing ultra-rare genetic disorders by leveraging advanced neural networks and facial feature recognition, facilitating faster and more reliable disease identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025057011_18092025_PF_FP_ABST
    Figure EP2025057011_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a computer-implemented method, comprising receiving an element of input image data representative of an image. The method further comprises an estimating step comprising estimating for the element of input image data by means of a machine learning model a set of facial feature(s) and at least one of (i) a set of disease classification confidences and (ii) a set of disease similarity estimation values, each of the disease classification confidences and / or each of the disease similarity estimation values corresponding to a respective disease from a list of diseases. Further, the method comprises outputting a result of the estimating step and a pre-training step. The pre-training step comprises pre-training the machine learning model with a face recognition data set comprising a plurality of training image data elements representative of a face photo. The method further comprises a fine-tuning step. The model further comprises a feature vector part computing a feature vector based on the input image data, at least one fully connected facial-feature estimation layer, and a disease estimation component. The method comprises the at least one fully connected facial-feature estimation layer estimating a set of classification confidences for a list of facial features based on the feature vector. Estimating the set of facial feature(s) comprises selecting the facial features from the list of facial features based on the estimated classification confidence. The method also comprises the disease estimation component estimating at least one of the set of disease classification confidences and the set of disease similarity estimation values. The fine-tuning step further comprises obtaining the at least one fully connected facial-feature estimation layer and the at least one fully connected disease estimation layer based on training data for fine-tuning. Also disclosed is a system comprising a data-processing system. The system is configured for carrying out the method. Further, a computer program product comprising instructions which, when the program is executed by a data-processing system, cause the data-processing system to carry out the method is disclosed. Also, a use of the method or the system to diagnose a genetic disease is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method, system and computer program product for improved analysis of image data

[0002] [1] The present invention relates to the field of image recognition, more particularly to the field of face analysis.

[0003] [2] In the past years, treatment and diagnosis of ultra-rare diseases have seen important progress. At present, it appears that most ultra-rare diseases are genetic, with some of them starting in childhood. While there is no general definition of ultra-rare diseases, in the European Union, ultra-rare diseases are defined as diseases with a prevalence of at most 5 per 10,000 inhabitants.

[0004] [3] Ferreira, Carlos R. "The burden of rare diseases." American journal of medical genetics Part A 179.6 (2019): 885-892, estimate that more than 6% of the global population is affected by rare genetic disorders. In view of the low prevalence and the overall diversity of genetic ultra-rare diseases, it may be difficult to efficiently reach a correct diagnosis.

[0005] [4] According to Hustinx, Alexander, et al. "Improving Deep Facial Phenotyping for Ultra- rare Disorder Verification Using Model Ensembles." Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision (2023), more than a third of patients suffering from ultra-rare diseases wait for over five years to receive a diagnosis, often referred to as the "diagnostic odyssey". Further, many ultra-rare genetic disorders are associated with distinctive dysmorphic facial features (gestalt), which may partially serve for diagnosis. Diagnosing based on said facial features may however be challenging for medical practitioners in case of ultra-rare diseases, as recognizing these features relies on a medical practitioner's experience and is more difficult if the medical practitioner has not seen the disease before.

[0006] [5] Previous work to Hustinx et al. (2023), such as GestaltMatcher, utilized representation vectors produced by a DCNN similar to AlexNet to match patients in high-dimensional feature space to support "unseen" ultra-rare disorders. However, the architecture and dataset used for transfer learning in GestaltMatcher have become outdated. Further, the disclosure argues that a way to train the model for generating better representation vectors for unseen ultra-rare disorders has not yet been studied. Because of the overall scarcity of patients with ultra-rare disorders, it is infeasible to directly train a model on them. Hence, the disclosure analyzes the influence of replacing GestaltMatcher DCNN with a state-of- the-art face recognition approach, iResNet with ArcFace. Additionally, the authors experimented with different face recognition datasets for transfer learning. Furthermore, the disclosure proposes test-time augmentation, and model ensembles that mix general face verification models and models specific for verifying disorders to improve the disorder verification accuracy of unseen ultra-rare disorders. [6] However, the results of Hustinx (2023) do not provide optimal results with respect to explainability of results, that is, with respect to generating at least one output or indicator based on which a user can determine which features were decisive for an estimation of the model.

[0007] [7] US10667689B2, US10470659B2 and US10327637B2 disclose systems, methods, and computer-readable media for performing image processing in connection with phenotypic analysis. For example, at least one processor may be configured to receive electronic numerical information corresponding to pixels reflective of at least one external soft tissue image of an individual and access geographically dispersed genetic information stored in a database. The geographically dispersed genetic information may include numerical data that correlates anomalies in pixels in soft tissue images of a plurality of geographically dispersed individuals to specific genes or to specific genetic variants. The at least one processor may also be configured to compare the electronic numerical information for the individual with the numerical data of the geographically dispersed genetic information stored in a database, to determine at least a likelihood that the individual has a specific genetic variant, and prioritize, based on the comparison, one or more genetic variants according to likelihood of pathogenicity.

[0008] [8] Robinson, Peter N., and Stefan Mundlos. "The human phenotype ontology." Clinical genetics 77.6 (2010): 525-534, suggest a standardized, controlled vocabulary allowing phenotypic information to be described in an unambiguous fashion in medical publications and databases. The Human Phenotype Ontology (HPO) is being developed in an effort to provide such a vocabulary. The use of an ontology to capture phenotypic information allows the use of computational algorithms that exploit semantic similarity between related phenotypic abnormalities to define phenotypic similarity metrics, which can be used to perform database searches for clinical diagnostics or as a basis for incorporating the human phenome into large-scale computational analysis of gene expression patterns and other cellular phenomena associated with human disease. However, Robinson (2010) does not discuss facial recognition for ultra-rare diseases or a suitable machine learning model.

[0009] [9] Hsieh, Tzung-Chien, et al. "GestaltMatcher facilitates rare disease matching using facial phenotype descriptors." Nature genetics 54.3 (2022): 349-357, discusses the use of computer-aided next-generation phenotyping tools for diagnosing monogenic disorders causing craniofacial abnormalities with characteristic facial morphology. Hsieh et al. particularly discuss using a deep convolutional neural network based on the DeepGestalt framework.

[0010]

[0010] While the prior art approaches may be satisfactory in some regards, they have certain shortcomings and disadvantages.

[0011] It is therefore an object of the invention to overcome or at least alleviate the shortcomings and disadvantages of the prior art. More particularly, it is an object of the present invention to provide a method, system and computer program product for improved analysis of image data.

[0011]

[0012] It is an optional objective of the present invention to provide a method, system and computer program product for improved feature recognition in facial image data.

[0012]

[0013] It is another optional objective of the present invention to provide a method, system and computer program product for improved estimation of a disease based on facial image data.

[0013]

[0014] In a first embodiment, a method is disclosed. The method comprises receiving an element of input image data representative of an image. The method further comprises an estimating step. The estimating step comprises estimating, for the element of input image data, by means of a machine learning model a set of facial feature(s), particularly facial feature(s) indicative of phenotypic abnormalities, and at least one of (i) a set of disease classification confidences and (ii) a set of disease similarity estimation values, each of the disease classification confidences and / or each of the disease similarity estimation values corresponding to a respective disease from a list of diseases.

[0014]

[0015] The method further comprises outputting a result of the estimating step, such as an estimated presence of a disease from the set of diseases.

[0015]

[0016] The model further comprises a feature vector part computing a feature vector based on the input image data.

[0016]

[0017] The feature vector part may, for example, comprise convolutional layers of a convolutional neural network and, for example, at least one or a plurality of fully connected layer(s) of such a convolutional neural network. Instead of or in addition to the fully connected layer(s), the feature vector part may also comprise a pooling layer.

[0017]

[0018] In other words, in case of a convolutional neural network, the feature vector part may perform a dimensionality reduction of data generated by a preceding part of the convolutional neural network.

[0018]

[0019] In other words, the estimating step may comprise estimating for the element of image data the set of facial feature(s) and for each of the diseases from the list of diseases at least one of a disease classification confidence and a disease similarity estimation value.

[0020] The term "estimating" is intended to refer to determining by means of a machine learning model. This may for example include associating different mutually exclusive options, such as "Disease A present" and "Disease A absent", with a measure for a probability, and outputting the option with the higher or highest probability.

[0019]

[0021] The term "facial features" is intended to refer to features of a face of a person, particularly to human-intelligible features indicative of a phenotypical abnormality. In other words, the term "facial features" is intended to abnormal facial features. For example, the facial features may be indicated by according to the Human Phenotype Ontology (HPO) as discussed in Robinson (2010) or a subset thereof. In other words, the term "facial features" may refer to HPO-terms, particularly to HPO leaf nodes.

[0020]

[0022] However, in other cases, non-leaf nodes, that is, intermediate nodes, may be considered indicative of facial features. This may be the case, e.g., if there are not enough cases of patients with HPO terms from leaf nodes under the intermediate nodes to train on these leaf nodes. There might be a lack of training data relating to leaf nodes.

[0021]

[0023] With reference to the HPO, it will be understood that the HPO has roughly a tree-like shape. For example, a classification may be described as: "All" (root I level 0) "Phenotypic abnormality" (level 1) "Abnormality of the head and neck" (level 2) "Abnormality of the head" (level 3) - "Abnormality of the face" (level 4) - "Abnormality of the nose" (level 5) ... "Bulbous nose" (leaf HPO / level 9). It will further be understood that, according to embodiments of the present invention, the root may not necessarily be the "level 0". In an example, the root may be "Abnormality of the face" (level 3). This may be at least because embodiments of the present invention are generally directed to facial photos and facial feature(s).

[0022]

[0024] The machine learning model may comprise a neural network. The neural network may be configured for image analysis. For example, the neural network may be a convolutional neural network, such as a deep convolutional neural network (DCNN).

[0023]

[0025] In particular, the neural network may be a residual neural network, as discussed in He, Kaiming, et al. "Deep residual learning for image recognition." Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.

[0024]

[0026] The neural network may also be an improved residual network, as discussed in Duta, lonut Cosmin, et al. "Improved residual networks for image and video recognition." 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021.

[0025]

[0027] The residual neural network, particularly the improved residual network, may yield improved image analysis and thus improved recognition and / or estimation.

[0028] Additionally or alternatively to the convolutional neural network, the machine learning model may comprise an image transformer instead.

[0026]

[0029] The machine learning model may further comprise a plurality of lookup tables. The plurality of lookup tables may comprise a lookup table for each facial feature and each disease of the set of diseases. In other words, there may be two lookup tables, namely a lookup table for the facial feature(s) and a lookup table for the disease(s).

[0027]

[0030] Determining the facial feature(s) and the presence of the disease from the set of diseases may be optionally advantageous, as the facial feature(s) may comprise inherent information relating to the disease. This may particularly apply in case of ultra-rare genetic diseases. Hence, using a common feature space for the facial feature(s) and the presence of the disease may optionally advantageously allow for improved model performance, as features comprising an underlying causal relation are estimated based on a common feature space.

[0028]

[0031] The model may comprise at least one fully connected facial-feature estimation layer. The method may comprise the at least one fully connected facial-feature estimation layer estimating a set of classification confidences for a list of facial features based on the feature vector. Estimating the set of facial feature(s) may comprise selecting the facial features from the list of facial features based on the estimated classification confidence. For example, estimating the set of facial feature(s) may comprise applying a threshold to the classification confidence and adding the feature(s) above the threshold to the set.

[0029]

[0032] The model may comprise a disease estimation component. The method may comprise the disease estimation component estimating at least one of the set of disease classification confidences and the set of disease similarity estimation values.

[0030]

[0033] The disease estimation component may comprise at least one fully connected disease estimation layer. The method may comprise the at least one fully connected disease estimation layer estimating the set of disease classification confidences.

[0031]

[0034] The disease estimation component may comprise a disease clustering component. The method may comprise the disease clustering component estimating the set of disease similarity estimation values.

[0032]

[0035] The method may comprise a pre-training step. The pre-training step may comprise pre-training the machine learning model, particularly the convolutional neural network, with a face recognition data set comprising a plurality of training image data elements representative of a face photo.

[0036] Obtaining the machine learning model, particularly the convolutional neural network, may comprise using additive angular margin loss (ArcFace) as loss function. The additive angular margin loss is defined as where N is the batch size, representing the number of samples in a mini-batch, n is the number of classes in the dataset, yi is the true class label for the i-th sample in the batch, s is the scaling factor that controls the spread of feature embeddings, 0 is the angle between the weight W and feature x, and m is the margin penalty, which may also be referred to as angular margin.

[0033]

[0037] The additive angular margin loss function is in an exemplary manner discussed in- depth in Deng, Jiankang, et al. "Arcface: Additive angular margin loss for deep face recognition." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019. and An, Xiang, et al. "Killing two birds with one stone: Efficient and robust training of face recognition cnns by partial fc." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022.

[0034]

[0038] Using ArcFace as loss function may be optionally advantageous, as it may adjust learning so that features are learnt in a discriminative and compact manner. In other words, optionally advantageously, the network may be able to differentiate better between different faces and may be able to cluster similar faces next to each other.

[0035]

[0039] In some embodiments, batch normalization may not be used when evaluating the loss function. In particular, batch normalization may not be performed after a representation layer and before fully connected layers of the network in case of neural networks, such as residual networks or improved residual networks.

[0036]

[0040] The face recognition data set may comprise at least 100.000 images corresponding to at least 10.000 individuals.

[0037]

[0041] Examples of such data sets are the CASIA data set or GLINT360K. However, diverse data sets appear to yield better results in pre-training.

[0038]

[0042] The method may comprise an aligning step. The aligning step may comprise identifying a plurality of facial landmarks in the input image data and transforming the input image data to align the facial landmarks with pre-defined positions before performing the estimating step.

[0043] The identifying of the plurality of facial landmarks in the input image data may be based on the RetinaFace algorithm disclosed in Deng, Jiankang, et al. "Retinaface: Single- stage dense face localisation in the wild." arXiv preprint arXiv: 1905.00641 (2019).

[0039]

[0044] The aligning step may comprise using at least five facial landmarks, such as left and right eyes, nose and left and right mouth corners.

[0040]

[0045] The aligning step may comprise using an affine transformation for transforming the input image data.

[0041]

[0046] The pre-training step may comprise further performing the aligning step for each training image data element of the face recognition data set.

[0042]

[0047] The method may comprise a fine-tuning step.

[0043]

[0048] The fine-tuning step may comprise obtaining the at least one fully connected facialfeature estimation layer based on training data for fine-tuning.

[0044]

[0049] The fine-tuning step may further comprise obtaining the at least one fully connected disease estimation layer based on training data for fine-tuning.

[0045]

[0050] The fine-tuning step may comprise at least obtaining the feature vector part of the machine learning model.

[0046]

[0051] Obtaining the feature vector part, the at least one fully connected facial-feature estimation layer and / or the at least one fully connected disease estimation layer may be intended to refer to obtaining adapted weights for these layers and / or part. As the person skilled in the art will easily understand, these layers and / or this part usually comprise initial values at a begin of the fine-tuning step that are modified or replaced during the fine- tuning step.

[0047]

[0052] The fine-tuning step may comprise using at least one of cross entropy, binary cross entropy and weighting.

[0048]

[0053] Below reproduced are the equations for weighted binary cross entropy loss: in,c= -wn,c[pcyn,c ■ log CT (xn c.) + (1 - ync) ■ log (1 - cr(nc))], (eq. 2) wherein wn edenotes the class weight, pcis the positive weight, xi denotes the predicted class and yi denotes the true class.

[0054] Further, the class weights can be calculated as where c is the class index and D is the distribution of the training set. Thus, Dcis the distribution for class c.

[0049]

[0055] The positive weights can be calculated as

[0050]

[0056] As in some embodiments, in the training data, an indication of a facial feature may either be present or not, the distribution may be a count of a number of positive cases, i.e., cases where the facial feature is present.

[0051]

[0057] The fine-tuning step may comprise obtaining the at least one fully connected facialfeature estimation layer by using binary cross entropy.

[0052]

[0058] The fine-tuning step may comprise obtaining the at least one fully connected disease estimation layer by using cross entropy.

[0053]

[0059] In the fine-tuning step, obtaining the at least one fully connected disease estimation layer may comprise positive weighting, particularly positive weighting and class weighting.

[0054]

[0060] The fine-tuning step may comprise training the machine-learning model to estimate the set of facial feature(s) and the set of disease classification confidences together. This approach may also be referred to as 1-stage training of a hybrid model.

[0055]

[0061] The fine-tuning step may comprise training the machine learning model to first learn the set of facial feature(s) or the disease, and to learn the respective other in a subsequent step. This approach may also be referred to as 2-stage training of a hybrid model.

[0056]

[0062] The fine-tuning step may comprise first training the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer, while not modifying weights of the at least one fully connected facia I -feature estimation layer. The fine-tuning step may additionally comprise to then train the at least one fully connected facial-feature estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer.

[0057]

[0063] The fine-tuning step may alternatively comprise first training the fully connected layer(s) of the feature vector part and the at least one fully connected facial-feature estimation layer, while not modifying weights of the at least one fully connected disease estimation layer, and then training the at least one fully connected disease estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one facial-feature estimation layer.

[0058]

[0064] The term "frozen" is intended to mean that during a respective training step, weights of the frozen element(s) of the machine learning model are not modified. In other words, when weights are not modified in a certain training step, they may be referred to as being frozen.

[0059]

[0065] The fine-tuning step may comprise using a Reduce-Learning-Rate-on-Plateau- Learning Rate Scheduler.

[0060]

[0066] The method may comprise determining, based on a result of the estimating step, particularly an estimated disease, and the set of facial feature(s) estimated in the estimating step, a similarity between the estimated set of facial feature(s) and a set of facial features typically associated with the estimated disease. For example, a similarity between estimated HPO terms and the estimated ultra-rare disease may be determined.

[0061]

[0067] Thus, verifying a result of the method by a health practitioner may optionally advantageously be simplified.

[0062]

[0068] Further, the similarity may optionally advantageously allow to estimate an uncertainty and / or probability of the presence of the estimated disease, thus further facilitating verification of the diagnosis.

[0063]

[0069] The similarity may be determined, e.g., based on CADA. CADA is discussed i.a. in Peng, Chengyao & Dieck, Simon & Schmid, Alexander & Ahmad, Ashar & Knaus, Alexej & Wenzel, Maren & Mehnert, Laura & Zirn, Birgit & Haack, Tobias & Ossowski, Stephan & Wagner, Matias & Brunet, Theresa & Ehmke, Nadja & Danyel, Magdalena & Rosnev, Stanislav & Kamphans, Tom & Nadav, Guy & Fleischer, Nicole & Frohlich, Holger & Krawitz, Peter. (2021). CADA: Phenotype-driven gene prioritization based on a case-enriched knowledge graph. 10.1101 / 2021.03.01.21251705, which is incorporated herein in its entirety.

[0064]

[0070] Determining the similarity may comprise retrieving the set of facial features typically associated with the estimated disease from a database comprising for each disease of the set of diseases at least one facial feature associated with the respective disease.

[0065]

[0071] The method may comprise generating at least one class activation map of the machine learning model. By providing a class activation map, optionally advantageously, features relevant for the model's estimation may be provided to the user, allowing for easier and / or more reliable verification of the estimation.

[0072] The at least one class activation map of the machine learning model may be at least one gradient-weighted class activation map. Thus, optionally advantageously, the activation map can be generated for a broader range of neural networks, allowing for a more flexible selection of the neural network.

[0066]

[0073] The generation of the gradient-weighted class activation map may be based on Selvaraju, Ramprasaath R., et al. "Grad-cam : Visual explanations from deep networks via gradient-based localization." Proceedings of the IEEE international conference on computer vision. 2017.

[0067]

[0074] The at least one class activation map may relate to the at least one fully connected facial-feature estimation layer.

[0068]

[0075] The at least one class activation map may relate to the at least one fully connected disease estimation layer.

[0069]

[0076] The at least one class activation map may relate to the at least one fully connected facial-feature estimation layer and the at least one fully connected disease estimation layer.

[0070]

[0077] The element of input image data representative of the image may be representative of a frontal face photo.

[0071]

[0078] The at least one or the plurality of facial feature(s) may be at least one or a plurality of HPO-terms.

[0072]

[0079] The disease whose presence is estimated by the machine learning model may be selected from the list of diseases comprising human genetic diseases, such as rare or ultra- rare human genetic diseases. In particular, the disease whose presence is estimated by the machine learning model may be selected from a set of ultra-rare human genetic diseases causing recognizable facial features.

[0073]

[0080] By also learning the facial feature(s), optionally advantageously, the feature space of the model is conditioned to consider features and / or regions in the image that are more relevant for a medical practitioner.

[0074]

[0081] The method may be a computer-implemented method.

[0075]

[0082] The element of input image data representative of the image may be representative of at least one among: a 2D face photo in a front view and / or a profile view and / or a top view and / or any view in between a front, profile, and top view; a 3D face photo.

[0083] The model may comprise a facial-feature estimation component. The facial-feature estimation component may comprise the at least one fully connected facial-feature estimation layer. The model may comprise a facial-feature clustering component. The facial-feature estimation component may comprise the facial-feature clustering component. The method may comprise the facial-feature clustering component estimating a set of facial-feature similarity estimation values within the set of facial-features.

[0076]

[0084] The at least one fully connected facial-feature estimation layer may comprise a plurality of fully connected facial-feature estimation layers. The plurality of fully connected facial-feature estimation layers may be, at least in part, a parallelly organized plurality of fully connected facial-feature estimation layers. The method may comprise each fully connected facial-feature estimation layer in the parallelly organized plurality of fully connected facial-feature estimation layers estimating a set of classification confidences for a distinct facial feature based on the feature vector; wherein estimating the set of facial feature(s) comprises selecting a facial feature based on the estimated classification confidence.

[0077]

[0085] In other words, in the parallelly organized plurality of fully connected facial-feature estimations layers, each fully connected facial-feature estimations layer may parallel to each other. Put differently, the method may comprise each fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers estimating a set of classification confidences for a distinct facial feature based on the feature vector. In still other words, the distinct feature of any fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facialfeature estimations layers may be different from the distinct feature of any other fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers. This may be optionally advantageous, since it may improve at least the efficiency and / or the accuracy and / or speed of the method and system running the method in that the complexity of each fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers is reduced. This may be at least because each fully connected facialfeature estimations layer in the parallelly organized plurality of fully connected facialfeature estimations layers may "specialize" on a distinct feature.

[0078]

[0086] The plurality of fully connected facial-feature estimation layers may be, at least in part, a cascaded-organized plurality of fully connected facial-feature estimation layers. The method may comprise each fully connected facial-feature estimation layer in the cascaded- organized plurality of fully connected facial-feature estimation layers estimating a set of classification confidences for a list of facial features of a term of classification based on the feature vector. The term of classification of the facial features for a fully connected facial- feature estimation layer in the cascaded-organized plurality of fully connected facialfeature estimation layers may be parent to the term of classification of the facial features for the following fully connected facial-feature estimation layer in the cascaded-organized plurality of fully connected facial-feature estimation layers. Estimating the set of facial feature(s) comprises selecting the facial features from the list of facial features based on the estimated classification confidence.

[0079]

[0087] In other words, in the cascaded-organized plurality of fully connected facial-feature estimations layers, each fully connected facial-feature estimations layer may be cascaded to another. That is, the fully connected facial-feature estimation layers may be consecutive to each other, thus forming a chain. Put differently, the method may comprise each node in the chain estimating a set of classification confidences for a list of facial features of a term of classification based on the feature vector. The term of classification for a node may be parent to the term of classification for the following node. This may be optionally advantageous, since it may improve at least the efficiency and / or the accuracy and / or speed of the method and system running the method in that the complexity of each fully connected facial-feature estimations layer in the cascaded-organized plurality of fully connected facial-feature estimations layers is reduced. This may be at least because each fully connected facial-feature estimations layer in the cascaded-organized plurality of fully connected facial-feature estimations layers may "specialize" on a term of classification.

[0080]

[0088] Also disclosed are a system, a computer program product and uses. Advantages and details discussed in the context of the method may respectively apply also in the context of the system, the computer program product and the uses.

[0081]

[0089] In a second embodiment, a system is disclosed. The system comprises a data- processing system. The system is configured for carrying out the method according to any of the preceding embodiments.

[0082]

[0090] The system may be configured for receiving the input image data and performing the estimation step according to any of the method embodiments.

[0083]

[0091] In a third embodiment, a computer program product is disclosed. The computer program product comprises instructions which, when the program is executed by a data- processing system, cause the data-processing system to carry out the method according to any of the method embodiments.

[0084]

[0092] In a fourth embodiment, a use of the method or the system to diagnose a genetic disease, particularly an ultra-rare disease, is disclosed.

[0093] In a fifth embodiment, a use of the method or the system to verify a diagnosis of a genetic disease, particularly an ultra-rare disease, is disclosed.

[0085]

[0094] In a sixth embodiment, a use of the method or the system for design of a personalized therapy of a genetic disease, particularly an ultra-rare disease, is disclosed.

[0086]

[0095] The following embodiments also form part of the invention.

[0087] Method embodiments

[0088]

[0096] Below, embodiments of a method will be discussed. The method embodiments are abbreviated by the letter "M" followed by a number. Whenever reference is herein made to the "method embodiments", these embodiments are meant.

[0089] Ml. A method, comprising

[0090] (a) receiving an element of input image data representative of an image,

[0091] (b) an estimating step comprising estimating for the element of input image data by means of a machine learning model a. a set of facial feature(s), particularly facial feature(s) indicative of phenotypic abnormalities, and b. at least one of (i) a set of disease classification confidences and (ii) a set of disease similarity estimation values, each of the disease classification confidences and / or each of the disease similarity estimation values corresponding to a respective disease from a list of diseases,

[0092] (c) outputting a result of the estimating step, such as an estimated presence of a disease from the set of diseases, wherein the model further comprises a feature vector part computing a feature vector based on the input image data, such as convolutional layers of a convolutional neural network, at least one or a plurality of fully connected layer(s) and / or at least one or a plurality of pooling layer(s) of such a convolutional neural network.

[0093] M2. The method according to the preceding embodiment, wherein the machine learning model comprises an image transformer and / or a convolutional neural network, particularly a residual neural network, such as an improved residual network.

[0094] M3. The method according to any of the preceding embodiments, wherein the model comprises at least one fully connected facial-feature estimation layer; wherein the method comprises the at least one fully connected facial-feature estimation layer estimating a set of classification confidences for a list of facial features based on the feature vector; and wherein estimating the set of facial feature(s) comprises selecting the facial features from the list of facial features based on the estimated classification confidence.

[0095] M4. The method according to any of the preceding embodiments, wherein the model comprises a disease estimation component; wherein the method comprises the disease estimation component estimating at least one of the set of disease classification confidences and the set of disease similarity estimation values.

[0096] M5. The method according to the preceding embodiment, wherein the disease estimation component comprises at least one fully connected disease estimation layer; and wherein the method comprises the at least one fully connected disease estimation layer estimating the set of disease classification confidences.

[0097] M6. The method according to any of the two preceding embodiments, wherein the disease estimation component comprises a disease clustering component; and wherein the method comprises the disease clustering component estimating the set of disease similarity estimation values.

[0098] M7. The method according to any of the preceding embodiments, wherein the method comprises a pre-training step, the pre-training step comprising pre-training the machine learning model with a face recognition data set comprising a plurality of training image data elements representative of a face photo.

[0099] M8. The method according to the preceding embodiment and with the features of M2, wherein the machine learning model comprises the convolutional neural network, and wherein the pre-training step comprises using additive angular margin loss (ArcFace) as loss function.

[0100] M9. The method according to the preceding embodiment, wherein batch normalization is not used when evaluating the loss function.

[0101] MIO. The method according to any of the preceding embodiments with the features of M7, wherein the face recognition data set comprises at least 100.000 images corresponding to at least 10.000 individuals.

[0102] Mil. The method according to any of the preceding embodiments, wherein the method comprises an aligning step, the aligning step comprising identifying a plurality of facial landmarks in the input image data and transforming the input image data to align the facial landmarks with pre-defined positions before performing the estimating step. M12. The method according to the preceding embodiment, wherein the aligning step comprises using at least five facial landmarks, such as left and right eyes, nose and left and right mouth corners.

[0103] M13. The method according to any of the two preceding embodiments, wherein the aligning step comprises using an affine transformation for transforming the input image data.

[0104] M14. The method according to any of the preceding embodiments with the features of M7, wherein the pre-training step comprises further performing the aligning step for each training image data element of the face recognition data set.

[0105] M15. The method according to any of the preceding embodiments, wherein the method comprises a fine-tuning step.

[0106] M16. The method according to any of the preceding embodiments and with the features of M2, M3 and M15, wherein the fine-tuning step further comprises obtaining the at least one fully connected facial-feature estimation layer based on training data for fine-tuning.

[0107] M17. The method according to any of the preceding embodiments and with the features of M2, M5 and M15, wherein the fine-tuning step further comprises obtaining the at least one fully connected disease estimation layer based on training data for fine-tuning.

[0108] M18. The method according to any of the preceding embodiments with the features of M15, wherein the fine-tuning step comprises at least obtaining the feature vector part of the machine learning model, such as the fully connected layer(s) of the feature vector part of the model.

[0109] M19. The method according to any of the preceding embodiments with the features of M15, wherein the fine-tuning step comprises using at least one of cross entropy, binary cross entropy and weighting.

[0110] M20. The method according to any of the preceding embodiments with the features of M16, wherein the fine-tuning step comprises obtaining the at least one fully connected facialfeature estimation layer by using binary cross entropy.

[0111] M21. The method according to any of the preceding embodiments with the features of M17, wherein the fine-tuning step comprises obtaining the at least one fully connected disease estimation layer by using cross entropy.

[0112] M22. The method according to any of the preceding embodiments with the features of M17, wherein in the fine-tuning step, obtaining the at least one fully connected disease estimation layer comprises positive weighting, particularly positive weighting and class weighting. M23. The method according to any of the preceding embodiments with the features of M16 and M17, wherein the fine-tuning step comprises training the machine-learning model to estimate the set of facial feature(s) and the a set of disease classification confidences together.

[0113] M24. The method according to any of the preceding embodiments but M23, and with the features of M16 and M17, wherein the fine-tuning step comprises first training the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer, while not modifying weights of the at least one fully connected facial-feature estimation layer, and then training the at least one fully connected facial-feature estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer, or first training the fully connected layer(s) of the feature vector part and the at least one fully connected facial-feature estimation layer, while not modifying weights of the at least one fully connected disease estimation layer, and then training the at least one fully connected disease estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one facialfeature estimation layer.

[0114] M25. The method according to any of the preceding embodiments with the features of M15, wherein the fine-tuning step comprises using a Reduce-Learning-Rate-on-Plateau-Learning Rate Scheduler.

[0115] M26. The method according to any of the preceding embodiments, wherein the method comprises determining, based on a result of the estimating step, particularly an estimated disease, and the set of facial feature(s) estimated in the estimating step, a similarity between the estimated set of facial feature(s) and a set of facial features typically associated with the estimated disease.

[0116] M27. The method according to the preceding embodiment, wherein determining the similarity comprises retrieving the set of facial features typically associated with the estimated disease from a database comprising, for each disease of the list of diseases, at least one facial feature associated with the respective disease.

[0117] M28. The method according to any of the preceding embodiments, particularly with the features of any of the two preceding embodiments, wherein the method comprises generating at least one class activation map of the machine learning model. M29. The method according to the preceding embodiment, wherein the at least one class activation map of the machine learning model is at least one gradient-weighted class activation map.

[0118] M30. The method according to any of the preceding embodiments with the features of M3 and M28, wherein the at least one class activation map relates to the at least one fully connected facial-feature estimation layer.

[0119] M31. The method according to any of the preceding embodiments with the features of M5 and M28, wherein the at least one class activation map relates to the at least one fully connected disease estimation layer.

[0120] M32. The method according to any of the preceding embodiments with the features of M3, M5 and M28, wherein the at least one class activation map relates to the at least one fully connected facial-feature estimation layer and the at least one fully connected disease estimation layer.

[0121] M33. The method according to any of the preceding embodiments, wherein the element of input image data representative of the image is representative of a frontal face photo.

[0122] M34. The method according to any of the preceding embodiments, wherein the at least one or the plurality of facial feature(s) are at least one or a plurality of HPO-terms.

[0123] M35. The method according to any of the preceding embodiments, wherein the disease whose presence is estimated by the machine learning model is selected from a list of diseases comprising human genetic diseases, such as rare human genetic diseases.

[0124] M36. The method according to any of the preceding embodiments, wherein the method is a computer-implemented method.

[0125] M37. The method according to any of the preceding embodiments, wherein the element of input image data representative of the image is representative of at least one among: a 2D face photo in a front view and / or a profile view and / or a top view and / or any view in between a front, profile, and top view; a 3D face photo.

[0126] M38. The method according to any of the preceding embodiments, wherein the model comprises a facial-feature estimation component.

[0127] M39. The method according to the preceding embodiment, with the features of embodiment M3 and M38, wherein the facial-feature estimation component comprises the at least one fully connected facial-feature estimation layer.

[0128] M40. The method according to any of the preceding embodiments, wherein the model comprises a facial-feature clustering component. M41. The method according to the preceding embodiment, with the features of embodiment M38 and M40, wherein the facial-feature estimation component comprises the facial-feature clustering component.

[0129] M42. The method according to the preceding embodiment, with the features of embodiment M40, wherein the method comprises the facial-feature clustering component estimating a set of facial-feature similarity estimation values within the set of facialfeatures.

[0130] M43. The method according to any of the preceding embodiments, with the features of embodiment M3, wherein the at least one fully connected facial-feature estimation layer comprises a plurality of fully connected facial-feature estimation layers.

[0131] M44. The method according to the preceding embodiment, wherein the plurality of fully connected facial-feature estimation layers is, at least in part, a parallelly organized plurality of fully connected facial-feature estimation layers.

[0132] M45. The method according to the preceding embodiment, wherein the method comprises each fully connected facial-feature estimation layer in the parallelly organized plurality of fully connected facial-feature estimation layers estimating a set of classification confidences for a distinct facial feature based on the feature vector; wherein estimating the set of facial feature(s) comprises selecting a facial feature based on the estimated classification confidence.

[0133] In other words, in the parallelly organized plurality of fully connected facial-feature estimations layers, each fully connected facial-feature estimations layer may be parallel to each other. Put differently, the method may comprise each fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers estimating a set of classification confidences for a distinct facial feature based on the feature vector. In still other words, the distinct feature of any fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facialfeature estimations layers may be different from the distinct feature of any other fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers. This may be optionally advantageous, since it may improve at least the efficiency and / or the accuracy and / or speed of the method and system running the method in that the complexity of each fully connected facial-feature estimations layer in the parallelly organized plurality of fully connected facial-feature estimations layers is reduced. This may be at least because each fully connected facialfeature estimations layer in the parallelly organized plurality of fully connected facialfeature estimations layers may "specialize" on a distinct feature. M46. The method according to any of the preceding embodiments, with the features of embodiment M43, wherein the plurality of fully connected facial-feature estimation layers is, at least in part, a cascaded-organized plurality of fully connected facial-feature estimation layers.

[0134] M47. The method according to the preceding embodiment, wherein the method comprises each fully connected facial-feature estimation layer in the cascaded-organized plurality of fully connected facial-feature estimation layers estimating a set of classification confidences for a list of facial features of a term of classification based on the feature vector.

[0135] M48. The method according to the preceding embodiment, wherein the term of classification of the facial features for a fully connected facial-feature estimation layer in the cascaded-organized plurality of fully connected facial-feature estimation layers is parent to the term of classification of the facial features for the following fully connected facial-feature estimation layer in the cascaded-organized plurality of fully connected facialfeature estimation layers.

[0136] M49. The method according to any of the preceding two embodiments, wherein estimating the set of facial feature(s) comprises selecting the facial features from the list of facial features based on the estimated classification confidence.

[0137] In other words, in the cascaded-organized plurality of fully connected facial-feature estimations layers, each fully connected facial-feature estimations layer is cascaded to another. That is, the fully connected facial-feature estimation layers are consecutive to each other, thus forming a chain. Put differently, the method may comprise each node in the chain estimating a set of classification confidences for a list of facial features of a term of classification based on the feature vector. The term of classification for a node may be parent to the term of classification for the following node. This may be optionally advantageous, since it may improve at least the efficiency and / or the accuracy and / or speed of the method and system running the method in that the complexity of each fully connected facial-feature estimations layer in the cascaded-organized plurality of fully connected facial-feature estimations layers is reduced. This may be at least because each fully connected facial-feature estimations layer in the cascaded-organized plurality of fully connected facial-feature estimations layers may "specialize" on a term of classification.

[0138] System embodiments

[0139]

[0097] Below, embodiments of a system will be discussed. The system embodiments are abbreviated by the letter "S" followed by a number. Whenever reference is herein made to the "system embodiments", these embodiments are meant. 51. A system comprising a data-processing system, wherein the system is configured for carrying out the method according to any of the preceding embodiments.

[0140] 52. The system according to the preceding embodiment, wherein the system is configured for receiving input image data and performing the estimation step according to any of the method embodiments.

[0141] Computer program product embodiments

[0142]

[0098] Below, embodiments of a computer program product will be discussed. These embodiments are abbreviated by the letter "C" followed by a number. Whenever reference is herein made to the "computer program product embodiments", these embodiments are meant.

[0143] Cl. A computer program product comprising instructions which, when the program is executed by a data-processing system, cause the data-processing system to carry out the method according to any of the method embodiments.

[0144] Use embodiments

[0145]

[0099] Below, embodiments of a use will be discussed. These embodiments are abbreviated by the letter "U" followed by a number. Whenever reference is herein made to the "use embodiments", these embodiments are meant.

[0146] Ul. Use of the method according to any of the method embodiments or the system according to any of the system embodiments to diagnose a genetic disease, particularly an ultra-rare disease.

[0147] U2. Use of the method according to any of the method embodiments or the system according to any of the system embodiments to verify a diagnosis of a genetic disease, particularly an ultra-rare disease.

[0148] U3. Use of the method according to any of the method embodiments or the system according to any of the system embodiments for design of a personalized therapy of a genetic disease, particularly an ultra-rare disease.

[0149]

[0100] Exemplary features of the invention are further detailed in the figures and the below description of the figures.

[0150] Brief description of the figures

[0151] Fig. 1 shows an embodiment of a method, comprising training a machine learning model

[0152] Fig. 2 shows another embodiment of the method, comprising training the machine learning model and generating an estimation Fig. 3 shows still another embodiment of the method, comprising generating the estimation and generating an activation map

[0153] Figs. 4a, 4b and 4c show still other exemplary embodiments of the method, particularly of training the machine learning model.

[0154] Figs. 5 and 6 show two embodiments of generating the estimation

[0155] Figs. 7a-7f show examples of the activation maps for different facial features

[0156] Figs. 8a, 8b show example activation maps for different models estimating a same disease

[0157] Fig. 9 shows an embodiment of part of the method, comprising the use of a plurality of parallelly organized fully-connected facial-feature estimations layer.

[0158] Fig. 10 shows an embodiment of part of the method, comprising the use of a plurality of cascaded-organized fully-connected facial-feature estimations layer.

[0159] Fig. 11 shows an embodiment of part of the method, comprising a combination of the embodiment of Fig. 9 and of Fig. 10.

[0160] Detailed figure description

[0161]

[0101] For the sake of clarity, some features may only be shown in some figures, and others may be omitted. However, also the omitted features may be present, and the shown and discussed features do not need to be present in all embodiments.

[0162]

[0102] Fig. 1 shows an embodiment of the method. The method comprises training a machine learning model 10 comprising a neural network, such as an improved residual network. The method comprises pre-training the machine learning model 10 by means of a face recognition data set 30, which in the example of Fig. 1 comprises a multitude of images relating to human faces comprising a great diversity. Thus, the network is pretrained for face verification and / or face recognition. As loss function in the pre-training step, in the example of Fig. 1, additive angular margin loss (ArcFace) is used. The face recognition data set may for example be GLINT360K. The pre-training step is usually unsupervised.

[0163]

[0103] As a result, the pre-trained model is suitable to generate a feature vector for an input image of a face. Usually, the feature vector comprises 1 by n fields, such as 1 by 512 fields. Each field usually comprises a scalar value. As the person skilled in the art will easily understand, the feature vector can be interpreted as specifying dimensions of the feature spaces, and the scalar values may then refer the position of a point in said feature space.

[0164]

[0104] In an alternative embodiment, the machine learning model comprises an image transformer which is pre-trained to generate a corresponding feature vector.

[0105] The part of the model configured to generate the feature vector may be referred to as feature vector part in the present disclosure.

[0165]

[0106] Further, the method comprises a fine-tuning step. In the fine-tuning step, the machine learning model 10 is trained to classify the input image data based on presence of facial feature(s) from a list of facial features as well as disease(s) from a list of diseases. In the example of Fig. 1, the machine learning model 10 is trained to estimate for the input image data a plurality of classification confidence levels, each corresponding to one class of facial features or one class of the diseases.

[0166]

[0107] In one exemplary embodiment, in production, that is, in use after training, the method comprises estimating a disease for an element of input image data based on a disease for which a highest classification confidence level was determined. Further, in production, the facial feature(s) may be estimated by selecting the facial features from the list of facial features whose classification confidence level is above a threshold, i.e., a cutoff.

[0167]

[0108] In the present example, the diseases in the set are ultra-rare genetic diseases. These diseases generally show an extremely low prevalence. Pre-training the machinelearning model with general face data and fine-tuning it with respect to ultra-rare genetic diseases that can be detected based on facial features may thus optionally advantageously help overcome a challenge because of too few training data.

[0168]

[0109] Further, in the present example, the facial feature(s) are abnormal facial features. As discussed above, an example for a list of such abnormal facial features is given, e.g., by "The human phenotype ontology" (HPO) (cf. citation above). Thus, for a "normal" face, no facial feature may be present according to HPO.

[0169] [HO] Returning to Fig 1, the fine-tuning step is performed using training data for fine tuning 40. The training data for fine-tuning 40 comprise training data with labels of facial feature(s) 42 and training data with labels of disease(s) 44. As, in Fig. 1, the diseases to which the training data with labels of disease(s) 44 relate are diseases that can be recognized based on facial features, such as the Cornelia de Lange syndrome or the Williams-Beuren syndrome, a common feature space for the training data with labels of facial feature(s) 42 and the training data with labels of disease(s) 44 may be used. The obtained machine learning model 10 may thus also be referred to as hybrid model. Using a hybrid model trained to estimate both, facial features and a disease, may optionally advantageously simplify result verification for a practitioner.

[0170] [Hl] The training data for fine-tuning 40 may for example based on the GestaltMatcher Database collection, cf. Hsieh et al (2022).

[0112] Additionally, or alternatively, that data used in any step of the training and / or fine- tuning may be based on a database of labelled data, wherein the labelling may be determined by at least a user. For example, the at least one user may see an image, such as patient's photo, and concurrently have access to at least a provided list of facial feature(s), such as HPO terms. The at least one user may be asked to annotate at least some of the listed facial feature(s) as either present or absent in the image, to the best of their abilities. Further, users may also annotate additional facial feature(s) that are not present in the provided list.

[0171]

[0113] In the example of Fig. 2, the model is shown in production, i.e., while processing an element of input image data 12 representative of a face of a person. The machine learning model 10 estimates for the element of input image data 12 an estimation result 20. In particular, the machine learning model 10 estimates a set of facial feature(s) 22, which may for example be HPO-terms in line with the ontology discussed in Robinson (2010). In case that there is no abnormality, the model may also determine that no HPO term is appropriate for the element of input image data 12. In this case, the set of facial features(s) may be an empty set {}.

[0172]

[0114] Further, the machine learning model 10 estimates for the element of input image data 12 a presence of a disease 24 from a list of diseases. In the example of Fig. 2, the diseases are ultra-rare genetic diseases. In the example of Fig. 2, the machine learning model 10 first estimates a set of disease classification confidences for diseases from a list of diseases (that the model has been trained with), and then selects a disease with a highest classification confidence. However, in case that the classification confidences are, e.g., all below a threshold, the machine learning model 10 may also yield another result, such as an indication that no disease was detected with sufficient confidence.

[0173]

[0115] In the example of Fig. 2, further, after training, a validation step is performed using an evaluation data set 50. The disease estimation performance may be evaluated using a mean accuracy over the diseases present in the evaluation data set 50, as described in Hustinx et al. (2023), which can be calculated as

[0174] (eq. 5) where mAk is the top- mean accuracy, C is the number of classes, c is the class index, and Ak,c is the top-k accuracy for class c. Optionally advantageously, the mean accuracy may better reflect the accuracy when considering an imbalanced evaluation set. In this context, the disease estimation may also be considered a classification problem.

[0116] For evaluating the estimation of the set of facial feature(s), mean accuracy is determined over the facial feature(s) present in the evaluation data set 50. Further, sensitivity and specificity are determined for the facial feature(s) present in the evaluation data set 50, that is, for the facial feature(s) present in the sets of facial feature(s) relating to the image data of the evaluation data set 50. Determining mean accuracy, sensitivity and specificity may optionally advantageously provide for more meaningful validation results, as the estimation of the facial features may be considered a multi-class multi-label task, where assigning no facial feature (at all) may result in a high mean accuracy (but is not the desired model behavior).

[0175]

[0117] Sensitivity and specificity may be determined using the below equations 6 and 7: number of true positives sensitivity = - - - - - - - > , - number of true positives + number of false negatives

[0176] (eq. 6) number of true negatives specificity = - - - - - - - > , - number of true negatives + number of false positives

[0177] (eq. 7)

[0178]

[0118] Fig. 3 shows still another embodiment of the method. The machine learning model 10 generates the estimation result 20 based on the element of input image data 12. In other words, the machine learning model estimates the facial feature(s) 22 and the estimated disease(s) associated with the element of input image data 12.

[0179]

[0119] Additionally, the method comprises generating an activation map 52 for the machine learning model 10, particularly for the neural network that the machine learning model 10 comprises in the example of Fig. 3, and the element of input image data 12. In the example of Fig. 3, the activation map 52 is generated by gradient-weighted class activation mapping (GradCAM). In other words, the method may comprise performing gradient-weighted class activation mapping. Optionally advantageously, the activation map may indicate how much different zones of the element of input image data contribute to the final estimation / prediction, hence, optionally advantageously allowing for easier verification of the estimation result 20 and better explainability of the estimation result 20, e.g., by a user such as a medical practitioner.

[0180]

[0120] In the example of the Figures, when the terms "Disorder" or "#Dis." are used, diseases in the sense of the present disclosure are meant. "#Dis" indicates an encoded representation of the disease, for example a single number or a vector with a plurality of scalar numbers, each scalar number corresponding to a disease or its classification confidence level.

[0121] In the context of this specification, "disorder" and the like may be construed as "disease" and the like.

[0181]

[0122] When the Figures comprise the term "HPO" or"#HPO", facial feature(s) in the sense of the present disclosure are meant. "#HPO" indicates an encoded representation of the facial feature(s), for example a list of single numbers or a vector with a plurality of scalar numbers, each scalar number corresponding to a facial feature or its classification confidence level.

[0182]

[0123] Figs. 4a, 4b and 4c show different embodiments of the method for the fine-tuning step. In these examples, the machine learning model 10 comprises a part computing a feature vector 60, that is, the feature vector part discussed above.

[0183]

[0124] In the example of a convolutional neural network, the feature vector part may for example comprise one or more pooling layers, one or more convolutional layers, and / or one or more fully connected layers. The feature vector part may, in the case of a convolutional neural network, reduce a dimensionality of a result of a preceding part of the machine learning model 10 to obtain the feature vector.

[0184]

[0125] For legibility, reference numerals have been added only to the left part of Fig. 4c, however, like features may also be present in the remainder of Fig. 4c. and Figs. 4a-4c.

[0185]

[0126] In the example of Figs. 4a-4c, the machine learning model 10 further comprises at least one fully connected facial feature estimation layer 62 and at least one fully connected disease estimation layer 64. These fully connected layers may each generate a set of classification confidence values for facial features from the list of facial features and diseases from a list of diseases.

[0186]

[0127] However, these classification confidence values may also be further processed. For example, in case of the set of facial feature(s), a threshold or cutoff-value may be applied, and for all fields comprising classification confidence values above the threshold, 1 may be outputted, and 0 for all fields comprising classification confidence values below said threshold. Further, a lookup table may be used to obtain human-intelligible terms corresponding to the facial feature(s) based on the set of classification confidence values or the exemplary set of ones and zeroes.

[0187]

[0128] With continued reference to the example, in case of the classification confidence values for the diseases, a disease comprising a highest classification confidence value may be estimated to be present. However, further criteria, such as a minimum confidence value, may be applied, too. Again, a lookup table may be used to obtain a human-intelligible term corresponding to the disease, e.g., converting an index of a field (like "23") to a name of the disease (like "Cornelia de Lange syndrome").

[0188]

[0129] It is to be noted that Figs. 4a-4c relate to training the machine learning model 10. A structure of the machine learning model 10 may be modified before the model 10 is used in production, i.e., before estimations of the model 10 is used for estimation purposes. (Cf. below section relating to Figs. 5 and 6)

[0189]

[0130] In the example of Fig. 4a, the fine-tuning step comprises fine-tuning the fully connected layer(s) of the feature vector part 61, the at least one fully connected facialfeature estimation layer 62 and the at least one fully connected disease estimation layer together 64. in on stage, that is, together. In other words, in this example, the model is hybrid and it is trained in one stage.

[0190]

[0131] In the example of Fig. 4b, in a first stage of the fine-tuning step, the at least one fully connected layer(s) of the feature vector part 61 and the at least one fully connected facial feature estimation layer 62 are trained are trained, while weights of the fully connected disease estimation layer are not modified, i.e., are "frozen". In a second stage of the fine-tuning step, the at least one fully connected disease estimation layer 64 is trained, while weights of the fully connected layer(s) of the feature vector part 61 and the at least one fully connected facial-feature estimation layer 62 are not modified ("frozen"). In other words, in the example of Fig 4b, the hybrid model is trained in two stages.

[0191]

[0132] In the example of Fig. 4c, in the first stage of the fine-tuning step, the fully connected layer(s) of the feature vector part 61 and the at least one fully connected disease estimation layer 64 are trained, while weights of the at least one fully connected facialfeature estimation layer 62 are not modified. In the second stage of the fine-tuning step, the at least one fully connected facial-feature estimation layer 62 is trained, while the weights of the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer 64 are not modified ("frozen").

[0192]

[0133] The person skilled in the art will easily understand that, in the fine-tuning step, other parts of the machine learning model 10, such as convolutional layers of a convolutional neural network, may be modified / re-trained, too. However, in case of (ultrafare diseases, there may be too few training data for fine-tuning 40 to obtain an improvement of these layers, hence, in some cases, it may be beneficial to not change their weights in the fine-tuning step.

[0193]

[0134] Figs. 5 and 6 show different embodiments of the machine learning model "in production", i.e., different embodiments of a step of the method in which the machine learning model is used to generate an estimation based on an element of input image data

[0194] 12.

[0195]

[0135] In the example of Fig. 5, the machine learning model 10 trained, e.g. as described in one of Figs. 4a, 4b or 4c is used.

[0196]

[0136] An element of input image data 12 is processed by the machine learning model 10, e.g., by convolutional layers of a convolutional neural network, as shown in Fig. 5, and by the fully connected layer(s) of the feature vector part 61 of the convolutional neural network. Thus, the feature vector 60 corresponding to the element of input image data 12 is generated.

[0197]

[0137] Based on the feature vector 60, the at least one fully connected facial-feature estimation layer 62 generates the classification confidence levels for the facial features from the list of facial features and thus generates the estimated set of facial features 22, as discussed above.

[0198]

[0138] The method may comprise using the feature vector 60, which is based on the input image data (e.g. of a patient), and comparing the feature vector 60 to further feature vectors, which are based on reference image data. In an example, the method may comprise calculating the distance between the feature vector 60 and further feature vectors. Additionally, or alternatively, the method may comprise averaging the feature vectors based on image data, which image data belong to a group (e.g. relating to a single patient and / or a single disorder), into an average feature vector. The method may then comprise comparing the feature vector 60 to the average feature vector.

[0199]

[0139] In the example of Fig. 5, further, the at least one fully connected disease estimation layer 64 may generate classification confidence levels for diseases from the list of diseases, and based thereon, an estimated disease 24 may be generated.

[0200]

[0140] However, the machine learning model 10 may also be used for estimating the disease 24 and / or the set of facial feature(s) 22 according to another embodiment shown in Fig. 6.

[0201]

[0141] The skilled person will understand that the method and / or models and / or components as discussed with respect to the facial-feature(s) may also apply, mutatis mutandis, with respect to the disease.

[0202]

[0142] In Fig. 6, the convolutional neural network, the fully connected layer(s) of the feature vector part of the network 61 and the corresponding feature vector 60 to generate the estimated set of facial feature(s) 22 are similar and / or the same as in the example of Fig. 5.

[0143] However, the machine learning model 10 may comprise a disease clustering component 66. The disease clustering component 66 clusters the element of input image data 12 based on the corresponding feature vector 60 against feature vectors of images from the training data for fine-tuning 66 and / or feature vectors corresponding to other images showing faces with rare or ultra-rare diseases.

[0203]

[0144] Thus, while the convolutional neural network and the at least one fully connected facial-feature estimation layer 62 and the at least one fully connected disease estimation layer 62 are trained according to one of the examples of Figs. 4a-4c, when estimating the facial feature(s) 22 and the disease 24 (or the respective classification confidence levels), the fully connected layer(s) of the feature vector part 61, the at least one fully connected facial-feature estimation layer 62 and a disease clustering component 66 are used. Hence, optionally advantageously, the obtained feature vector 60 (which also the disease clustering component 66 uses) is adapted to rare diseases that may be detected based on facial features, optionally resulting in improved clustering / estimation accuracy and / or clustering by dimensions that allow for easier verification of the clustering result by humans.

[0204]

[0145] Additionally, or alternatively, it will be clear to the skilled person that the clustering discussed with regard to the disease clustering component 66 may also be encompassed by embodiments of the present invention with regard to a facial-feature clustering component. This is not shown in the figures. Put differently, the machine learning model 10 may comprise a facial-feature clustering component. The facial-feature clustering component may cluster the element of input image data 12 based on the corresponding feature vector 60 against feature vectors of images from the training data for fine-tuning 66 and / or feature vectors corresponding to other images showing faces with given facial feature(s). The model may comprise a facial-feature estimation component which, in turn, may comprise the at least one fully connected facial-feature estimation layer 62 and / or the facial-feature clustering component.

[0205]

[0146] For illustrative purposes, Fig. 7a shows a computer-generated example of an element of input image data imitating a facial photo of a patient suffering from Cornelia- de-Lange-syndrome.

[0206]

[0147] Figs. 7b-7e show exemplary class activation maps of the machine learning model 10 for different estimated facial features 22, in particular gradient-weighted class activation maps generated based on the at least one fully connected facial-feature estimation layer 62. In other words, Figs. 7a-7e show regions of the exemplary element of input image data 12 that had an increased weight for estimation of the respective facial feature 22. As can be seen, the class activation maps shown in Figs. 7b-7e allow for improved verification of the model's estimation of the facial feature(s) 22.

[0207]

[0148] Fig. 7f shows a class activation map corresponding to the estimated facial feature(s) 22 corresponding to the element of input image data 12, thus allowing for easier verification of regions of the element of input image data 12 based on which the machine learning 10 estimated the presence of the (abnormal) facial feature(s) 22.

[0208]

[0149] Figs. 8a-8b show class activation maps relating to the estimation of the disease(s) 24 and / or the respective classification confidence levels.

[0209]

[0150] Fig. 8b show class activation maps relating to a first example of the at least one fully connected disease estimation layer 64. An example of a machine learning model 10 focusing on regions of the element of input image data 12 that reasonably correspond to regions showing facial abnormalities. In the example of Fig. 8b, the at least one fully connected disease estimation layer 64 was trained according to one of the examples of Figs. 4a-4c.

[0210]

[0151] Fig. 8a shows a class activation map relating to second example of the at least one fully connected disease estimation layer 64. In this case, it is apparent that the regions of the element input image data 12 that were weighted higher by the model 10 do not correspond to regions that reasonably indicate a present of a disease that (also) results in abnormal facial features, as several zones with high weighting / class activation are located outside of the face, and the zones with high class activation within the image do not correspond to zones where abnormal facial feature(s) would be expected with the corresponding disease.

[0211]

[0152] The method may be carried out by a data-processing system. For example, the data-processing system may be used for the fine-tuning step, the estimating step and the pre-training step.

[0212]

[0153] The data processing system may comprise one or more processing units configured to carry out computer instructions of a program (i.e. machine readable and executable instructions). The processing unit(s) may be singular or plural. For example, the data processing system may comprise at least one of CPU, GPU, DSP, APU, ASIC, ASIP or FPGA. The data processing system may comprise memory components, such as, main memory (e.g. RAM), cache memory (e.g. SRAM) and / or secondary memory (e.g. HDD, SDD). The data processing system may comprise volatile and / or non-volatile memory such an SDRAM, DRAM, SRAM, Flash Memory, MRAM, F-RAM, or P-RAM. The data processing system may comprise internal communication interfaces (e.g. busses) configured to facilitate electronic data exchange between components of the data processing system, such as, the communication between the memory components and the processing components. The data processing system may comprise external communication interfaces configured to facilitate electronic data exchange between the data processing system and devices or networks external to the data processing system. For example, the data processing system may comprise network interface card(s) that may be configured to connect the data processing system to a network, such as, to the Internet. The data processing system may be configured to transfer electronic data using a standardized communication protocol. The data processing system may be a centralized or distributed computing system.

[0213]

[0154] The data processing system may comprise user interfaces, such as: output user interface, such as: o screens or monitors configured to display visual data (e.g. displaying graphical user interfaces of the questionnaire to the user), o speakers configured to communicate audio data (e.g. playing audio data to the user), input user interface, such as: o a camera configured to capture visual data (e.g. capturing images and / or videos of the user), o a microphone configured to capture audio data (e.g. recording audio from the user), o a keyboard configured to allow the insertion of text and / or other keyboard commands (e.g. allowing the user to enter text data and / or other keyboard commands by having the user type on the keyboard) and / or o a trackpad, mouse, touchscreen, joystick - configured to facilitate the navigation through different graphical user interfaces.

[0214]

[0155] To put it simply, the data processing system may be a processing unit configured to carry out instructions of a program. The data processing system may be a system-on- chip comprising processing units, memory components and busses. The data processing system may be a personal computer, a laptop, a pocket computer, a smartphone, a tablet computer. The data processing system may be a server, a server system, a portion of a cloud computing system or a system emulating a server, such as a server system with an appropriate software for running a virtual machine. The data processing system may be a processing unit or a system-on-chip that may be interfaced with a personal computer, a laptop, a pocket computer, a smartphone, a tablet computer and / or user interfaces (such as the upper-mentioned user interfaces).

[0215]

[0156] Fig. 9 shows an embodiment of part of the method, comprising the use of a plurality of parallelly organized fully-connected facial-feature estimation layers 62, 62'.

[0216]

[0157] The at least one fully-connected facial-feature estimations layer, put differently, may comprise, at least in part, at least fully-connected facial-feature estimations layers 62, 62'. In other words, in the parallelly organized plurality of fully connected facial-feature estimations layers, each fully connected facial-feature estimations layer 62, 62' is parallel to each other. The method may comprise layer 62 estimating a set of classification confidences for a first distinct facial feature based on the feature vector 60. Based on said classification confidence, estimation result 22 may be determined. Analogously, the method may comprise layer 62' estimating a set of classification confidences for a second distinct facial feature based on the feature vector 60. Based on said classification confidence, estimation result 22' may be determined. In other words, the first distinct facial feature may be different from the second distinct facial feature. That is, estimation result 22 and estimation result 22 may relate to different facial features.

[0217]

[0158] Put differently, a method according to embodiments may comprise utilizing of a number of fully connected facial-feature estimations layers, said layers being non- consecutive, where the number of fully connected facial-feature estimations layers may coincide with a number of distinct and / or specific facial features to be predicted and / or estimated, e.g. of HPO terms to be predicted and / or estimated. Each fully connected facialfeature estimations layer in the number of fully connected facial-feature estimations layers may in other words predict and / or estimate a distinct, specific facial feature.

[0218]

[0159] Fig. 10 shows an embodiment of part of the method, comprising the use of a plurality of cascaded-organized fully-connected facial-feature estimation layers 62, 62'.

[0219]

[0160] The at least one fully-connected facial-feature estimations layer, put differently, may comprise, at least in part, at least fully-connected facial-feature estimations layers 62, 62'. In other words, in the cascaded-organized plurality of fully connected facialfeature estimations layers, each fully connected facial-feature estimations layer 62, 62' is consecutive to each other. The method may comprise layer 62 estimating a set of classification confidences for a list of facial features of a first term of classification. Based on said classification confidences, estimation result 22 may be determined. Analogously, the method may comprise layer 62' estimating a set of classification confidences for a list of facial features of a second term of classification. Based on said classification confidences, estimation result 22' may be determined. Layer 62' may be downstream of layer 62 and the second term of classification may be child to the first term of classification, so that estimation result 22' may relate to a facial feature that is child, in terms of classification, to the facial feature that estimation result 22 relates to.

[0220]

[0161] Put differently, embodiments of the present invention may relate to a cascade of, at least some, fully-connected facial-feature estimations layers, for example as a form of hierarchical classification. That is, in an example, a fully connected layer 62 may have L outputs which relate to "Abnormal facial shape", a fully-connected linear layer 62' may have M outputs which relate to "Abnormality of the nose", where "Abnormality of the nose" is a child term of "Abnormal facial shape". The skilled person will understand that the cascade may comprise more than two fully-connected facial-feature estimations layers 62, 62'.

[0221]

[0162] Fig. 11 shows an embodiment of part of the method, comprising a combination of the embodiments of Fig. 9 and of Fig. 10.

[0222]

[0163] The at least one fully-connected facial-feature estimation layer may comprise at least two pluralities of cascaded-organized fully-connected facial-feature estimation layers. The first plurality may comprise at least fully-connected facial-feature estimation layers 62, 62' and the second plurality may comprise at least fully-connected facial-feature estimation layers 62", 62'". The at least two pluralities of cascaded-organized fully- connected facial-feature estimation layers may be parallelly organized with respect to each other. Put differently, the first plurality may be organized parallelly with respect to the second plurality.

[0223]

[0164] Each of the pluralities may estimate a set of classification confidences for a distinct list of facial features in a hierarchical manner, as detailed above with regard to the embodiment of Fig. 10. In other words, estimation result 22' may relate to a facial feature that is child, in terms of classification, to the facial feature that estimation result 22 relates to. Estimation result 22'" may relate to a facial feature that is child, in terms of classification, to the facial feature that estimation result 22". The list of facial features estimated by the first plurality may be distinct from the list of facial features estimated by the second plurality. Generally, the number of pluralities may coincicde with the number of distinct lists of facial features that are estimated.

[0224]

[0165] It will be understood that, with regard to the fine-tuning step and / or to any training step, that said step(s) may be carried out for one or more fully connected facial-feature estimation layers in the parallelly organized or cascaded-organized plurality of layers at a time. In other words, one or more fully connected facial-feature estimation layers in the plurality can undergo the fine-tuning step and / or to any training step at a first time and one or more other fully connected facial-feature estimation layers in the plurality can undergo the fine-tuning step and / or to any training step at a second time, different from the first time.

[0225]

[0166] While in the above, a preferred embodiment has been described with reference to the accompanying drawings, the skilled person will understand that this embodiment was provided for illustrative purpose only and should by no means be construed to limit the scope of the present invention, which is defined by the claims.

[0226]

[0167] Whenever a relative term, such as "about", "substantially" or "approximately" is used in this specification, such a term should also be construed to also include the exact term. That is, e.g., "substantially straight" should be construed to also include "(exactly) straight".

[0227]

[0168] Whenever steps were recited in the above or also in the appended claims, it should be noted that the order in which the steps are recited in this text may be accidental. That is, unless otherwise specified or unless clear to the skilled person, the order in which steps are recited may be accidental. That is, when the present document states, e.g., that a method comprises steps (A) and (B), this does not necessarily mean that step (A) precedes step (B), but it is also possible that step (A) is performed (at least partly) simultaneously with step (B) or that step (B) precedes step (A). Furthermore, when a step (X) is said to precede another step (Z), this does not imply that there is no step between steps (X) and (Z). That is, step (X) preceding step (Z) encompasses the situation that step (X) is performed directly before step (Z), but also the situation that (X) is performed before one or more steps (Yl), ..., followed by step (Z). Corresponding considerations apply when terms like "after" or "before" are used.

[0228] Reference numerals

[0229] 10 Machine learning model

[0230] 12 Element of input image data

[0231] 20 Estimation result

[0232] 22 Estimated set of facial feature(s)

[0233] 24 Estimated disease(s)

[0234] 30 Face recognition data set

[0235] 40 Training data for fine-tuning

[0236] 42 Training data with labels of facial feature(s)

[0237] 44 Training data with labels of disease(s)

[0238] 50 Evaluation data set

[0239] 52 Activation map

[0240] 60 Feature vector

[0241] 61 Fully connected layer(s) of the feature vector part

[0242] 62 Fully connected facial-feature estimation layer

[0243] 64 Fully connected disease estimation layer

[0244] 66 Disease clustering component

Claims

Claims1. A computer-implemented method, comprising(a) receiving an element of input image data representative of an image,(b) an estimating step comprising estimating for the element of input image data by means of a machine learning model a. a set of facial feature(s), particularly facial feature(s) indicative of phenotypic abnormalities, and b. at least one of (i) a set of disease classification confidences and (ii) a set of disease similarity estimation values, each of the disease classification confidences and / or each of the disease similarity estimation values corresponding to a respective disease from a list of diseases,(c) outputting a result of the estimating step, such as an estimated presence of a disease from the set of diseases,(d) a pre-training step, the pre-training step comprising pre-training the machine learning model with a face recognition data set comprising a plurality of training image data elements representative of a face photo, and(e) a fine-tuning step, wherein the model further comprises a feature vector part computing a feature vector based on the input image data, at least one fully connected facial-feature estimation layer, and a disease estimation component; wherein the method comprises the at least one fully connected facial-feature estimation layer estimating a set of classification confidences for a list of facial features based on the feature vector; wherein estimating the set of facial feature(s) comprises selecting the facial features from the list of facial features based on the estimated classification confidence, wherein the method comprises the disease estimation component estimating at least one of the set of disease classification confidences and the set of disease similarity estimation values, and wherein the fine-tuning step further comprises obtaining the at least one fully connected facial-feature estimation layer and the at least one fully connected disease estimation layer based on training data for fine-tuning.

2. The method according to the preceding claim, wherein the fine-tuning step comprises using a Reduce-Learning-Rate-on-Plateau-Learning Rate Scheduler.

3. The method according to any of the preceding claims, wherein the fine-tuning step comprisesfirst training the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer, while not modifying weights of the at least one fully connected facial-feature estimation layer, and then training the at least one fully connected facial-feature estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one fully connected disease estimation layer, or first training the fully connected layer(s) of the feature vector part and the at least one fully connected facial-feature estimation layer, while not modifying weights of the at least one fully connected disease estimation layer, and then training the at least one fully connected disease estimation layer while not modifying weights of the fully connected layer(s) of the feature vector part and the at least one facial-feature estimation layer.

4. The method according to any of the preceding claims, wherein the method comprises determining, based on a result of the estimating step, particularly an estimated disease, and the set of facial feature(s) estimated in the estimating step, a similarity between the estimated set of facial feature(s) and a set of facial features typically associated with the estimated disease.

5. The method according to the preceding claim, wherein determining the similarity comprises retrieving the set of facial features typically associated with the estimated disease from a database comprising, for each disease of the list of diseases, at least one facial feature associated with the respective disease.

6. The method according to any of the preceding claims, particularly according to any of any of the two preceding claims, wherein the method comprises generating at least one class activation map of the machine learning model.

7. The method according to the preceding claim, wherein the class activation map of the machine learning model is at least one gradient-weighted class activation map.

8. The method according to any of the preceding claims, wherein the element of input image data representative of the image is representative of at least one among: a 2D face photo in a front view and / or a profile view and / or a top view and / or any view in between a front, profile, and top view; a 3D face photo.

9. The method according to any of the preceding claims, wherein the disease whose presence is estimated by the machine learning model is selected from a list of diseases comprising human genetic diseases, such as rare human genetic diseases.

10. The method according to any of the preceding claims, wherein the method comprises a pre-training step, the pre-training step comprising pre-training the machine learning model with a face recognition data set comprising a plurality of training image dataelements representative of a face photo, wherein the machine learning model comprises the convolutional neural network, and wherein the pre-training step comprises using additive angular margin loss (ArcFace) as loss function.

11. The method according to any of the preceding claims, wherein the method comprises an aligning step, the aligning step comprising identifying a plurality of facial landmarks in the input image data and transforming the input image data to align the facial landmarks with pre-defined positions before performing the estimating step.

12. The method according to any of the preceding claims, wherein the fine-tuning step comprises obtaining the at least one fully connected facial-feature estimation layer by using binary cross entropy.

13. The method according to any of the preceding claims, wherein the model comprises a facial-feature clustering component, wherein the method comprises the facial-feature clustering component estimating a set of facial-feature similarity estimation values within the set of facial-features.

14. A system comprising a data-processing system, wherein the system is configured for carrying out the method according to any of claims 1-13.

15. A computer program product comprising instructions which, when the program is executed by a data-processing system, cause the data-processing system to carry out the method according to any of claims 1-13.

16. Use of the method according to any of claims 1-13 or the system according to claim 13 to diagnose a genetic disease, particularly an ultra-rare disease.

Citation Information

Patent Citations

  • Systems, methods, and computer-readable media for patient image analysis to identify new diseases

    US10327637B2

  • Systems, methods, and computer-readable media for gene and genetic variant prioritization

    US10470659B2

  • Systems, methods, and computer-readable media for patient image analysis to identify new diseases

    US10667689B2

  • Methods for visual identification of cognitive disorders

    US20220165425A1