Method for assessing a risk of breast cancer recurrence

A computer-implemented method using machine learning to analyze breast cancer tissue images predicts recurrence by bypassing gene expression analysis, offering a faster, cheaper, and more accurate alternative to existing tests.

WO2025159640A1PCT designated stage Publication Date: 2025-07-31AGENDIA NV

Patent Information

Application Number
PCT/NL2025/050040
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current diagnostic tests for assessing breast cancer recurrence, such as MammaPrint, are time-consuming and costly, and existing AI methods do not accurately predict breast cancer recurrence based on image data.

Method used

A computer-implemented method using machine learning models to analyze images of stained breast cancer tissues, performing pre-processing, patch extraction, feature encoding, and aggregation to predict breast cancer recurrence, bypassing microarray-based gene expression analysis.

Benefits of technology

Accurately predicts breast cancer recurrence with a reliable and cost-effective method, providing a digital MammaPrint Index that outperforms traditional methods in predicting distant metastasis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NL2025050040_31072025_PF_FP_ABST
    Figure NL2025050040_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for assessing a risk of breast cancer recurrence, to a computer program having instructions which when executed by a computing device or system cause the computer device or system to perform the method as well as to a data-processing system comprising means for carrying out the method for assessing a risk of breast cancer recurrence. The method of the invention is for example performed by whole slide image (WSI) processing, wherein the risk of breast cancer recurrence is assessed from a WSI of a tumor tissue.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]P135360PC00Title: Method for assessing a risk of breast cancer recurrenceFIELD OF THE INVENTION The invention relates to a computer-implemented method for assessing a risk of breast cancer recurrence based on image data. Further, the invention relates to a computer program and a data-processing system for performing the method of assessing the risk of breast cancer recurrence based on image data. BACKGROUND OF THE INVENTION Worldwide, breast cancer is the leading type of cancer in women, accounting for 25% of all cases. In 2018, it resulted in 2 million new cases and 627,000 deaths. Outcomes for breast cancer vary depending on the cancer type, the extent of disease, and the person's age. The five-year survival rates in England and the United States are between 80 and 90%. In developing countries, five-year survival rates are lower. Breast cancer patients with the same stage of disease can have markedly different treatment responses and overall outcome. The strongest predictors for metastases fail to classify accurately breast tumors according to their clinical behavior. Chemotherapy or hormonal therapy reduces the risk of distant metastases by approximately one-third; however, 70 – 80 % of patients receiving this treatment would have survived without it. MammaPrint is a routine prognostic test used to assess the risk of breast cancer recurrence (van ’t Veer et al., 2002. Nature 415: 530–536). MammaPrint comprises a set of distinct breast cancer marker genes. The expression levels of such marker genes may be analyzed in a microarray-based assay for each tumor sample resulting in a specific MammaPrint index predicting a risk of breast cancer recurrence. Gene-expression based test assays such as microarrays are time and cost extensive. Therefore, there is a need to find new ways of accurate diagnostic tests to assess a risk of a recurrence in breast cancer patients. Artificial intelligence (AI) encompasses a variety of computer aided techniques that are used to solve highly complex tasks. More specifically, AI typically refers to computers excelling in tasks that are commonly associated with human cognition / behaviors. Especially the advancement of deep learning in the recent years has led to great progress in a variety of tasks. Some examples are image and speech recognition, natural language processing and pattern recognition. In the healthcare domain, AI is increasingly used to help physicians to develop more personalized treatment plans for their patients. With the use of AI the invention solves the problem to provide a reliable and accurate as well as fast and cost-effective prognostic test to assess a risk of breast cancer recurrence. The invention refers to an AI-based method to extract diagnostic and prognostic information from imaged slides of breast cancer tissue and to output a breast cancer recurrence score. Hence, the invention circumvents the use of microarray-based analysis of gene expression levels in breast cancer samples. SUMMARY OF THE INVENTION The present invention provides reliable and accurate assessment of the risk of breast cancer recurrence by analyzing images of stained tissues comprising breast cancer cells. As is shown in the examples, using a method according to the invention, breast cancer recurrence, for example the occurrence of distant recurrence which is also called metastasis, can be accurately predicted. In a first aspect, the invention provides a computer-implemented method for assessing a risk of breast cancer recurrence, the method comprises the steps of a) inputting an image of a stained tissue sample comprising breast cancercells from an individual into a processor; b) pre-processing the image in order to achieve a tissue segment image ofrelevant image areas by the processor; c) patch extraction, wherein N patches of the pre-processed tissue segmentimage are generated by the processor; d) encoding using at least one machine learning data processing model theextracted N patches into a set of feature vectors, wherein the input image is associated with a set of N feature vectors by the processor; and e) outputting using at least one further machine learning data processingmodel an overall label output predicting a breast cancer recurrence status by aggregating the contribution of each of the N feature vectors to the overall label output by the processor. For example, the image used in the computer-implemented method of the invention is a whole slide image (WSI). For example, pre-processing of the image according to the invention comprises extracting uninformative image content comprising background and / or adipose tissue. Preferably the pre-processing of the image comprises the steps of: a) down-sampling of the image by a factor s;b) transforming the image to greyscale;c) transforming each pixel-value of the image to optical density (OD);d) creating a mask of pixels by defining an OD threshold level,preferably the OD threshold is > 0.1; e) excluding pixels at the boundary, preferably the boundary region is defined by 0.03*min (width, height); f) excluding objects which are smaller than 0.1 – 1* area largest object in mask; g) excluding objects of a certain length, preferably wherein objects are excluded of a length of (Lmax – Lmin) / Lmax > 1.5 to 0.5, more preferably of a length of (Lmax – Lmin) / Lmax > 0.95; and / or h) excluding marking, preferably wherein the marking is defined by a variance across RGB channels < 0.001. Further, the down-sampling factor s for example matches one dimension of the size of the extracted patches, such as having a down-sampling factor s of 256 for extracted patches of 256 x 256. For example, the computer-implemented method comprises that each patch is transformed to a feature vector of a certain length. Preferably each feature vector encodes textural and / or morphological information of each patch. For example, the computer-implemented method comprises that after patch extraction, an additional quality control step is performed. Preferably said additional quality control step comprises, for every patch, a tissue type classification and / or nuclei detection. For example, the encoding of the feature vectors is performed using a convolutional neural network (CNN) comprising an input layer, one or more hidden layers and an output layer, such as a ResNet50. Preferably the encoding of the feature vectors is performed by using a Resnet50 having a length 1024. For example, the encoding of the feature vectors is performed using self- supervised learning. For example, the at least one further machine learning data processing model used for outputting of the overall label output is at least one fully-connected network (FCN) comprising a gated attention mechanism. Preferably the at least one further machine learning data processing model is an adapted version of the attention-based multiple Instance learning (A-MIL) model. For example, the feature vectors that receive high attention from the at least one fully-connected network (FCN) comprising a gated attention mechanism originate from tumor regions in the stained tissue sample. For example, outputting the overall label output according to the invention comprises the steps of: a) inputting the N feature vectors; b) propagating the input N feature vectors through a linear layer of a certain length, such as a length is 512; c) performing Rectified Linear Unit (ReLU) activation and Dropout; d) processing the resulting features by at least two parallel streams of gated attention mechanism, preferably wherein the first stream consists of a linear layer of length 384, followed by tangent hyperbolic (Tanh) activation and Dropout and the second stream consists of a linear layer of length 384, followed by Sigmoid activation and Dropout; e) multiplying the output of the at least two attention layers element- wise with each other; f) passing the multiplied result through the last layer of the gated attention, wherein the last layer is a linear layer, preferably wherein the length is 1; g) providing input for the last classification layer by matrix multiplication between the hidden features that were the output of the first linear layer and the last attention layer; h) propagating a slide-level feature to a single node, preferably wherein the slide-level feature is of length 512; i) providing the overall label output. Dropout is for example performed as described in Srivastava et al., 2014. J Machine Learning Res 15: 1929-1958. For example, the computer-implemented method further comprises a) providing a slide-level prediction index predicted by a Patch-based End-to-End Module from the pre-processed image patches; and b) combining the overall label output and the output of step a) to generate a refined overall label output by using a further machine learning data processing model, preferably wherein the further machine learning data processing model used for combining is at least one further fully-connected network (FCN), more preferably wherein the at least one further FCN is an ensemble model. For example, combining of the overall label output and the image output of a low-resolution image version comprises the steps of: a) inputting the overall label output and the output of the Patch-based End-to-End Module in an ensemble model, comprising at least three linear layers, wherein each layer except the last layer is followed by ReLU activation and Dropout; and b) generating a refined overall label output. For example, the staining of the tissue sample which is used as image input in the method of the invention is heamatoxylin and eosin (H&E). In another aspect, the invention also comprises training of the at least one machine learning data processing model used for encoding of feature vectors, outputting the overall label output, outputting the image output of a low-resolution image version and / or outputting the refined overall label output. For example, the training comprises the steps of:a) obtaining as ground truth a set of N patches having corresponding encoded Nfeature vectors;b) inputting as example data the set of N patches into a processor comprisingthe machine learning data processing model used for encoding the extracted N patches into a set of N feature vectors;c) performing the method of the invention comprising at least one machinelearning data processing model of claim 1D on the input set of extracted N patches;d) measuring the error between the generated feature vectors of the extractedpatches and the ground truth data; ande) updating the weights of the machine learning data processing model used forfeature vector encoding for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for feature vector encoding are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing. Further, training the method according to the invention for example comprises training the at least one machine learning data processing model used for encoding the extracted N patches into a set of feature vectors, wherein the training comprises the steps of:a) obtaining as ground truth data a set of N feature vectors having acorresponding label providing information about a risk of breast cancer recurrence in a tissue sample, preferably the label is a result of gene expression profiling; b) inputting as example data the set of N feature vectors s into a processor comprising the machine learning data processing model used for outputting the overall label output; c) performing the method according to step e) of claim 1 on the input set of N feature vectors; d) measuring the error between the generated overall label output and the ground truth data; and e) updating the weights of the at least one machine learning data processing model used for used for outputting the overall label output for performing the training of the machine learning data processing model, preferably wherein the weights of the at least one further machine learning data processing model used for outputting the overall label output are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing. Further, training the method according to the invention for example comprises training the at least one further machine learning data processing model used for outputting the output of the Patch-based End-to-End Module, wherein the training comprises the steps of:a) obtaining as ground truth data an output of the Patch-based End-to-EndModule of a stained tissue sample comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in the tissue sample, preferably the label is a result of gene expression profiling;b) inputting as example data the output of the Patch-based End-to-End Moduleof a stained tissue sample into a processor comprising the machine learning data processing model used for outputting the output of the Patch-based End-to-End Module;c) performing the method according to step a) of claim 12;d) measuring the error between the generated output of the Patch-based End-to-End Module and the ground truth data; ande) updating the weights of the machine learning data processing model used foroutputting output of the Patch-based End-to-End Module for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for outputting output of the Patch-based End-to-End Module are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing. Further, training the method according to the invention for example comprises training the machine learning data processing model used for outputting the refined overall label output, wherein the training comprises the steps of:a) obtaining as ground truth data an overall label output and an output of thePatch-based End-to-End Module having a corresponding label providing information about a risk of breast cancer recurrence of an image a stained tissue sample comprising breast cancer cells, preferably the label is a result of gene expression profiling;b) inputting as example data the overall label output and the output of thePatch-based End-to-End Module of stained tissue sample;c) performing the method according to step b) of claim 12 on the input overalllabel output and the input of the output of the Patch-based End-to-End Module of stained tissue sample;d) measuring the error between the generated refined overall label output andthe ground truth data; ande) updating the weights of the machine learning data processing model used foroutputting the refined label output for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for outputting the refined label output are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing. In another aspect, the invention is also directed to a computer program having instructions which when executed by a computing device or system cause the computer device or system to perform the computer-implemented method according to the invention. In a further aspect, the invention is also directed to a data-processing system comprising means for carrying out the computer-implemented method according to the invention. In another aspect, the invention also comprises a method of treating an individual who has been assessed as having a high risk of breast cancer recurrence according to the computer-implemented method, computer program product and / or data-processing system of the invention. For example, the method of treatment according to the invention comprises assessing the risk of breast cancer recurrence in an individual according to the computer-implemented method of the invention, and providing a therapy for preventing and / or treating breast cancer to the individual assessed as having a high risk of breast cancer recurrence. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1: Exemplary schematic overview of methods of the invention. A: Whole slide images (WSIs) are first pre-processed to obtain N patches and encoded feature vectors corresponding to these N patches. The feature vectors form the input of the Core AI model comprising an Attention-based Multiple instance learning (A-MIL) model. Also, a down-sampled low-resolution image of the WSI is passed through a ResNet50 classifier (“Tiny WSI model”). A-MIL and the ResNet50 classifier are both trained to output a label predicting a low / high risk of breast cancer recurrence. The output of the Core AI model (overall label output) and the output of the Tiny WSI model (image output of a low-resolution image version) together form the input of a final ensemble model. Here, the multimodal aspect of the ensemble model is leveraged into a final overall label (refined overall label output) predicting a refined breast cancer recurrence status (“Digital Mamma Print Index”). B: Whole slide images (WSIs) are first pre-processed to obtain N patches and encoded feature vectors corresponding to these N patches. The encoded feature vectors comprises Pathology self-supervised features obtained via Self-Supervised Learning (SSL). The feature vectors form the input of the Core AI model comprising an Attention-based Multiple instance learning (A-MIL) model. Also, a Patch-based End-to-End module was performed directly on the pre-processed image patches. A-MIL and the output of the Patch-based End-to-End module are both trained to output a label predicting a low / high risk of breast cancer recurrence. The output of the Core AI model and the output of the Patch-based End-to-End module together form the input of a final ensemble model. Here, the multimodal aspect of the ensemble model is leveraged into a final overall label (refined overall label output) predicting a refined breast cancer recurrence status (“Digital Mamma Print Index”). Fig. 2: Exemplary pre-processing of an WSI Tissue segmentation. Stroma and epithelial tissue, encapsulated by the contours is selected for further analysis. Adipose tissue and background areas are excluded. Fig. 3: Exemplary patch extraction of a WSI. Tissue segmentation and patch extraction of a whole-slide-image (WSI). Fig. 4: Exemplary network architecture of the Attention-based Multiple instance learning (A-MIL) model generating an overall label output Every feature vector which has been encoded to each of the extracted N patches is multiplied by its corresponding weight (attention score). The resulting low- dimensional embedding is then fed into a fully-connected network (FCN). Fig. 5: Exemplary ensemble model architecture The respective outputs of the Tiny WSI (image output of a low-resolution image version) and of the Core AI models (overall label output) are fed as input into a fully-connected network (FCN), such as an ensemble model. The FCN generates a (final) refined overall output predicting a breast cancer recurrence status (Digital MammaPrint Index). Fig. 6: Exemplary distribution of the final Digital Mamma Print Index (refined overall label output ) predicting a risk of breast cancer recurrence according to the method of the invention. The horizontal axis represents the Digital Mamma Print Index which is the (final) refined overall label output of the ensemble model. The vertical axis denotes density. The continuous line represents the smooth probability density function. Fig. 7: Schematic diagram of a neural network used for vector encoding A patch of an input image of a stained tissue samples obtained from an individual having breast cancer undergoes a series of convolutional and pooling layers with varying sizes. The output of the 2-D average pooling layer is an encoded feature vector derived from the input image. Fig. 8: Tissue patches at different magnifications. 256x256 pixel tissue patch at 20x (A), 10x (B) and 5x(C) magnification. Fig. 9: An example input and output of the tissue type classification model. Each marker represents the predicted label for a 256x256 pixel tissue patch (circle: invasive, diamond: non-invasive, triangle: stroma, cross: artifact, square: infiltrated-invasive, pentagon: tumor-stroma). Fig. 10: Nuclei detection. Extracted patches were processed by the nuclei detector to detect the nuclei and classify their type (circle: tumor cells, triangle: stroma cells, star: lymphocyte cells). Fig. 11: Kaplan-Meier survival curve. On the y-axis the probability of distant recurrences is represented. DETAILED DESCRIPTION Throughout this specification and the claims, unless the context requires otherwise, the word ’comprise’, and variations such as ’comprises’ and ’comprising’, will be understood to imply the inclusion of a stated member, integer or step or group of members, integers or steps but not the exclusion of any other member, integer or step or group of members, integers or steps. The terms ’a’ and ’an’ and ’the’ and similar reference used in the context of describing the invention (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., ’such as’, ’for example’), provided herein is intended merely to better illustrate the invention and does not pose a limitation on the scope of the invention otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the invention. All documents cited or referenced herein (’herein cited documents’), and all documents cited or referenced in herein cited documents, together with any manufacturer's instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the invention. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference. In the following, the features of the invention will be described in more detail. It should be understood that embodiments may be combined in any manner and in any number to create additional embodiments. The variously described examples and embodiments should not be construed to limit the invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed features. Furthermore, any permutations and combinations of all described features in this application should be considered disclosed by the description of the present application unless the context indicates otherwise. It will be appreciated that the invention for example includes computer implemented steps. For example, all steps of a method of the invention are computer-implemented steps. Embodiments for example comprise a computer apparatus, wherein a method is of the invention is performed in said computer apparatus. The invention for example extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the invention into practice. The program is for example in the form of source or object code or in any other form suitable for use in the implementation of a method according to the invention. For example, the carrier is any entity or device capable of carrying the program. For example, the carrier comprises a storage medium, such as a ROM, for example a semiconductor ROM or hard disk. Further, the carrier for example is a transmissible carrier such as an electrical or optical signal which is e.g. conveyed via electrical or optical cable or by radio or other means, e.g. via the internet or cloud. The invention solves the problem of providing a novel, fast and cost-efficient method for predicting a risk of breast cancer recurrence. The inventors of the invention surprisingly identified a method of assessing the risk of breast cancer recurrence by inputting an image of a stained tissue sample comprising breast cancer cells from an individual, pre-processing the image in order to achieve a tissue segment of relevant image areas, performing patch extraction, wherein N patches of the tissue segment are generated, encoding the extracted patches into a set of N feature vectors, and outputting an overall label output predicting breast cancer recurrence status by aggregating the contribution (weight / attention score) of each feature vector to the overall label output. The invention is for example implemented using a machine or tangible computer-readable medium or article which optionally stores an instruction or a set of instructions that, if executed by a machine, cause the machine to perform a method and / or operations in accordance with the invention. The invention is for example implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements include processors, microprocessors, circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, microchips, chip sets, etc.. Examples of software for example include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, mobile apps, middleware, firmware, software modules, routines, subroutines, functions, computer implemented methods, procedures, software interfaces, application program interfaces (API), methods, instruction sets, computing code, computer code, etc.. Inputting of the image The method of the invention uses images of stained tissue samples obtained from an individual having breast cancer. Said tissue sample comprises breast cancer cells, or is suspected to comprise breast cancer cells. For example, the tissue sample is stained with a staining selected from hematoxylin and eosin (H&E), eriochrome and eosin, modified Papanicolaou or combinations thereof. For example, the tissue sample is derived from an individual during or after a surgery for removing the breast cancer tissue or derived from a biopsy. For example, the image is obtained by virtual microscopy. For example, the image is obtained by whole-slide imaging. For example, the image is a whole-slide image (WSI). For example, the WSI has a rectangle or square format. For example, the WSI is in rectangle format having a size of 1000 x 750; 2000 x 1500; 3000 x 2250; 4000 x 3000; 5000 x 3750; 6000 x 4500; 7000 x 5250; 10,000 x 7500; 20,000 x 15,000; 30,000 x 22,500; 40,000 x 30,000; 40,000 x 30,000; 50,000 x 37,500; 60,000 x 45,000; 70,000 x 52,500; 80,000 x 60,000; 90,000 x 67,500; 100,000 x 75,000 or 200,000 x 150,000 pixels. For example, the WSI is in square format having a size of 1000 x 1000, 2000 x 2000; 3000 x 3000; 4000 x 4000; 5000 x 5000; 6000 x 6000; 7000 x 7000; 10,000 x 10,000; 20,000 x 20,000; 30,000 x 30,000; 40,000 x 40,000; 50,000 x 50,000; 60,000 x 60,000; 70,000 x 70,000; 80,000 x 80,000; 90,000 x 90,000; 100,000 x 100,000 or 200,000 x 200,000 pixels. For example, the WSI has a size up to 100,000 x 100,000 pixels. For example, the image, optionally a WSI, has a size that does not directly fit into the available GPU memory of the hardware system or data-processing system performing a method of the invention. The term ‘color channel of a pixel’, as used herein, refers to one component of the color information associated with a certain pixel. For example, in a grayscale image, every pixel is represented by a single channel which is representing the intensity of light. Typically in a grayscale image, the pixel values are within a range of 0 to 255, corresponding with black and white respectively. In a RGB color image, a pixel’s color may be determined by three channels, i.e. a red (R)-channel, green (G)-channel, blue (B)-channel. Each pixel then has a value between 0 and 255 for every color-channel, expressing the intensity of the corresponding color of the channel. Together, these values define the overall color of the pixel. The term ‘optical density (OD) of a pixel’, as used herein, refers to a measure of the extent to which light is absorbed or attenuated by a material, as represented by that pixel. Image pre-processing The method of the invention comprises a step of pre-processing the image. For example, the pre-processing is performed in order to achieve a tissue segment image comprising relevant tissue areas, preferably comprising mainly or only relevant tissue areas. For example, the image is pre-processed which results in a tissue segment image which comprises only tissue areas informative for predicting a breast cancer recurrence status. The term ‘uninformative tissue area’, as used herein, refers to a tissue area which is not informative for predicting a breast cancer recurrence, for example a tissue area comprising background and / or adipose tissue. For example, the tissue segment image obtained from pre-processing does not contain background and / or adipose tissue. For example, the image, optionally a WSI, is pre-processed using a tissue segmentation pipeline. Optionally, the pre- processing is performed by an algorithm. For example the pre-processing comprises the steps of a) down-sampling of the image by a factor s; b) transforming the image to greyscale; c) transforming each pixel-value of the image to optical density (OD);d) creating a mask of pixels by defining an OD threshold level; e) excluding pixels at the boundary; f) excluding objects which are smaller than 0.1 – 1* area largest object in mask; g) excluding objects of a certain length; and / or h) excluding markings. The term “down-sampling”, as used herein, refers to the process of reducing an image’s spatial resolution or dimensions by decreasing the number of pixels contained in the image. The term “down-sampling factor”, as used herein, refers to the factor by which an image’s resolution or dimensions are decreased during a down-sampling process. For example, the down-sampling factor s is from 8 – 1024. For example, the down-sampling factor s is 8; 16; 32; 64; 128; 192; 224; 256; 384; 512; 768; 1024; or 2048. For example, the down-sampling factor s is 256. As an example, if an image has size 1024 x 1024 and the down-sampling factor s is 256, the down-sampled image has a resolution of 4 x 4. Each pixel in the 4 x 4 down- sampled image corresponds to a 256 x 256 region in the original image. From the original image, patches of 256 x 256 can be extracted (vide infra), aligning with said regions in the down-sampled image. As such, in embodiments, the down- sampling factor matches a size of extracted patches. For example, the down- sampling factor s is 256 and the extracted patches are of size of 256 x 256 x 3, wherein “x 3” indicates the presence of three color channels (RGB). In embodiments, in step c, each pixel-value of the image is transformed to optical density (OD). For example, each pixel value in a grayscale image is transformed by computing the negative base-10 logarithm of each pixel, with the exception of pixels with a value of zero. For example, in step d, a mask, for example a binary mask, of pixels may be created based on two conditions: a) Pixels in the optical density-transformed image have values greater than a defined OD threshold level; and b) the variance of pixel values across the color channels of the original image is greater than a defined level. For example, the OD threshold level is in the range of > 1 - > 0.001, optionally > 1; > 0.8; > 0.7; > 0.6; > 0.5; > 0.4; > 0.3; > 0.2; > 0.1; > 0.09; > 0.08; > 0.07; > 0.06; > 0.05; > 0.04; > 0.03; > 0.02; > 0.01 or > 0.005. For example the OD threshold level is > 0.1. For example, the boundary region of which pixels are excluded is defined as being in a range of 0.001 – 0.5*min (width, height). For example, the boundary region of which pixels are excluded is defined by 0.5; 0.4; 0.3; 0.2; 0.1; 0.09; 0.08; 0.07; 0.06; 0.05; 0.04; 0.03; 0.02; 0.01; 0.009; 0.008; 0.007; 0.006; 0.005; 0.004; 0.003; 0.002 or 0.001*min (width, height). For example, the boundary region of which pixels are excluded is defined by at least 0.03*min (width, height). For example, objects are excluded which are smaller than 1 – 0.001*area largest object in mask. For example, objects are excluded which are smaller than 0.9; 0.8; 0.7; 0.6; 0.5; 0.4; 0.3; 0.2; 0.1; 0.09; 0.08; 0.07; 0.06; 0.05; 0.04; 0.03; 0.02 or 0.01*area largest object in mask. A person skilled in the art knows how an object in a mask can be defined. For example an object may be defined as all pixels with the same value, e.g. ‘1’ in a binary mask, that are directly connected such as horizontally, vertically and / or diagonally connected. To achieve this, the person skilled in the art may for example use the function ‘skimage.measure.label’, which is part of the skimage.measure module in the scikit-image library (available at scikit-image.org / docs / stable / api / skimage.measure.html#skimage.measure.label). For example, object of a certain length are excluded wherein the length is (Lmax – Lmin) / Lmax > 1.5 to 0.5, optionally wherein the length is (Lmax – Lmin) / Lmax > 1.5; > 1.4; > 1.3; > 1.2; > 1.1; > 1.0; > 0.95; > 0.9; > 0.85; > 0.8; > 0.75; > 0.7; > 0.65; > 0.6; > 0.55; > 0.5; >0.45; > 0.4; > 0.35; > 0.3; > 0.25; > 0.2; > 0.15; or > 0.1. For example, markings which are excluded are pen markings, dirt, dust, droplets or any further element which is not relevant for performing a method of the invention. For example, any marking is excluded having a variance across RGB channels < 0.1 – 0.0001, optionally < 0.1; < 0.09; < 0.08; < 0.07; < 0.06; < 0.05; < 0.04; < 0.03; < 0.02; < 0.01; < 0.009; < 0.008; < 0.007; < 0.006; < 0.005; < 0.004; < 0.003; < 0.002; < 0.001; < 0.0009; < 0.0008; < 0.0007; < 0.0006; or < 0.0005. As is used herein, “Lmax” and “Lmin” of an object refer to the lengths of the major and minor axes, respectively, of an equivalent ellipse that has the same covariance as the object. Pre-processing may further include the steps of tissue type classification and / or nuclei detection. In embodiments, tissue type classification may be performed by a convolutional neural network (CNN) such as a pre-trained CNN such as, for example, ResNet50, ResNet34, EfficientNet, DenseNet, RegNet, ResNeXt or a combination thereof. Nuclei detection may be performed using an AI model such as ‘You Only Look Once’ version 8 (YOLOv8), as is detailed herein below. Separate therefrom, or in addition, histopathological determination of slides or WSI by, e.g., a pathologist, may be used to provide an estimate of the tumor percentage scoring of a slide or WSI. Slides or WSI with an estimate of less than 50 % of invasive tumor cells. such as less than 40 %, less than 30 %, less than 20 %, or less than 10 % of invasive tumor cells, may be excluded from further analyses. Patch extraction The method of the invention comprises a step of patch extraction. The term “patches” as used herein refers to parts or pieces of the image, for example to parts or pieces of the pre-processed image. For example, the patch extraction comprises the generation of N patches of the pre-processed tissue segment image. N matches number of extracted patches and for example depends on the size of the specimen. For example, the N patches are extracted from a pre-processed tissue segment WSI. For example, patch extraction is performed by extracting patches of the entire pre-processed tissue segment image, e.g., pre-processed tissue segment WSI. For example, patch extraction is performed by extracting only a selection, i.e., specific areas, of the pre-processed tissue segment image, e.g. pre-processed tissue segment WSI. For example, the pre-processed tissue segment image used for patch extraction has full resolution. Full-resolution means that the resolution of the obtained image is not changed for further image processing. For example, the full- resolution image which is used for patch extraction is not down-sampled. For example, patch extraction comprises mapping of the indices of the positive binary mask pixels to the coordinates of the image, optionally to the coordinates of the image such as WSI, wherein at each set of xy-coordinates a patch is extracted. The extracted patches are for example of any size smaller than the input image, such as WSI. For example, the patch size is from 8 x 8 to 2048 x 2048 pixels. Further, the extracted patches for example comprise 3 color channels (RGB). For example, the extracted patches are of size 8 x 8; 16 x 16; 32 x 32; 64 x 64; 128 x 128; 196 x 196; 256 x 256; 384 x 384; 512 x 512; 768 x 768; 1024 x 1024; or 2048 x 2048 pixels. For example, patches are extracted having a size of 256 x 256 x 3, wherein “x 3” indicates the presence of three color channels (RGB). After patch extraction, an additional quality control step may be implemented. In said quality control step images with certain characteristics may be excluded from further analysis. For example, for every patch a tissue type classification and / or nuclei detection may be performed. The information obtained by the tissue type classification and / or nuclei detection may be used to exclude images from further analysis. For example, slides may be excluded by connected component analysis, if based on the analysis of only invasive tissue patches, the largest connected component consists of fewer than 100 patches and the second largest connected component has fewer than 50 patches. For example, if a slide has only 40 invasive patches and nothing else, the tumor percentage will be 100%. But such slide may be rejected because of the low number of invasive patches. Therefore, before calculating a tumor percentage, a connected component analysis may be performed and a slide may be rejected if the above conditions are not met. Furthermore, or in addition, slides with less than 50 % of invasive patches, such as less than 40 %, less than 30 %, less than 20 %, or less than 10 % of invasivepatches may be excluded, whereby tumor percentage is defined as:^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^  =^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^ . In tissue type classification, for every patch information about the tissue type is obtained. These tissue types may be: invasive, infiltrated invasive, non- invasive, lymphocyte, stroma, tumor stroma, and other tissue types such as red blood cells, necrosis and artifacts. For the tissue type classification task, a Multi- Scale Feature Combination Model may be used, for example developed usingPyTorch. To improve the model’s generalization to new, unseen data, the modelmay incorporate techniques such as dropout and batch normalization. Additionally, the architecture may include an embedding layer to reduce the dimensionality of the feature representations and a SoftMax-based classifier for tissue type prediction. Furthermore, or in addition, a patch selection criteria may be used in which the number of patches per tissue type per WSI are maximalized or capped. For example, the number of invasive patches may be maximalized to 1024, the number of infiltrated invasive patches to 512, the number of non-invasive patches to 256, the number of lymphocyte patches to 256, the number of stroma patches to 128, the number of tumor stroma patches to 256. Adjusting these numbers allows flexibility in emphasizing certain tissue types that might contribute more significantly to the prediction task. This tailored selection ensures that the sampled subset of N patches provides a balanced and informative representation of the slide while ignoring the uninformative part of the slide (e.g. fatty tissue), termed “enriched” slide. After performing nuclei detection, information on detected nuclei may be obtained for every patch. For nuclei detection, the advanced AI model called ‘You Only Look Once’ version 8 (YOLOv8) may be used to process each patch extracted during the patch extraction step. For example, the small version of YOLOv8 (YOLOv8s, Jocher et al., 2023. Ultralytics, available at github.com / ultralytics / ultralytics) may be used. Nuclei detection may also be performed using Mask R-CNN (He et al., 2017. arXiv:1703.06870), HoVer-Net (Graham et al. 2019. Medical Image Analysis 58: 101563) or RetinaNet (Lin et al., 2017. arXiv:1708.02002).In preferred embodiments, for every patch a tissue type classification and nuclei detection is performed. In such embodiments, the information about the tissue type and the detected nuclei obtained for the patches in an image may be combined into one measure for a certain image. Said measure is for example a digital tumor percentage (dTu%) which can be calculated by dividing the total number of invasive nuclei as observed in all patches of an image by the total number of nuclei as observed in all patches of an image. An image which has a dTu% below a certain threshold may be excluded from further analyses. For example, a dTu% below 20%, such as below 10% may be excluded from further analysis. Feature vector encoding The method of the invention further contains a step of encoding the extracted N patches into a set of feature vectors, wherein the input image is associated with N feature vectors. For example, the encoding comprises that each of the N patches is transformed to a feature vector of a certain length. For example, each of the N patches is encoded with a feature vector, i.e., the number of feature vectors resembles the number of patches (N).Optionally, the length of each of the N feature vectors is from 1-10000 such as from 8 – 2048. Optionally the length of each of the N feature vectors is 8; 16; 32; 64; 128; 256; 384; 512; 768; 1024 or 2048. Optionally, the length of each of the N feature vectors is 1024. The encoding is performed for example by using a convolutional neural network (CNN). A CNN comprises for example and input layer, one or more hidden layers, such as convolutional, pooling and / or flattening layers, and an output layer. An input layer for example receives the input image with its pixel values. The input layer for example has a stack of convolutional layers that are extracting the hierarchical features from the input image (e.g., LeNet-5 has 3 convolutional layers and ResNet50 has 49 convolutional layer). Each convolutional layer for example consists of filters to capture patterns and spatial hierarchies in different scales. Non-linear activation functions like ReLU (Rectified Linear Unit) are for example used after each convolutional operation to introduce non-linearity and enhance the model's capacity to learn complex features (ResNet50 has 32 activation layers). Pooling layers (e.g., MaxPooling or AveragePooling) down-sample (i.e., reduce) the spatial dimensions of the feature maps, reducing computational complexity and retaining essential information (e.g., ResNet50 has 16 pooling layers). Down- sampling for example refers to a process in which the spatial resolution or dimensions of an image or a feature map are reduced. It involves for example reducing the number of pixels or data points along the width and height of the image or feature map. In the context of pooling layers, down-sampling is for example achieved through operations like MaxPooling or AveragePooling. These operations involve dividing the input into non-overlapping or partially overlapping regions and summarizing each region by taking the maximum or average value, respectively. Flattening layers are for example used to flatten the output from the convolutional layers into a one-dimensional vector, preparing it for the subsequent fully connected layers which are connected to the output layer. For example, a feature vector encodes textural and morphological information of the respective patch. For example, the CNN, such as ResNet50, is designed to learn complex hierarchical features providing feature vectors which are rich of descriptive textural and morphological information. For example, the feature vectors comprise information regarding the form and / or structure of the tissues or cells within the image. For example, feature vectors comprise morphological details, such as shape, size, and arrangement of cells. For example, the feature vectors further comprise spatial, structural, texture and / or contextual information. For example, the feature vectors further comprise information how different structures are spatially arranged in the image, information on details about the texture or patterns within different regions of the image, information on specific structural patterns and abnormalities and / or information on relationships between different components in the image. The CNN is for example a pre-trained CNN. For example, the weights of the CNN are obtained / trained by training the CNN to classify natural images from a database. For example, the CNN is selected from ResNet50, ResNet34, EfficientNet, DenseNet, RegNet, ResNeXt or a combination thereof. For example, the CNN is ResNet50, optionally a pre-trained ResNet50. The ResNet50 used in the method of the invention is for example as described in He K., Zhang, X.; Shaoqing R. and Sun, J.; 2015; arXiv:1512.03385, Deep Residual Learning for Image Recognition. The ResNet50 architecture for example comprises: 1) Input Layer: Standard input layer, wherein the image data is fed into the CNN. 2) Initial Convolutional Layer: The initial convolutional layer has 64 filters with a kernel size of 7x7 and a stride of 2. This layer is responsible for the initial feature extraction. 3) Residual Blocks: ResNet50 is characterized by the use of residual blocks, each containing multiple convolutional layers. There are two main types of residual blocks with different numbers of layers. Basic Block containing two convolutional layers for smaller-sized blocks and Bottleneck Block containing three convolutional layers for larger-sized blocks. 4) Global Average Pooling Layer: After the stack of residual blocks, ResNet50 employs global average pooling, reducing the spatial dimensions to 1x1. For example, the ResNet50 has been pre-trained on ImageNet. For example, the pre-trained ResNet50 is obtained from Pytorch. For example, ResNet50 is pretrained on ImageNet50. For example, the weights of the ResNet50 are obtained by training it to classify natural images in the ImageNet database. For example, the N feature vectors are obtained by extracting the 3rdresidual block of a total of 5 blocks of a ResNet50 followed by an Average Pooling layer. For example, each of the N patches is transformed to a feature vector of length selected from 512; 768; 1024; 1280; 1536 or 2048. For example, each of the N patches is transformed to a feature vector of length 1024. For example, the N feature vectors are stored in PyTorch. In embodiments, the feature vector encoding is performed by a feature extractor that is trained using self-supervised learning (SSL). The resulting feature vector may comprise pathology self-supervised features. These features are representations extracted from pathology images using an SSL feature extractor. Examples of SSL methods that can be applied are the DINOv2 framework (Vorontsov et al., 2024. Nat Med. 30: 2924:2935), Barlow Twins (Zbontar et al., 2021 ICML 12310-12320, PMLR) and SwAV (Caron et al., 2020, NeurIPS, 33:9912– 9924). Further augmentations on top the default augmentations used in these frameworks may be applied such as methods described by Kang et al. (Kang et al., 2023. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.3344-3354) such as stain augmentation combined with vertical and horizontal flips. Providing an overall label output The method of the invention further comprises outputting an overall label output predicting a breast cancer recurrence status. For example, the overall label output is calculated and / or determined by aggregating the contribution (weight / attention score) of each of the N feature vectors to the overall label output. For example, the outputting of the overall label output is performed by using at least one fully-connected network (FCN) comprising a regression model, such as gated attention mechanism. For example, the gated attention mechanism is an adapted version of the attention-based multiple Instance learning (A-MIL) model. The A-MIL used in the method of the invention is for example as described in Ilse, M.; Tomczak J. M.; and Welling, M.; 2018; arXiv:1802.04712; Attention-based Deep Multiple Instance Learning. For example, such a model has a gated attention module and a fully connected network. The gated attention module has for example three fully connected layers. Two branches for example process the input features differently using tanh and sigmoid activation functions, and the outputs are for example element-wise multiplied. The result is for example then passed through a fully connected layer. The fully connected network for example takes the input features, processes them through a fully connected layer and uses the weights / attention score to produce a final output. For example, the A-MIL takes N feature vectors of a certain length as input. For example, the length of the N feature vectors for input is selected from 64; 128; 256; 384; 512; 768; 1024 and 2048. For example, the length of the N feature vectors for input is 1024. For example, the N feature vectors are propagated through a linear layer. The linear layer is for example of a length of from 8 – 2048. For example, the length of the linear layer is selected from 8; 16; 32; 64; 128; 256; 384; 512; 768; 1024 and 2048. The linear layer is for example of length 512. For example, the propagation is followed by Rectified Linear Unit (ReLU) activation and Dropout. For example, the resulting features are processed in two parallel streams of the gated attention mechanism. For example, the first stream comprises a linear layer. The linear layer of the first stream is for example of a length of from 8 – 2048 followed by Tanh activation and Dropout. For example, the length of the linear layer of the first stream is selected from 8; 16; 32; 64; 128; 256; 384; 512; 768; 1024 and 2048 followed by Tanh activation and Dropout. For example, the linear layer of the first stream is of length 384, which is followed by Tanh activation and Dropout. For example, the second stream of the gated attention mechanism is a linear layer. The linear layer of the second stream is for example of a length of from 8 – 2048 followed by Sigmoid activation and Dropout. For example, the length of the linear layer of the second stream is selected from 8; 16; 32; 64; 128; 256; 384; 512; 768; 1024 and 2048 followed by Sigmoid activation and Dropout. For example, the linear layer of the second stream is of length 384 followed by Sigmoid activation and Dropout. For example, Dropout of both layers is performed by multiplying each element with each other. For example, the result of the multiplication is passed through the last attention layer of the gated attention mechanism which is e.g. a linear layer of length 1. This is for example followed by a matrix multiplication between the hidden features that were the output of the first linear layer and the last attention layer from the input for the final classification layer. For this, the side level feature is for example of a length of from 8 – 2048 and propagated to a single node. For example, the side level feature is of a length selected from 8; 16; 32; 64; 128; 256; 384; 512; 768; 1024 and 2048 and propagated to a single node. For example, the side level feature is of length 512 and propagated to a single node. For example, the output of the single node is the overall output label predicting a breast cancer recurrence status (i.e., risk of breast cancer recurrence). The gated attention module of the A-MIL for example calculates a contribution of each of the N feature vectors by outputting a weight (attention score) for each feature vector assigned to each of the N patches, wherein the sum of all weights (attention scores) per input image is 1. For example, the higher the weight (attention score) of a feature vector is, the higher is the importance of the patch of said feature vector over all patches of a input image. Patches having a feature vector of high attention score (weight) are more important in predicting the overall label output. For example, high-attention scores (weights) originate from patches of a input image of a stained tissue sample which comprise tumor regions. For example, patches having a feature vector of high attention score (weight) are above the 75th; 80th; 85th; 90th; or 95thpercentile of all patches of an input image. For example, patches having a feature vector of high attention score (weight) are above 90thpercentile of all patches of an input image. Further, the contribution (weight / attention score) of each of the N feature vectors is for example aggregated by the A-MIL to the overall label output. Providing an image output of a low-resolution image version The outputting of a method according to the invention for example further comprises providing an image output of a low-resolution image version of the stained tissue sample. The low-resolution image version is for example achieved by down-sampling the original image of the stained tissue sample. For example, the low-resolution image version is down-sampled to a size / resolution where individual cells are no longer visible. For example, the low-resolution image version is down- sampled to a size selected from 8 x 8; 16 x 16; 32 x 32; 64 x 64; 128 x 128; 196 x 196; 256 x 256; 512 x 512; 768 x 768; 1024 x 1024; 2048 x 2048; or 4096 x 4096 pixels. Further, the low-resolution image version for example comprises 3 colorchannels (RGB). For example, the low-resolution image version is down-sampled toa size of 512 x 512 x 3, wherein “x 3” indicated the presence of three color channels (RGB). The low-resolution image version of the stained tissue sample may be used as input to a machine learning data processing model to obtain the image output of the low-resolution image version. For example, the machine learning data processing model relies on the global appearance of the stained tissue sample. For example, the machine learning data processing model used is a convolutional neural network (CNN). For example, the machine learning data processing model is selected from network architectures suitable for classification and regression, e.g. ResNet50, ResNet34, EfficientNet, ResNeXt, DenseNet, RegNet. For example, the machine learning data processing model is selected from vision transformers, e.g. ViT, DeiT, CaiT, Swin Transformer, T2T-ViT. For example, the classifier is a ResNet50. ResNet50 is a deep convolutional neural network architecture with for example 50 layers. It has for example residual blocks, each containing two convolutional layers and utilizing skip connections to address the vanishing gradient problem. The architecture for example comprises global average pooling and fully connected layers at the end, making it suitable for image classification tasks. The ResNet50 used in the method of the invention is for example as described in He K., Zhang, X.; Shaoqing R. and Sun, J.; 2015; arXiv:1512.03385, Deep Residual Learning for Image Recognition. Providing a slide-level prediction index by a Patch-based End-to-End Module The outputting of a method according to the invention for example further comprises providing a slide-level prediction index directly predicted from the pre- processed image patches or a subset thereof. In embodiments, a so called “Patch- based end-to-end module” (Fig. 1B) may be used to learn to predict a slide-level prediction index from the pre-processed image patches or a subset thereof. In embodiments, the input of the patch based end-to-end model is a subset of the pre-processed image patches. In embodiments, said subset of pre-processed image patches is selected from the pre-processed image patches based on tissue type information in a process called tissue-based patch selection. As explained herein above, using tissue type classification, for every patch information about the tissue type is obtained. These tissue types may be: invasive, infiltrated invasive, non-invasive, lymphocyte, stroma, tumor stroma, and other tissue types such as red blood cells, necrosis and artifacts. For tissue-based patch selection, the inputcomprises: (^^ × ^^ × ^^ × ^^) dimensional patches sampled from pre-processed WSIs,where M is the total number of patches within the whole slide, and (H x W x C) are the height-width and the number of color channels in each patch, respectively (such as 256 x 256 x 3). The tissue-based patch selection process takes an initial setof ^^ patches and reduces it to a subset of ^^ patches where ^^ ≪ ^^ based onpredefined criteria. To achieve this, it is assumed that access to a ^^ × ^^ discretetissue classification result for each patch, where ^^ represents the number of unique tissue categories. Then, a patch selection criteria may be devised based on the relative population of patches based on their tissue type. The selection isguided by K-tuple criteria (^^, ^^) where ^^ denotes the relative percentage of patchesto be sampled from each tissue category. The sum of all percentages satisfies the condition: ^^ Here, ^^^^is the percentage of tissue category ^^. The parameter ^^ is a hyperparameter that can be tuned to optimize the overall performance of the downstream prediction model. Adjusting these percentages allows flexibility in emphasizing certain tissue types that might contribute more significantly to the prediction task. This tailored selection ensures that the sampled subset of ^^ patches provides a balanced and informative representation of the slide while ignoring the uninformative part of the slide (e.g. fatty tissue). Three nested building blocks are typically used for a patch-based end-to-end module. A first block is an embedding block to represent the patches of the input image such as the input WSI. This may be in the form of a deep CNN or a Vision- Transformer. A second block is an aggregation layer to aggregate the representation of all the patches to emphasize the important parts of the slide. This may be a Multiple Instance Learning (MIL)-based approach, such as ABMIL (Ilse et al., 2018. arXiv:1802.04712), TransMIL (Shao et al., 2021. Advances in Neural Information Processing Systems 34: 2136-2147) or CLAM (Lu et al., 2021. Nat Biomed Eng 5:555-570). A third block is a regression layer such as a 1-layer multi-layer perception (MLP) to map the aggregated slide-level embedding to the slide-level prediction index. For example, a detailed exemplary Patch-based end-to-end module is described herein below, which takes as input the subsampled patches ^^ and then predicts a slide-level prediction index ^^^^^^^^^^using three main components, namely an embedding layer (embed), feature aggregation layer (agg) and regression layer (reg): = In this module, the patch embedding block (^^ × ^^) feature embeddings E areextracted from the extracted patches, represented by (^^ × ^^ × ^^ × ^^) , usingfollowing formula: ^^ = ^^ ( ) ^^×^^×^^×^^ ^^×^^embed ^^ , ^^ ∈ ^^ , ^^ ∈ ^^ ,where ^^ is the number of features from the last layer of the chosen arbitrary patch embedder. In this module, the aggregation layer is an attention-based aggregation layer,which aggregates ^^ × ^^ patch embeddings into a 1 × ^^ slide-level embedding S,using following formula: ^^ = ^^agg(^^), ^^ ∈ ^^^^×^^In this module, the regression layer maps the slide-level embedding to a one- dimensional slide-level prediction index ypred, using following formula: ^^ = ^^ (^^), ^^ ∈wherein he regression layer Perceptron with weightdimensions ^^ × 1 as the target prediction index is a scalar value. Providing a refined overall label output For example, the outputting of a method of the invention further comprises combining the overall label output, e.g. generated by the gated attention mechanism, such as A-MIL, and the image output of a low-resolution image version of the stained tissue sample for generating a refined overall label output. For example, the combining for generating a refined overall label output comprises inputting the overall label output and the image output of a low-resolution image version of the stained tissue sample in a (simple) fully-connected network (FCN), such as an ensemble model. In embodiments, the outputting of a method of the invention further comprises combining the overall label output, e.g., generated by the gated attention mechanism, such as A-MIL, and the output of the patch-based end-to-end module for generating a refined overall label output. For example, the combining for generating a refined overall label output comprises inputting the overall label output and the output of the patch-based end-to-end module in a (simple) fully- connected network (FCN), such as an ensemble model. In embodiments, the output of the patch-based end-to-end module used for generating a refined overall label output is at least one of the patch embeddings (E), the slide-level embeddings (S) and / or the patch-based predicted index. For example the output of the patch-based end-to-end module used for generating a refined overall label output is at least two of the patch embeddings (E), the slide-level embeddings (S) and the patch-based predicted index, for example the output of the patch-based end-to-end module used for generating a refined overall label output is the patch embeddings (E), the slide- level embeddings (S) and the patch-based predicted index. For example, the output of the patch-based end-to-end module is concatenated with the feature vectors of other models, such as the self-supervised feature vector. For example, the ensemble model is a machine learning algorithm suitable for regression, such as linear regression, decisions trees, gradient boosting etc.. The ensemble model for example comprises 2 – 10 layers. For example, the ensemble model comprises at least two; at least three; at least four; at least five; at least six; at least seven; least eight; at least nine; at least 10; at least 15; at least 20; at least 25; at least 30; at least 35; at least 40; at least 45; at least 50; at least 60; at least 70; at least 80; at least 90; or at least 100 layers. For example, the input layer of the ensemble model has two neurons, i.e. is of length 2. For example, the input layer of the ensemble model has two neurons (length 2), wherein the overall label output and / or the image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module serve as input of the ensemble model. For example, the ensemble model has 1 – 100 hidden layers. The number of hidden layers of the ensemble model may be adjusted to the complexity of the performed task. For example, the ensemble model has a single hidden layer with 64 neurons. For example, the output of the ensemble model is of length 1. For example, the output layer of length 1 of the ensemble model outputs a refined overall label output. For example, each layer except the last layer is followed by at least one activation function such as ReLU, ELU, GLU, and / or Leaky ReLU activation and Dropout. Activation functions which are for example used in the method of the invention are for example further described in Dubey, S.; Singh, S. and Chaudhuri, B.; 2022; Neurocomputing, Activation functions in deep learning: A comprehensive survey and benchmark. For example, each layer except the last layer is followed by ReLU activation and Dropout. For example, if the refined overall label output, the overall label output and / or the image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module is > 0 the risk of breast cancer recurrence is low. For example, if the refined overall label output, the overall label output and / or the image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module is ≤ 0 the risk of breast cancer recurrence is high. For example, the high risk of breast cancer recurrence group is further divided into 2 groups by taking the median of the high risk images of a stained tissue samples comprising breast cancer cells. High risk group 2 is defined as having a refined overall label output, overall label output and / or image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module higher than the median value of the high risk images, while High risk group 1 is defined as having a refined overall label output, overall label output and / or image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module lower than the median value of the high risk images. Training the models The method according to the invention is for example a pretrained method, i.e. no training of the machine learning data processing models used in the method of the invention is required. The method according to the invention for example comprises training of at least one or all machine learning data processing models used in the method of the invention. Training of at least one machine learning data processing model according to the invention for example comprises a) obtaining ground truth data; b) inputting example data into at least one machine learning data processing model used in the method of the invention; c) performing the at least one machine learning processing model according to the invention on the input example data; d) measuring the error between the generated output and the ground truth data; and e) updating the weights of the at least one or all machine learning data processing model according to the invention for performing the training of the machine learning data processing model. Optionally, the steps b) to e) are repeated until the error is no longer decreasing. For example, the weights of at least one machine learning data processing model according to the invention are updated using backpropagation. For example, ground truth data used for training at least one machine learning data processing model used in the method of the invention comprises a set of encoded N feature vectors corresponding with a set of N patches, a label providing information about a risk of breast cancer recurrence in a tissue sample corresponding with a set of N feature vectors, a label providing information about a risk of breast cancer recurrence in the tissue sample corresponding with a low- resolution image version of an image, and / or an output of the patch-based end-to end module of a stained tissue sample comprising breast cancer cells, a label providing information about a risk of breast cancer recurrence of an image of a tissue sample comprising breast cancer cells corresponding with the image of a stained tissue sample comprising breast cancer cells and / or a label providing information about a risk of breast cancer recurrence of an image of a tissue sample comprising breast cancer cells corresponding with the overall label output and the image output of a low-resolution output of the image and / or an output of the patch- based end-to end module of a stained tissue sample comprising breast cancer cells. The label is for example the result of a gene expression signature measured. For example, the label is the output of an analyzed MammaPrint gene set, i.e. a Mamma Print index. For example, example data used for training at least one machine learning data processing model used in the method of the invention comprises a set of N patches of an image of a stained tissue sample comprising breast cancer cells having corresponding feature vectors, wherein the corresponding feature vectors serve as ground truth data. For example, example data used for training at least one machine learning data processing model used in the method of the invention comprises a set of encoded N feature vectors having a corresponding label providing information about a risk of breast cancer recurrence in a tissue sample, wherein the corresponding label serves as ground truth data. For example, example data used for training at least one machine learning data processing model used in the method of the invention comprises a low- resolution image version of an image and / or an output of the patch-based end-to end module of a stained tissue sample comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in a tissue sample of the image, wherein the corresponding label serves as ground truth data. For example, example data used for training at least one machine learning data processing model used in the method of the invention comprises an overall label output and an image output of a low-resolution image version of an image and / or an output of the patch-based end-to end module of a stained tissue sample comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in a tissue sample of the image, wherein the corresponding label serves as ground truth data. For example, example data used for training at least one machine learning data processing model used in the method of the invention comprises an image of a stained tissue sample comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in a tissue sample of the image, wherein the corresponding label serves as ground truth data. For example, training a machine learning data processing model used for encoding N patches of an image of a stained tissue sample into a set of N feature vectors according to the invention comprises a) obtaining as ground truth data a set of encoded N feature vectors corresponding to a set of N patches; b) inputting as example data the set of the N patches into a processor comprising the machine learning data processing model used for feature vector encoding; c) performing the machine learning processing model on the input set of N patches; d) measuring the error between the generated feature vectors and the ground truth data; and e) updating the weights of the machine learning data processing model. Optionally, the steps b) to e) are repeated until the error is no longer decreasing. Training a method of the invention for example comprises inputting as example data an image of stained tissue samples comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in said tissue sample. The corresponding label providing information about a risk of breast cancer recurrence in said tissue sample for example provides ground truth data for measuring the error and updating the weights of the machine learning data processing model used in the method of the invention. For example, an image of stained tissue samples comprising breast cancer cells having a corresponding labels providing information about a risk of breast cancer recurrence in said tissue sample is used to train at least one or all machine learning data processing model used in the method of the invention. For example, the training comprises separating example data into three different subsets: a training, a validation and a test set. The test set is for example kept aside during the training of a method of the invention. To the example data, examples of stained tissue samples comprised in the training and validation, a set k-fold cross validation to train k models is for example applied. A method of the invention is for example trained k = 1 – 100. For example, a method of the invention is trained k = 1; 2; 3; 4; 5; 6; 7; 8; 9; 10; 11; 12; 13; 14; 15; 16; 17; 18; 19; 20; 25; 30; 35; 40; 45; 50; 55; 60; 65; 70; 75; 80; 85; 90; 95; 100. During each fold example data of stained tissue samples are for example randomly selected from a pool of example data of stained tissue samples. For example, 60,000 example data of stained tissue samples are selected for the training set and 5,000 example data of stained tissue samples are selected for the validation set. For example, training a method of the invention is performed on every example data of stained tissue samples from the training set, loss is calculated and backpropagation is used to update the weights of the machine learning data processing model, such as neural networks (NN), used in the method of the invention. Loss functions are for example used to measure the error between the generated output of at least one machine learning data processing model, such as neural networks (NN), used in the method of the invention and the ground truth data, such as a predefined label corresponding to the example data. A predefined label is for example obtained by gene expression analysis. For example, the output by gene expression analysis is the output of the MammaPrint Index measured. For example, Adam, AdamW, SGD, Adagrad, RMSprop and / or SparseAdam optimization is used to update the NNs’ weights of a method of the invention. For example, Adam optimization is used to updates the NNs’ weights of a method of the invention. Once at least one or all machine learning data processing model used in the method of the invention is performed on all example data of stained tissue samples from the training set, the weights of at least one or all performed machine learning data processing model are for example frozen and the performance of the at least one or all machine learning data processing model is for example evaluated on all example data of stained tissue samples from the validation set. The average loss on the example data of stained tissue samples from the validation samples is for example reported after every epoch, after every second epoch, after every third epoch or only once. The average loss is for example calculated as Mean Squared Error (MSE) loss which measures the average squared difference between the predicted values and the actual values. n is the number of samples in the dataset. yi is the actual (ground truth) value for the i-th sample. ŷi is the predicted value for the i-th sample. For example, the validation process is repeated until the validation loss is no longer decreasing, i.e., until the model is not improving. Alternatively, the average loss is calculated as Mean-absolute Error (MAE), Huber Loss, Quantile Loss, Weighted MSE and / or MAE. For example, once the validation loss is no longer decreasing the weights of the machine learning data processing model, e.g., NNs used are stored. For example, Adam, AdamW, SGD, Adagrad, RMSprop and / or SparseAdam optimization is used. For example, Adam optimization is used. For example, Adam optimization is used with a learning rate of 1xe-5; 1xe-4; 2xe-4or 5xe-4. For example, Adam optimization is used with a weight decay of from 1xe-5- 1xe-3, such as 1xe-5; 1xe-4or 1xe-3. For example, Adam optimization is used with a learning rate of 2e-4and a weight decay of 1e-5. The classification model for outputting an overall label output predicting a breast cancer recurrence status, such as A-MIL, is for example trained with an example data batch size of 1 – 100. For example, the classification model for outputting an overall label output predicting a breast cancer recurrence status, such as A-MIL, is trained with example data a batch size of at least 1; at least 3; at least 5; at least 10; at least 15; at least 20; at least 25; at least 30; at least 35; at least 40; at least 45; at least 50; at least 60; at least 70 or at least 80. A batch for example refers to all of the N feature vectors of one image data example of stained tissue samples. The training of the machine learning data processing model to classify the image output of a low-resolution image version of the stained tissue sample is similar to the training of the classification model for outputting an overall label output predicting a breast cancer recurrence status as described above. For example, training of the machine learning data processing model to classify the image output of a low-resolution image version of the stained tissue sample comprises a) obtaining as ground truth data a low-resolution image version of a stained tissue sample having a corresponding label providing information about a risk of breast cancer recurrence in the tissue sample of the low-resolution image; b) inputting as example data the low-resolution image version into a processor comprising the machine learning data processing model used for outputting the image output of a low-resolution image version; c) performing the machine learning data processing model; d) measuring the error between the generated image output of a low-resolution image version and the ground truth data; and e) updating the weights of the machine learning data processing model used for outputting the image output of a low-resolution image version for performing the training of the machine learning data processing model. Optionally the weights of the machine learning data processing model used for outputting the image output of a low- resolution image version are updated using backpropagation. Optionally the training is repeated until the error is no longer decreasing. The machine learning data processing model to classify the image output of a low-resolution image version of the stained tissue sample is trained for example by using cross-entropy loss, Binary Cross-Entropy Loss and / or Focal loss. The classification model to classify the image output of a low-resolution image version of the stained tissue sample is trained for example by using cross-entropy loss. For example, Adam, AdamW, SGD, Adagrad, RMSprop and / or SparseAdam optimization is used. For example, Adam optimization is used. For example, Adam optimization is used with a learning rate of 1xe-5; 1xe-4; 2xe-4or 5xe-4. For example, Adam optimization is used with a weight decay of from 1xe-5– 1xe-3, such as 1xe-5; 1xe-4or 1xe-3. For example, Adam optimization is used with a learning rate of 2e-4and a weight decay of 1e-5. For example, Adam optimization with a learning rate of 2e-4and a weight decay of 1e-5is used. The training of the machine learning data processing model to classify the output of the patch-based end-to end module of the stained tissue sample is similar to the training of the classification model for outputting an overall label output predicting a breast cancer recurrence status as described above. For example, training of the machine learning data processing model to classify the output of the patch-based end-to end module of the stained tissue sample comprises a) obtaining as ground truth data the output of the patch-based end-to end module of a stained tissue sample having a corresponding label providing information about a risk of breast cancer recurrence in the tissue sample image; b) inputting as example data the output of the patch-based end-to end module into a processor comprising the machine learning data processing model used for outputting the output of the patch-based end-to end module; c) performing the machine learning data processing model; d) measuring the error between the generated output of the patch-based end-to end module and the ground truth data; and e) updating the weights of the machine learning data processing model used for outputting the output of the patch-based end-to end module for performing the training of the machine learning data processing model. Optionally the weights of the machine learning data processing model used for outputting the output of the patch-based end-to end module are updated using backpropagation. Optionally the training is repeated until the error is no longer decreasing. The machine learning data processing model to classify the output of the patch-based end-to end module of the stained tissue sample is trained for example by using cross-entropy loss, Binary Cross-Entropy Loss and / or Focal loss. The classification model to classify the output of the patch-based end-to end module of the stained tissue sample is trained for example by using cross-entropy loss. For example, Adam, AdamW, SGD, Adagrad, RMSprop and / or SparseAdam optimization is used. For example, Adam optimization is used. For example, Adam optimization is used with a learning rate of 1xe-5; 1xe-4; 2xe-4or 5xe-4. For example, Adam optimization is used with a weight decay of from 1xe-5– 1xe-3, such as 1xe-5; 1xe-4or 1xe-3. For example, Adam optimization is used with a learning rate of 2e-4and a weight decay of 1e-5. For example, Adam optimization with a learning rate of 2e-4and a weight decay of 1e-5is used. During training image augmentation is for example applied to enhance model generalizability. Image augmentation for example comprises ColorJitter, RandomVerticalFlip, RandomHorizontalFlip and / or RandomRotation. The ColorJitter for example comprises brightness = 0.1, contrast = 0.05, saturation = 0.05 and / or hue = 0.2. The RandomRotation is for example 180°. For example all images are transformed to PyTorch tensor and normalized in each color channel. The images are for example normalized in each color channel to yield mean=(0.485; 0.456; 0.406), std=(0.229; 0.224; 0.225). For example, during validation only the transformation to tensors and / or normalization is applied to the images. For example, no augmentation is applied. The classification model to classify the image output of a low-resolution image version and / or the output of the patch-based end- to end module of the stained tissue sample is trained with a batch size of 1 – 100, such as at least 1; at least 3; at least 5; at least 10; at least 15; at least 20; at least 25; at least 30; at least 35; at least 40; at least 45; at least 50; at least 60; at least 70 or at least 80. A batch for example refers to all of the N feature vectors of one image data example of stained tissue samples. Training the machine learning data processing model for outputting a refined overall label output, such as an ensemble model, is for example similar to the training of the machine learning data processing model for outputting an overall label output predicting a breast cancer recurrence status and / or training of the machine learning data processing model to classify the image output of a low- resolution image version and / or the output of the patch-based end-to end module of the stained tissue sample as described above. For example, training of the machine learning data processing model for outputting a refined overall label output comprises a) obtaining as ground truth data an overall label output and an image output of a low-resolution image and / or the output of the patch-based end-to end module having a corresponding label providing information about a risk of breast cancer recurrence of an image a stained tissue sample comprising breast cancer cells, preferably the label is a result of gene expression profiling; b) inputting as example data the overall label output and the image output of a low-resolution image version of the image and / or the output of the patch-based end-to end module of a stained tissue sample; c) performing the machine learning data processing model for outputting a refined overall label output on the input overall label output and the input image output of a low-resolution image version and / or the output of the patch-based end-to end module, d) measuring the error between the generated refined overall label output and the ground truth data; and e) updating the weights of the machine learning data processing model used for outputting the refined label output for performing the training of the machine learning data processing model. Optionally the weights of the machine learning data processing model used for outputting the refined label output are updated using backpropagation. Optionally the steps b) to e) of the training are repeated until the error is no longer decreasing. Training the ensemble model for example comprises mean-squared-error loss. For example, training the ensemble model comprises Adam optimizer using a learning rate of 1e-3. For example, ensemble is trained with a batch size of 1 – 5000; 50 – 3000; 100 – 2000 or 200 – 1500. For example, ensemble is trained with a batch size of at least 1; at least 10; at least 50; at least 75; at least 100; at least 150; at least 200; at least 250; at least 300; at least 450; at least 500; at least 550; at least 600; at least 700; at least 800; a at least 850; at least 900; at least 1000; at least 1500; at least 2000; at least 2500; at least 3000. A batch for example refers to all of the N feature vectors of one image data example of stained tissue samples. The training of a method is for example performed for all models used in the method separately. The training of the method is for example performed for all models used in a method of the invention together. Computer program product The invention for example comprises a computer program or computer program product having instructions which when executed by a computing device or system to perform each of the steps of a method of the invention. Further, the invention comprises for example a computer program product embodied on a non- transitory computer readable medium comprising instructions stored thereon to cause one or more processors to perform a method of the invention. The computer program or computer program product is for example in the form of source or object code or in any other form suitable for use in the implementation of a method according to the invention. For example, the computer program or computer program product is for example stored on a carrier. A carrier is for example any entity or device capable of carrying the computer program or computer program product. For example, the carrier comprises a storage medium, such as a ROM, for example a semiconductor ROM or hard disk. Further, the carrier for example is a transmissible carrier such as an electrical or optical signal which is e.g. conveyed via electrical or optical cable or by radio or other means, e.g. via the internet or cloud. Device or apparatus The invention for example comprises a data-processing system, an apparatus and / or an device comprising means for carrying out a method of the invention. The data-processing system, apparatus and / or device for example comprises processors, microprocessors, circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, microchips, chip sets, et cetera. The data-processing system, apparatus and / or device is for example a computer, laptop, tablet, smartphone and / or a combination thereof. The data-processing system, apparatus and / or an device comprising means for carrying out a method of the invention for example further comprises a imaging device, such as a microscope, for imaging stained tissue samples obtained from individuals. For example, the stained tissue samples obtained from individuals are prepared on slides. The microscope is for example a stereomicroscope, compound microscope, fluorescence microscope, inverted microscope, a confocal microscope, scanning probe microscope, lightsheet microscope, spinning disk microscope and / or a combination thereof. Methods of treatment The invention is also directed to a method of treating an individual who has been assessed as having a high risk of breast cancer recurrence according to the computer-implemented method, computer program product and / or data-processing system of the invention. For example, the individual is assessed as being at high risk of breast cancer recurrence if the refined overall label output, the overall label output and / or the image output of a low-resolution image version of the stained tissue sample and / or the output of the patch-based end-to-end module is ≤ 0. The method of treatment according to the invention for example comprises assessing the risk of breast cancer recurrence in an individual according to the computer-implemented method of the invention, and providing a therapy for preventing and / or treating breast cancer to the individual assessed as having a high risk of breast cancer recurrence. The therapy for preventing and / or treating breast cancer for example includes any therapy of the state of the art which is known to be capable of preventing and / or treating breast cancer. The therapy is for example selected from chemotherapy, immunotherapy, stem cell therapy, hormone therapy, such as endocrine therapy, radiotherapy, targeted therapy, performance of surgery, or a combination thereof. For example, the method of treatment comprises administrating to an individual assessed as having a high risk of breast cancer recurrence any active compound, drug and / or pharmaceutical compound of the state of the art which is known to be capable of preventing and / or treating breast cancer. For example, assessing the risk of breast cancer recurrence according to the computer- implemented method of the invention is performed prior to administration of an active compound, drug and / or pharmaceutical compound. For example, an active compound, drug and / or pharmaceutical compound for preventing and / or treating breast-cancer is selected from selective oestrogen receptor modulators (SERMs) such as acolbifene (Endoceutics), afimoxifene (BHR Pharma, Atossa Therapeutics), arzoxifene (Eli Lilly and company), bazedoxifene (Pfizer), clomifene (Sanofi), droloxifene (Pfizer), endoxifen (Atossa Therapeutics), lasofoxifene (Pfizer), ospemifene (Osphena), pipindoxifene (LEAPChem), raloxifene (Daiichi Sankyo), tamoxifen (Rosemont Pharmaceuticals) and toremifene (Orion Corporation); selective oestrogen receptor down-regulators (SERDs) such as amcenestrant (also called SAR439859, Sanofi), AZD-9496 (AstraZeneca), AZD-9833 (AstraZeneca), brilanestrant (also called ARN-810 or GDC-0810, Genentech), D- 0502 (InventisBio), elacestrant (also called RAD-1901 or ER-306323, Radius Pharmaceuticals), etacstil (also called GW-5638 or DPC974, Bristol Myers Squibb), fulvestrant (AstraZeneca), giredestrant (also called GDC-9545, Roche), LSZ102 (Novartis), LY3484356 (Eli Lilly and company), rintodestrant (G1 Therapeutics), SHR9549 (Jiangsu HengRui Medicine) and ZN-c5 (Zeno Alpha); aromatase inhibitors (AI) such as 1,4,6-androstatriene-3,17-dione (ATD), 4-androstene-3,6,17- trione (4-AT), aminoglutethimide (Novartis), anastrozole (AstraZeneca), exemestane (Pfizer), fadrozole (Novartis), formestane (Novartis), letrozole (Novartis), testolactone (Bristol Myers Squibb) and vorozole (Janssen Pharmaceutica); gonadotropin-releasing hormone agonists to induce ovarian suppression such as leuprolide acetate (also called leuprolin, Sanofi and Astellas) and goserelin (AstraZeneca); immune checkpoint inhibitors inhibiting immune checkpoint molecules such as CTLA4, PD-1 and PD-L1, A2AR, CD276, B7-H4, CD272 and Herpesvirus Entry Mediator (HVEM), LAG3, NOX2, TIM-3, V-domain Ig suppressor of T cell activation (VISTA), and CD328; chemotherapeutic agents such as alkylating agents, anthracyclines, taxanes, histone deacetylase inhibitors, topoisomerase inhibitors and platinum-based agents; targeted therapy, preferably comprising treatment with a PARP inhibitor, with a PI3K / AKT / mTOR inhibitor, or with both a PARP inhibitor and a PI3K / AKT / mTOR inhibitor; or a combination thereof. Immune checkpoint inhibitors used in the method of treatment according to the invention are for example selected from following non-limiting examples: CTLA-4 inhibitors such as antibodies, including ipilimumab (Bristol-Myers Squibb) and tremelimumab (MedImmune); PD1 / PDL1 inhibitors such as antibodies, including pembrolizumab (Merck), sintilimab (Eli Lilly and Company), tislelizumab (BeiGene), toripalimab (Shangai Junshi Biosciense Company), spartalizumab (Novartis), camrelizumab (Jiangsu HengRui Medicine C), nivolumab and MDX- 1105 (Bristol-Myers Squibb), pidilizumab (Medivation / Pfizer), MEDI0680 (AMP- 514; AstraZeneca), cemiplimab (Regeneron) and PDR001 (Novartis); fusion proteins such as a PD-L2 Fc fusion protein (AMP-224; GlaxoSmithKline); atezolizumab (Roche / Genentech), avelumab (Merck / Serono and Pfizer), durvalumab (AstraZeneca), KN035 (Jiangsu Alphamab Biopharmaceuticals Company), Cosibelimab (CK-301; Checkpoint Therapeutics), BMS-936559 (Bristol- Myers Squibb), BMS-986189 (Bristol-Myers Squibb); and small molecule inhibitors such as PD-1 / PD-L1 Inhibitor 1 (WO2015034820; (2S)-1-[[2,6-dimethoxy-4-[(2- methyl-3-phenylphenyl)methoxy]phenyl] methyl]piperidine-2-carboxylic acid), BMS202 (PD-1 / PD-L1 Inhibitor 2; WO2015034820; N-[2-[[[2-methoxy-6-[(2- methyl[1,1'-biphenyl]-3-yl)methoxy]-3-pyridinyl]methyl]amino] ethyl]-acetamide), PD-1 / PD-L1 Inhibitor 3 (WO / 2014 / 151634; (3S,6S,12S,15S,18S,21S,24S,27S,30R,39S,42S,47aS)-3-((1H-imidazol-5-yl)methyl)- 12,18-bis((1H-indol-3-yl)methyl)-N,42-bis(2-amino-2-oxoethyl)-36-benzyl-21,24- dibutyl-27-(3-guanidinopropyl)-15-(hydroxymethyl)-6-isobutyl-8,20,23,38,39- pentamethyl-1,4,7,10,13), CA-170 (Curis) and ladiratuzumab vedotin (Seattle Genetics). Chemotherapeutic agents used in the method of treatment according to the invention are for example selected from following non-limiting examples: alkylating compounds such as bendamustine (Mundipharma Pharmaceuticals), busulfan (Pierre Fabre), carmustine (Bristol-Myers Squibb), chlorambucil (Aspen), cyclophosphamide (Baxter), dacarbazine (Pfizer), estramustine (Pfizer), ifosfamide (Baxter), lomustine (Kyowa Kirin Pharma), melphalan (GlaxoSmithKline), nimustine (Sankyo), procarbazine (Leadiant Biosciences), streptozotocin (Keocyt), temozolomide (Merck & Co), thiotepa (Adienne), treosulfan (Lamepro) and trofosfamide (Baxter); anthracyclines such as daunorubicin (Medac), doxorubicin (Pfizer), epirubicin (Pfizer), idarubicin (Pfizer), mitoxantrone (Pfizer), pirarubicin (Sanofi), pixantrone (Servier) and valrubicin (Endo Pharmaceuticals); anti-tumor antibiotics (not anthracyclines) such as bleomycin (Inovio Pharmaceuticals), dactinomycin (Ovation Pharmaceutical) and mitomycin (UroGen Pharma); platinum compounds such as cisplatin (Bristol Myers Squibb), carboplatin (Bristol Myers Squibb), oxaliplatin (Pfizer) and satraplatin (Yakult Honsha), antimetabolites such as azacitidine (Pfizer), capecitabine (Roche), cytarabine (Pfizer), cladribine (Janssen Pharmaceutica), clofarabine (Sanofi), decitabine (Janssen Pharmaceutica), fludarabine (Bayer), (5-)fluorouracil (FivepHusion), 5- fluoro-2´-deoxyuridine (Sigma-Aldrich), gemcitabine (Eli Lilly and Company), (6- )mercaptopurin (Aspen), methotrexate (Aldeyra Therapeutics), nelarabine (Novartis), pemetrexed (Eli Lilly and Company), pentostatin (Pfizer) and (6- )tioguanine (Aspen); anti-mitotic cytostatics such as vinblastine (Teva), vincristine (Teva), vindesine (EG), vinflunine (Pierre Fabre) and vinorelbine (Pierre Fabre), taxanes such as cabazitaxel (Sanofi), docetaxel (Sanofi), paclitaxel (Celgene) and tesetaxel (Odonate Therapeutics); non-taxane microtubule inhibitors such as eribulin (Eisai), indibulin (Baxter), ixabepilone (R-PHARM), patupilone (Novartis) and sagopilone (Bayer HealthCare); topo-isomerase inhibitors such as camptothecin (RTI International), etoposide (Bristol-Myers Squibb), irinotecan (Pfizer), teniposide (Bristol-Myers Squibb), tretinoin (Roche) and topotecan (Novartis); and histone deacetylase inhibitors such as chidamide (Chipscreen Bioscience, HUYA Bioscience International), entinostat (Syndax), mocetinostat (Mirati therapeutics), tacedinaline (Pfizer), domatinostat (4SC), romidepsin (Celgene), abexinostat (Xynomic Pharmaceuticals), belinostat (Onxeo), nanatinostat (CHR-3996, Chroma Therapeutics, Viracta Therapeutics), givinostat (ITALFARMACO), MPT0E028 (3-(1-benzenesulfonyl-2,3-dihydro-1H-indol-5-yl)-N- hydroxy-acrylamide), panobinostat (Secura Bio Limited), pracinostat ((E)-3-[2- butyl-1-[2-(diethylamino)ethyl]benzimidazol-5-yl]-N-hydroxyprop-2-enamide), quisinostat (Janssen Pharmaceuticals, NewVac), resminostat (4SC, Yakult Honsha), ricolinostat (Regenacy Pharmaceutica), trichostatin A (Vanda Pharmaceuticals), vorinostat (suberanilohydroxamic acid or SAHA, Merck & Co), butyric acid, 4-phenylbutyric acid, pivanex (pivaloyloxymethyl butyrate), valproic acid (2-propylpentanoic acid), cambinol (5-[(2-Hydroxynaphthalen-1-yl)methyl]-6- phenyl-2-thioxo-2,3-dihydropyrimidin-4(1H)-one), selistat (EX-527, AOP Orphan Pharmaceuticals AG), nicotinamide (pyridine-3-carboxamide) and sirtinol (2-[[(2- hydroxy-1-naphthalenyl)methylene]amino]-N-(1-phenylethyl)-benzamide). A PARP inhibitor used in the method of treatment according to the invention are for example selected from following non-limiting examples: Olaparib (3-aminobenzamide, 4-(3-(1-(cyclopropanecarbonyl)piperazine-4-carbonyl)-4- fluorobenzyl)phthalazin-1(2H)-one; AZD-2281; AstraZeneca), rucaparib (6-fluoro-2- [4-(methylaminomethyl)phenyl]-3,10-diazatricyclo[6.4.1.04,13]trideca-1,4,6,8(13)- tetraen-9-one; Clovis Oncology, Inc.); niraparib tosylate ((S)-2-(4-(piperidin-3- yl)phenyl)-2H-indazole-7-carboxamide hydrochloride; MK-4827; GSK); talazoparib (11S,12R)-7-fluoro-11-(4-fluorophenyl)-12-(2-methyl-1,2,4-triazol-3-yl)-2,3,10- triazatricyclo[7.3.1.05,13]trideca-1,5(13),6,8-tetraen-4-one; BMN-673; Pfizer); veliparib (2-[(2R)-2-methylpyrrolidin-2-yl]-1H-benzimidazole-4-carboxamide dihydrochloride benzimidazole carboxamide; ABT-888; Abbvie); pamiparib (2R)-14- fluoro-2-methyl-6,9,10,19-tetrazapentacyclo[14.2.1.02,6.08,18.012,17]nonadeca- 1(18),8,12(17),13,15-pentaen-11-one; BGB-290; BeiGene); CEP-8983, and CEP 9722, a small-molecule prodrug of CEP-8983, a 4-methoxy-carbazole inhibitor (CheckPoint Therapeutics); E7016 (Eisai), PJ34 (2-(dimethylamino)-N-(6-oxo-5H- phenanthridin-2-yl)acetamide;hydrochloride) and 3-aminobenzamide. EXAMPLES Example 1: A precise example of performing the method of the invention using pretrained models and generating a refined overall label output The method is applied on a digitized whole slide image (WSI) file of a patient with early stage breast cancer. The first step is Tissue Detection. The initial phase involves the detection of relevant tissue regions within whole slide images. This is achieved by performing the following steps: 1) Down-sampling (s = 256): down-sampling of the image by the factor s = 256,e.g. an image of size (18500, 27800) pixel will be down-sampled to an image of size (73, 109) pixel. 2) Greyscale Transformation: Conversion of the image to greyscale for uniformrepresentation. 3) Optical Density Conversion: Transformation of pixel values to opticaldensity (OD). 4) Thresholding Criteria: Pixels with OD values > 0.1 are retained. Variance ofpixel values across color channels must exceed 0.001. 5) Boundary Exclusion: Removal of pixels at the boundary (0.001 x (thesmaller dimension)). 6) Object Exclusion: Elimination of objects smaller than 0.1 x (the area of thelargest object). 7) Elongation Factor Exclusion: Discarding objects with an elongation factor(R) greater than 0.95. 8) Tissue Mask Creation: Generation of the final tissue mask for subsequentanalysis. Following tissue detection, relevant patches are extracted based on the acquired tissue mask. See, an examples, Figures 2 and 3. 1) Tissue Mask Utilization: Leveraging the tissue mask to extract patches ofsize (256 x 256 x 3) from the full-resolution slide image.2) Patch Details: The number of extracted patches (N) is determined by thequantity of tissue present, resulting in a dataset of dimensions [N, 256, 256, 3]. Following patch extraction, the N patches are encoded into N feature vectors in a step called Feature Extraction. A pre-trained ResNet50 model on the ImageNet dataset for feature extraction is used. 1) Model Utilization: Each extracted patch is fed into the ResNet50 model.2) Feature Vector Generation: Extraction of the output of the 3rd residualblock followed by an Average Pooling Layer. 3) Feature Vector Transformation: Conversion of each of the N patches intoa feature vector of length 1024, resulting in a dataset of dimensions [N, 1024]. 4) These extracted feature vectors serve as input to the Core AI model.Next step is to create a Tiny Whole Slide Image that serves as input to the Tiny WSI model. 1) Transformation of the entire whole slide image (WSI) through re-sizing ordown-sampling, resulting in an image of dimensions 512 x 512 pixels and 3 color channels. Next step is to generate with the use of the extracted feature vectors an overall label output in the Core AI model and with the use of the down-sampled low- resolution image an image output of a low-resolution image version in the Tiny WSI model. Application of the trained Core AI and Tiny WSI models involves: 1) Tiny WSI model is a ResNet50 model where the input is the Tiny WSI (512x 512 x 3 pixel) and the output is the image output of a low-resolution image version of the stained tissue sample. 2) Core AI model is visualized in Fig. 4. The input to the model is the slide’sencoded feature vectors, [N, 1024] and the output is the overall label output. Next step is to generate a refined overall label output (Digital Mamma Print Index) by the ensemble model. The ensemble model is visualized in Fig. 5. The input to the ensemble model are the output of the Tiny WSI model (image output of a low- resolution image version of the stained tissue sample) and the output of the Core AI model (overall label output). The output is digital MammaPrint Index (refined overall label output). This is the final reported digital MammaPrint Index. See, for example, Fig.6. Example 2: One example of training the models with exemplary training data For training, digitized whole slide image (WSI) files of a patient with early stage breast cancer and MammaPrint Indices obtained by performing MammaPrint genomic test are used. The reported MammaPrint values are coupled with the slides and there is one reported value per patient. a) Training the Tiny WSI model The training of the Tiny WSI model requires the utilization of whole slide images in conjunction with reported MammaPrint Indices. The training procedure follows the outlined steps: Image Preprocessing: 1) Downsampling of whole slide images to dimensions of (512 x 512 x 3) pixels.2) Direct utilization of unmodified MammaPrint indices.Model Selection: ResNet50 architecture with a single output channel for model training. Training Parameters: The training of the Tiny WSI model involves the utilization of typical training parameters, listed in table 1: Table 1. Typical training parameters. Parameter ValueLoss Mean Squared ErrorOptimizer AdamLearning Rate 1e-4Batch Size 16b) Training the Core AI model The training of the Core AI model requires the encoded whole slide image feature vectors in conjunction with reported Microarray MammaPrint Indices. As explained before, the whole slide image has to go through the following steps in order to get the encoded feature vectors: 1) Tissue detection2) Patch extraction3) Feature extractionInput Specifications: 1) Encoded feature vectors per slide2) Utilization of unaltered MammaPrint Indices for model training.Model Configuration: Deployment of the Core AI model, as visualized in Fig. 4, for the training procedure. Training Parameters: The training of the Core AI model involves the utilization of typical training parameters, listed in table 2: Table 2. Typical training parameters. Parameter ValueLoss Mean Squared ErrorOptimizer AdamLearning Rate 2e-4Batch Size 1Weight Decay 1e-5c) Training the Ensemble model The Ensemble model is visualized in Fig. 5. Model Inputs: 1) The input to the ensemble model are the predicted MammaPrint Indices ofthe two Tiny WSI and Core AI models (outputs of these models). 2) Utilization of unaltered MammaPrint Indices for model trainingModel Configuration: Deployment of the Core AI model, as visualized in Fig. 4, for the training procedure. Training Parameters: The training of the ensemble model involves the utilization of typical training parameters, listed in table 3: Table 3. Typical training parameters. Parameter ValueLoss Mean Squared ErrorOptimizer AdamLearning Rate 1e-4Batch Size 64Example 3: Another precise example of performing the method of the invention using pretrained models and generating a refined overall label output The method was applied on digitized whole slide image (WSI) files of patients with early stage breast cancer. Patients included female breast cancer patients with stage I or stage II disease who are lymph node negative or lymph node positive with up to 3 positive nodes, with a tumor size less than or equal to 5.0 cm, and for patients with stage III disease. An overview of the method is visualized in Fig. 1b and Fig.7. The first steps include Tissue Detection and Patch Extraction, which were performed according to Example 1. Hereafter, a tissue type classification task and nuclei detection was performed. Tissue type classification A Multi-Scale Feature Combination Model was developed using PyTorch. The model utilized ResNet50 as its backbone convolutional neural network (CNN) architecture. To improve its generalization to new, unseen data, the model incorporated techniques such as dropout and batch normalization. Additionally, the architecture included an embedding layer to reduce the dimensionality of the feature representations and a SoftMax-based classifier for tissue type prediction. The model classifies the detected tissue to the following seven tissue types: invasive, infiltrated invasive, non-invasive, lymphocyte, stroma, tumor stroma, and rest (e.g., red blood cells, necrosis, artifacts, etc.). The input to the model comprised three 256×256 pixel patches captured at different magnifications: 20x, 10x, and 5x (see Figure 8). By integrating patches from multiple magnifications, the model gains a more comprehensive understanding of tissue structures and features, enabling more accurate and informed tissue type classification. This multiscale approach effectively captures both fine-grained details and larger-scale patterns within histopathology images, enhancing classification accuracy and robustness. The model's output is the predicted tissue type for the 256×256 pixel patch at 20x magnification. 191 partially annotated slides were used to train the algorithm. During training, the hyperparameters were set as following: batch size = 64, dropout = 0.40 and learning rate = 0.0001. In Figure 9, an example of an input and output of the tissue type classification model is shown, wherein each dot represents the predicted label for a 256x256 pixel tissue patch. Nuclei Detection For nuclei detection, the advanced AI model called ‘You Only Look Once’ version 8 (YOLOv8) was used to process each patch extracted during the Patch Extraction step. To balance inference time and performance, the small version of YOLOv8 (YOLOv8s, Jocher et al., 2023. Ultralytics, available at github.com / ultralytics / ultralytics) with 11.2 million parameters was chosen for the task. YOLOv8s is capable of detecting and classifying nuclei into four categories: tumor, stroma, lymphocyte and others, as shown in Figure 10. A dataset consisting of around 180k patches, extracted from 82 H&E slides was used to train the model. The dataset was randomly split into training and validation sets, with 80% of the data allocated for training and 20% for validation. During training, the hyperparameters were set as follows: input tensor size of 256, batch size of 60, and 50 epochs. Digital tumor percentage module Tissue type classification and nuclei detection models were applied to each identified tissue patches. As a result, for every tissue patch, information about its tissue type and the count and types of nuclei within it was obtained. The following formula was used to calculate the so-called digital tumor percentage (dTu%): Following criteria were set, when a certain slide met these criteria, it was excluded from further analysis: ^connected component analysis was performed using only the invasive tissuepatches; ^If the largest connected component consisted of fewer than 100 patches andthe second largest connected component had fewer than 50 patches. In addition thereto, slides with a tumor percentage of less than 10 % were excluded, whereby ^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^  =^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^ . As such, an accurate representation of tissue quality was provided. It is noted that in this example no images were excluded, as slides had been reviewed by a pathologist and were confirmed to encompass more than 30 % of tumor cells. Furthermore, a patch selection criteria was applied in this example as follows: “invasive": 1024; "infilterated_invasive": 512, "noninvasive": 256, "lymphocyte": 256, "stroma": 128, "tumor_stroma": 256, "rest": 0. Pathology self-supervised features In contrast to the ImageNet pre-trained feature extractor of Example 1, in this example a pathology self-supervised feature extraction was performed. Self-supervised learning (SSL), which leverages large amounts of unlabeled data, has shown to have superior performance in pathology specific tasks over supervised ImageNet pre-training in certain instances. A well-known SSL framework is DINOv2 (Vorontsov et al., 2024. Nat Med. 30: 2924:2935) was used for feature extraction of pathology images. Based on Kang et al. (Kang et al., 2023. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.3344-3354) two augmentations were applied during pretraining to the input images within the DINOv2 framework, i.e. stain augmentation and using vertical flip together with horizontal flip. Model training: To ensure robust and generalized feature extraction, the dataset was balanced across important characteristics, including the MammaPrint index, year of acquisition, country of origin of the sample, and center in which the scanning took place (Irvine and Amsterdam). Furthermore, slides scanned on a different scanner (e.g. Roche scanner) were used to make the features robust to the choice of scanner. Patches of size 256x256 pixels were extracted from the dataset, serving as the input images for the feature extractor pretraining phase. The training of the SSL feature extractor involves certain parameters, that can be seen in Table 4. Table 4. Typical training parameters for the SSL feature extractor. Parameter ValueLoss DINO loss + iBOT loss (both based onCross entropy loss). Optimizer AdamWLearning Rate 0.004Batch Size 128Patch-based end-to-end module Next a so called “Patch-based end-to-end module” (Fig. 1B) was used to learn to predict a slide-level prediction index directly from the pre-processed whole slide image patches. The input comprised: (^^ × ^^ × ^^ × ^^) dimensional patches sampled fromWSIs, where M is the total number of patches within the whole slide, and (H x W x C) are the height-width and the number of color channels in each patch, respectively (such as 256 x 256 x 3). The tissue-based patch selection process takes an initial set of ^^ patchesand reduces it to a subset of ^^ patches where ^^ ≪ ^^ based on predefined criteria(see, for example, Fig. 7). To achieve this, access to a ^^ × ^^ discrete tissueclassification result for each patch was assumed, where ^^ represents the number of unique tissue categories. Then, a patch selection criteria was devised based on the relative population of patches based on their tissue type. The selection is guided byK-tuple criteria (^^,^^) where ^^ denotes the relative percentage of patches to besampled from each tissue category. The sum of all percentages satisfies the condition: ^^ Here, ^^^^is the percentage of tissue category ^^. The parameter ^^ is a hyperparameter that can be tuned to optimize the overall performance of the downstream prediction model. The Patch-based end-to-end module takes as input the subsampled patches ^^ and then predicts slide-level prediction index ^^^^^^^^^^using three main components, namely embedding layer (embed), feature aggregation layer (agg) and the regression layer (reg): = These three blocks are 1) Patch Embedding Block: Patch embedding block is responsible for representing the appearance of a givenpatch. Given (^^ × ^^ × ^^ × ^^) tissue-selected patches, this block extracts (^^ × ^^)feature embeddings: ^^ = ^^embed(^^), ^^ ∈ ^^ ^^×^^×^^×^^ , ^^ ∈ ^^^^×^^where ^^ is the number of features from the last layer of the chosen arbitrary patch embedder. 2) Attention-Based Aggregation Layer:This layer aggregates ^^ × ^^ patch embeddings into a 1 × ^^ slide-level embedding:^^ = ^^agg(^^), ^^ ∈ ^^^^×^^ 3) Regression Layer: The final layer maps the slide-level embedding to a one-dimensional index: ^^pred = ^^reg(^^), ^^pred ∈ ^^The regression layer is a simple 1-layer Multi-Layer Perceptron with weightdimensions ^^ × 1 as our target prediction index is a scalar value.After the model makes a prediction, then this prediction is used to calculate the loss function, and finally update the respective layer parameters. The Patch-based end-to-end module model was trained using a Mean Squared Error (MSE) loss to minimize the discrepancy between the predicted index and the ground truth: ^^ where B is the batch size, and the ^^^^^^^^^^,^^is the ground truth predicted index of a given whole slide image of a patient ^^.The patch embedding backbone is decomposed into layers {layer0, layer1, … , layer^^}.During training, a selective combination of these layers (e.g. last few layers of the embedder) is updated alongside the randomly initialized patch aggregation and regression layers. This flexibility allows optimization of critical features while preventing over-fitting. The model using ADAMW (Loshchilov, 2017. arXiv:1711.05101) optimizer with a learning rate of 0.0001 decayed every 25 epochs with a factor of 0.1 for a total of 100 epochs. An effective batch size of 16 was used to train the model. For all parameters, please refer to the below Table. Table 5. Typical training parameters for patch-based end-to-end module. Parameter ValueLoss Mean Squared ErrorOptimizer AdamWLearning Rate 1e-4Learning Rate Decay Linear (Factor of 0.1, every 25 epochs) Trained Backbone ResNet-50Batch Size 16Epoch 100Weight Decay 1e-5Ensembling In the next step a refined overall label output (Digital Mamma Print Index, dMPI) was generated by the ensemble model. The ensemble model is the same as in previous examples, only the input differs. Here, as input to the ensemble model theslide-level index prediction ypred obtained as output of patch-based end-to-endmodule is used. The local patch-level features are concatenated to the patch-level features of the other models, here from the self-supervised model. The global slide- level features and the slide-level index prediction ypred is concatenated to the slide- level features of the Core AI module before predicting the final digital MammaPrint index. This is the final reported digital MammaPrint Index. Example 4: Prediction of recurrence of breast cancer using the digital MammaPrint Index in an independent dataset. The analysis is done using the model described in Example 3 on whole-slide images and the clinical outcomes of 261 patients, obtained from two separate cancer centers (NorthShore and Fox Chase Cancer Center) in the United States from 1992-2010. Of the 261 analyzed patients, 145 were classified as dMP High-Risk and 116 were classified as dMP Low-Risk. From the 261 patients, 15 developed distant recurrence, i.e. metastasis (event). A Kaplan-Meier survival analysis is shown in Figure 11, showing that the digital MammaPrint Index is a very accurate predictor of breast cancer recurrence. As can be seen, there was only one distant recurrence in the low risk group, while 14 distant recurrences were observed in the high risk group. In Fig.11 the probability of survival shown on the y-axis represents occurrences of breast cancer recurrence, more specifically distant recurrence, also called metastasis. Statistics: Log Rank (Mantel-Cox), Chi-Square = 19.659, df=1 and significance < 0.00001.

Claims

Claims 1. A computer-implemented method for assessing a risk of breast cancer recurrence, the method comprising the steps of a) inputting an image of a stained tissue sample comprising breast cancer cells from an individual into a processor; b) pre-processing the image in order to achieve a tissue segment image of relevant image areas by the processor; c) patch extraction, wherein N patches of the pre-processed tissue segment image are generated by the processor; d) encoding using at least one machine learning data processing model the extracted N patches into a set of feature vectors, wherein the input image is associated with a set of N feature vectors by the processor; and e) outputting using at least one further machine learning data processing model an overall label output predicting a breast cancer recurrence status by aggregating the contribution of each of the N feature vectors to the overall label output by the processor.

2. The computer-implemented method of claim 1, wherein the image is a whole slide image (WSI).

3. The computer-implemented method of any one of the preceding claims,wherein the pre-processing comprises extracting uninformative image content comprising background and / or adipose tissue, preferably wherein the pre-processing comprises the steps of: a) down-sampling of the image by a factor s;b) transforming the image to greyscale;c) transforming each pixel-value of the image to optical density(OD); d) creating a mask of pixels by defining an OD threshold level,preferably the OD threshold is > 0.1; e) excluding pixels at the boundary, preferably the boundaryregion is defined by 0.03*min (width, height);f) excluding objects which are smaller than 0.1 – 1* area largestobject in mask; g) excluding objects of a certain length, preferably wherein objectsare excluded of a length of (Lmax – Lmin) / Lmax > 1.5 to 0.5, more preferably of a length of (Lmax – Lmin) / Lmax > 0.95; and / or h) excluding marking, preferably wherein the marking is definedby a variance across RGB channels < 0.001.

4. The computer-implemented method according to claim 3, wherein the down-sampling factor s matches one dimension of the size of the extracted patches, preferably the down-sampling factor matches one dimension of the size of the extracted patches of 256 x 256, more preferably the down-sampling factor is 256.

5. The computer-implemented method according to any one of the preceding claims, wherein each patch is transformed to a feature vector of a certain length, preferably wherein each feature vector encodes textural and / or morphological information of each patch.

6. The computer-implemented method according to any one of the precedingclaims, wherein after patch extraction, an additional quality control step is performed, preferably wherein for every patch a tissue type classification and / or nuclei detection is performed.

7. The computer-implemented method according to any one of the preceding claims, wherein the encoding of the feature vectors is performed using a convolutional neural network (CNN) comprising an input layer, one or more hidden layers and an output layer, preferably wherein the CNN is ResNet50, more preferably wherein the encoding of feature vectors is performed by using a Resnet50 having a length 1024.

8. The computer-implemented method according to any one of claims 1-6, wherein the encoding of the feature vectors is performed using self-supervised learning.

9. The computer-implemented method according to any one of the preceding claims, wherein the at least one further machine learning data processing model used for outputting of the overall label output is at least one fully-connected network (FCN) comprising a gated attention mechanism, preferably wherein the at least one further machine learning data processing model is an adapted version of the attention-based multiple Instance learning (A- MIL) model.

10. The computer-implemented method according to claim 9, wherein feature vectors that receive high attention from the at least one fully-connected network (FCN) comprising a gated attention mechanism originate from tumor regions in the stained tissue sample.

11. The computer-implemented method according to claim 9 or claim 10, wherein the outputting comprises the steps of: a) inputting the N feature vectors; b) propagating the input N feature vectors through a linear layer of a certain length, preferably the length is 512; c) performing Rectified Linear Unit (ReLU) activation and dropout; d) processing the resulting features by at least two parallel streams of gated attention mechanism, preferably wherein the first stream consists of a linear layer of length 384, followed by Tanh activation and dropout and the second stream consists of a linear layer of length 384, followed by Sigmoid activation and dropout; e) multiplying the output of the at least two attention layers element- wise with each other; f) passing the multiplied result through the last layer of the gated attention, wherein the last layer is a linear layer, preferably wherein the length is 1;g) providing input for the last classification layer by matrix multiplication between the hidden features that were the output of the first linear layer and the last attention layer; h) propagating a slide-level feature to a single node, preferably wherein the slide-level feature is of length 512; i) providing the overall label output.

12. The computer-implemented method according to any one of the preceding claims, further comprising a) providing a slide-level prediction index predicted by a Patch-based End-to-End Module from the pre-processed image patches; and b) combining the overall label output and the output of step a) to generate a refined overall label output by using a further machine learning data processing model, preferably wherein the further machine learning data processing model used for combining is at least one further fully-connected network (FCN), more preferably wherein the at least one further FCN is an ensemble model.

13. The computer-implemented method according to claim 12, wherein the combining comprises the steps of: a) inputting the overall label output and the output of the Patch-based End-to-End Module in an ensemble model, comprising at least three linear layers, wherein each layer except the last layer is followed by ReLU activation and dropout; and b) generating a refined overall label output.

14. The computer-implemented method according to any one of the preceding claims, wherein the staining of the tissue sample is heamatoxylin and eosin (H&E).

15. The computer-implemented method according to any one of the preceding claims, further comprising training the machine learning data processing model used for encoding the extracted N patches into a set of feature vectors, wherein the training comprises the steps of:a) obtaining as ground truth data a set of N patches havingcorresponding encoded N feature vectors; b) inputting as example data the set of N patches into a processorcomprising the machine learning data processing model used for encoding the extracted N patches into a set of N feature vectors; c) performing the method according to step d) of claim 1 on the inputset of extracted N patches; d) measuring the error between the generated feature vectors of theextracted patches and the ground truth data; and e) updating the weights of the machine learning data processing modelused for feature vector encoding for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for feature vector encoding are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing.

16. The computer-implemented method according to any one of the preceding claims, further comprising training the at least one further machine learning data processing model used for outputting the overall label output wherein the training comprises the steps of: a) obtaining as ground truth data a set of N feature vectors having acorresponding label providing information about a risk of breast cancer recurrence in a tissue sample, preferably the label is a result of gene expression profiling; b) inputting as example data the set of N feature vectors into aprocessor comprising the at least one further machine learning data processing model used for outputting the overall label output; c) performing the method according to step e) of claim 1 on the input setof N feature vectors; d) measuring the error between the generated overall label output andthe ground truth data; and e) updating the weights of the at least one further machine learningdata processing model used for outputting the overall label output forperforming the training of the machine learning data processing model, preferably wherein the weights of the at least one further machine learning data processing model used for outputting the overall label output are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing.

17. The computer-implemented method according to any one of the preceding claims, further comprising training the machine learning data processing model used for outputting the output of the Patch-based End-to-End Module, wherein the training comprises the steps of: a) obtaining as ground truth data an output of the Patch-based End-to-End Module of a stained tissue sample comprising breast cancer cells having a corresponding label providing information about a risk of breast cancer recurrence in the tissue sample, preferably the label is a result of gene expression profiling; b) inputting as example data the output of the Patch-based End-to-EndModule of a stained tissue sample into a processor comprising the machine learning data processing model used for outputting the output of the Patch- based End-to-End Module; c) performing the method according to step a) of claim 12d) measuring the error between the generated output of the Patch-based End-to-End Module and the ground truth data; and e) updating the weights of the machine learning data processing modelused for outputting output of the Patch-based End-to-End Module for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for outputting output of the Patch-based End-to-End Module are updated using backpropagation; optionally wherein steps b) to e) are repeated until the error is no longer decreasing.

18. The computer-implemented method according to any one of the preceding claims, further comprising training the machine learning data processing model used for outputting the refined overall label output, wherein the training comprises the steps of: a) obtaining as ground truth data an overall label output and an outputof the Patch-based End-to-End Module having a corresponding label providing information about a risk of breast cancer recurrence of an image a stained tissue sample comprising breast cancer cells, preferably the label is a result of gene expression profiling; b) inputting as example data the overall label output and the output ofthe Patch-based End-to-End Module of stained tissue sample; c) performing the method according to step b) of claim 12 on the inputoverall label output and the input of the output of the Patch-based End-to- End Module of stained tissue sample, d) measuring the error between the generated refined overall labeloutput and the ground truth data; and e) updating the weights of the machine learning data processing modelused for outputting the refined label output for performing the training of the machine learning data processing model, preferably wherein the weights of the machine learning data processing model used for outputting the refined label output are updated using backpropagation, optionally wherein steps b) to e) are repeated until the error is no longer decreasing.

19. A computer program having instructions which when executed by a computing device or system cause the computer device or system to perform the method according to any one of claims 1 – 18.

20. A data-processing system comprising means for carrying out the method according to any one of claims 1 – 18.

Citation Information

Patent Citations

  • Macrocyclic inhibitors of the PD-1 / PD-l1 and CD80(b7-1) / PD-l1 protein / protein interactions

    WO2014151634A1

  • Compounds useful as immunomodulators

    WO2015034820A1

Cited By

  • Method and system for predicting harvesting loss of corn ear harvester based on deep learning

    CN121074689A

  • Textile defect detection method based on improved double backbone architecture network

    CN121391855A