Method for determining a risk score associated with a pathology, and corresponding device and program

The method uses a CNN with 'ConvNext' blocks and ABMIL for melanoma risk scoring, addressing the challenge of analyzing tissue and cellular characteristics in histological images to improve patient stratification and treatment decisions.

WO2026154030A1PCT designated stage Publication Date: 2026-07-23DIADEEP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DIADEEP
Filing Date
2026-01-14
Publication Date
2026-07-23

Smart Images

  • Figure EP2026050860_23072026_PF_FP_ABST
    Figure EP2026050860_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining a risk score (RS) associated with a pathology, the method comprising the following steps: - obtaining (S01) at least one image representative of a tissue section of a tumour; - determining (S02), via a tissue component (CompTS), at least one first vector (v1) representative of tissue characteristics of the tumour; - determining (S03), via a cellular component (CompCC), at least one second vector (v2) representative of cellular characteristics of the tumour; - calculating (S04), by an aggregation component (CompAG), and as a function of the at least one first vector (v1) representative of tissue characteristics of the tumour and of the at least one second vector (v2) representative of cellular characteristics of the tumour, the risk score (RS) associated with the tumour-related pathology.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Title: Method for determining a risk score associated with a pathology, corresponding device and program.

[0003] Scope of the invention

[0004] The invention relates to the field of digital oncology and artificial intelligence applied to pathology. More specifically, the invention concerns the prediction of survival outcomes for patients with primary melanoma based on histological images, for example, hematoxylin and eosin (H&E) images. The invention proposes a technique for implementing neural networks to analyze both the tissue and cellular characteristics of the images in order to generate an overall risk score.

[0005] Previous art

[0006] Cutaneous melanoma represents a significant public health concern. In Europe, for example, an increasing incidence is observed. In 2020, this type of cancer accounted for 4% of all new cancer diagnoses and was responsible for 1.3% of cancer-related deaths, and the trend is rising sharply. This prevalence underscores the importance of understanding the epidemiological trends and survival factors associated with melanoma. The incidence of melanoma in Europe follows a marked North-South gradient, primarily attributable to differences in skin phototypes and genetic predisposition among populations. Thus, Nordic countries such as Denmark, the Netherlands, and Sweden have the highest incidence rates, while Mediterranean countries such as Greece and Portugal have lower rates.

[0007] Regarding survival rates, these vary considerably depending on the stage of the disease at diagnosis. Existing statistics underscore the importance of early diagnosis and prompt treatment. In this context, it is crucial to have a survival / death risk score for melanoma patients. Such a tool allows for more precise patient stratification, thus facilitating personalized and optimized care. It also improves doctor-patient communication by providing more accurate information on individual prognosis and supporting shared decision-making regarding treatment options.

[0008] The determination of risk scores and survival rates for melanomas can currently be segmented into four types of techniques. First, there are molecular and genetic biomarkers. These techniques use genetic profiling tests and molecular biomarkers to assess the risk of recurrence and survival. For example, gene expression profiling (GEP) tests and circulating tumor DNA (ctDNA) are used to provide information on the biological behavior of melanomas and their potential for recurrence.

[0009] A typology can be defined using statistical models and prediction algorithms. These techniques include the application of regression models and other predictive tools to estimate survival rates based on clinical and pathological factors. These models allow for the quantification of mortality risk based on variables such as Breslow thickness, the presence of ulceration, and other clinical characteristics.

[0010] Another type of technique involves histopathological image analysis methods to extract morphological characteristics of tumor tissues from histological images. These characteristics are then used to predict clinical outcomes, often in combination with machine learning algorithms.

[0011] Deep learning and artificial intelligence techniques represent the most recent revolution in this field. The development of convolutional neural networks (CNNs) and other deep learning architectures allows for more in-depth and precise analysis of histological slide images.

[0012] These techniques, however, suffer from several shortcomings, particularly difficulties in effectively determining patient risk scores. Therefore, there is a need to provide better risk scores that meet the expectations of both practitioners and patients.

[0013] Summary of the Invention: The invention addresses at least some of the aforementioned drawbacks. More particularly, the invention relates to a method for determining a risk score associated with a pathology. It is implemented by a determination device comprising at least one processing unit and at least one memory unit, the method comprising the following steps:

[0014] obtaining at least one representative image of a tissue section of a tumor from a patient's tissue;

[0015] determination, by means of a tissue component to which said at least one representative image of a tissue section is provided, of at least one first representative vector of tissue characteristics of the tumor;

[0016] determination, by means of a cellular component to which said at least one representative image of a tissue section is provided, of at least one second representative vector of cellular characteristics of the tumor;

[0017] calculation, by an aggregation component, as a function of said at least a first representative vector of tissue characteristics of the tumor and said at least a second representative vector of cellular characteristics of the tumor, of the risk score associated with the pathology related to the tumor.

[0018] According to a particular characteristic, the tissue component implements a convolutional neural network, more specifically a convolutional neural network comprising at least one "ConvNext" block.

[0019] According to a particular characteristic, the convolutional neural network comprises at least four processing blocks, each processing block comprising at least one "ConvNext" block and each block performing a reduction in the spatial resolution of the image data it receives.

[0020] According to a particular characteristic, the convolutional neural network further comprises at least one fully connected layer delivering said at least one first representative vector of tumor tissue characteristics.

[0021] According to a particular characteristic, the step of determining, by the cellular component, said at least one second vector representative of the cellular characteristics of the tumor includes:

[0022] a cutting step of said at least one representative image of the tissue section into a set of tiles of predetermined size;

[0023] for at least some of the tiles in the tile set, implementation of a neural network using a multiple learning approach with attention, delivering a plurality of feature vectors;

[0024] from the plurality of characteristic vectors, calculate said at least a second vector;

[0025] According to a particular feature, the calculation step, by an aggregation component, based on said at least one first vector representing tissue characteristics of the tumor and said at least one second vector representing cellular characteristics of the tumor, of the risk score associated with the pathology related to the tumor implements a synthetic neural network of the "multi-layer perceptron" type.

[0026] According to a specific characteristic, the risk score calculated by the aggregation component represents the probability of death for the patient within a predefined time range. The risk score can also be a probability of recovery, a probability of survival, and / or a probability of recurrence. Specifically, the nature of the score is determined during component training. Furthermore, this training is implemented through meta-learning, according to the disclosure.

[0027] The risk score is more specifically associated with melanoma.

[0028] In another aspect, the invention also relates to a device for determining a risk score associated with a pathology, said determination device comprising at least one processing unit and at least one memory. Such a device is configured to implement the following steps:

[0029] obtaining at least one representative image of a tissue section of a tumor from a patient's tissue; determining, through a tissue component to which said at least one representative image of a tissue section is provided, at least one first representative vector of tissue characteristics of the tumor;

[0030] determination, by means of a cellular component to which said at least one representative image of a tissue section is provided, of at least one second representative vector of cellular characteristics of the tumor;

[0031] calculation, by an aggregation component, as a function of said at least a first representative vector of tissue characteristics of the tumor and said at least a second representative vector of cellular characteristics of the tumor, of the risk score associated with the pathology related to the tumor.

[0032] According to a preferred implementation, the various steps of the processes according to the proposed technique are implemented by one or more software or computer programs, comprising software instructions intended to be executed by a data processor of a risk scoring module according to the proposed technique and designed to control the execution of the various steps of the processes.

[0033] Consequently, the proposed technique also aims at a program, capable of being executed by a computer or a data processor, this program comprising instructions to control the execution of the steps of a process as mentioned above. This program can use any programming language and be in the form of source code, object code, or intermediate code. The proposed technique also aims at an information storage medium readable by a data processor, and comprising instructions for a program as mentioned above. The information storage medium can be any entity or device capable of storing the program, for example, a storage device, a microelectronic circuit, or a magnetic or optical recording medium.On the other hand, the information medium can be a transmissible medium such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio, or by other means. The program, according to the proposed technique, can in particular be downloaded from a network such as the Internet.

[0034] Alternatively, the information carrier can be an integrated circuit in which the program is embedded, the circuit being adapted to execute or be used in the execution of the process in question. According to one embodiment, the proposed technique is implemented using software and / or hardware components. In this context, the term "module" in this document can refer to a software component, a hardware component, or a set of hardware and software components. Each component of the device or process described above naturally implements its own modules. The various embodiments mentioned above can be combined for the implementation of the proposed technique.

[0035] Brief description of the figures

[0036] Other features and advantages of the invention will become clearer upon reading the following description of an embodiment of the invention, given by way of simple illustrative and non-limiting example, and the accompanying drawings, among which:

[0037] Figure 1 illustrates the architecture of the risk scoring solution according to the invention;

[0038] Figure 2 illustrates a simplified architecture of the tissue component;

[0039] Figure 3 illustrates a simplified architecture of the cellular component;

[0040] Figure 4 illustrates a method for determining a risk score associated with a pathology based on disclosure.

[0041] Description of a method of implementation

[0042] As previously stated, the disclosure relates to a system and method for determining a risk score, for example, in the form of survival outcomes for patients with primary melanoma, using hematoxylin and eosin (H&E) stained histological slide images and deep learning techniques. The proposed technique leverages two distinct levels of analysis: a tissue level and a cellular level. At the tissue level, a first component for determining a feature vector, called the tissue component and comprising a convolutional neural network (CNN), analyzes the overall tumor image to extract macroscopic morphological features.At the cellular level, a second scoring module, called the cellular component, comprises a neural network using Attention-Based Multiple Instance Learning (ABMIL). This module divides the overall image into smaller tiles and analyzes each tile to extract microscopic features related to the cells composing the tumor. The features (tissue and cellular characteristics) generated by these two scoring modules are then combined by a synthesis network to produce an overall risk score, enabling, for example, the stratification of patients according to their mortality risk.

[0043] More specifically, the component architecture in the proposed technique is described in relation to Figure 1. Overall, the system comprises a tissue component (CompTS), a cell component (CompCC), and an aggregation component (CompAG). Under operational conditions, a slide representative of a tumor is selected and scanned (e.g., with 40x zoom) using scanners (e.g., Ventana and Hamamatsu), yielding a whole-slide image (WSI) with a resolution of 0.25 pm / px. This image is examined to eliminate any potential scanning artifacts. It is then provided to the risk scoring system. Potentially, multiple images corresponding to a single patient can be provided sequentially.

[0044] The CompTS tissue component of the invention is configured to analyze the global WSI image of the tumor in order to extract macroscopic morphological features. As previously mentioned, the input data for this component are, for example, images of hematoxylin and eosin (H&E) stained histological slides. This tissue component captures tumor-wide morphological patterns present in its image(s) (when multiple tumor images are used for a patient). To achieve this, a pre-trained convolutional neural network (CNN), such as ConvNeXt, is used within the tissue component to extract features at different levels of granularity. The output data of this tissue component are aggregated feature vectors representing the tissue information of the global WSI image of the tumor.

[0045] More specifically, the inventors determined that, depending on the operational implementation conditions, the "ConvNeXt Small" architecture offers a favorable performance-to-implementation-resources ratio for a suitable deployment. This architecture belongs to the "ConvNeXt" family of convolutional neural networks (CNNs) designed to achieve good performance compared to transformer-based models such as vision transformers (ViTs), while maintaining the simplicity and efficiency of CNN architectures. "ConvNeXt" builds upon the fundamental principles of traditional CNNs and incorporates design choices inspired by ViTs. In the implementation example, the inventors used a scaled-down version designed to achieve a balance between accuracy and computational efficiency. "ConvNeXt Small" incorporates several features that enhance its efficiency.It uses depth-separable convolutions, which significantly reduce the number of parameters and computations required compared to standard convolutions, while preserving accuracy. Pre-normalization with LayerNorm, applied before the convolutions and inspired by vision transformers, further improves optimization by addressing vanishing gradient issues, particularly in deeper networks. The model also minimizes nonlinearity by reducing the use of ReLU activations, enabling faster inference and smoother gradients. ConvNeXt blocks are streamlined to eliminate parameter redundancy, maintaining high representational power without unnecessary complexity. The hierarchical design (in which spatial dimensions are progressively reduced) optimizes memory usage and computational load.Furthermore, larger convolutional kernels, such as 7x7, enable efficient spatial aggregation without requiring additional layers, thus allowing for the effective capture of broader contextual information. Therefore, "ConvNeXt Small" leverages the design principles of "ViT" (Virtual Image Technology), such as hierarchical structure and attention-inspired normalization techniques, while utilizing convolutional layers for feature extraction. This is well-suited and advantageous for the present implementation, particularly on systems with limited computing resources.

[0046] In an example implementation, illustrated in Figure 2, the architecture comprises four (4) stages (Stg1 to Stg4), also called processing blocks, with the output of one processing block being used for the next. Each stage constitutes a block comprising three (3) "Convnext" blocks. The first stage, "Stg1", includes an input component, "Stm". The following three stages include, before the three "Convnext" blocks, a reduction block, "DwnSmp".

[0047] More specifically, the layers of the model are as follows:

[0048] Stem block: The initial WTI image is reduced to a 512x512 pixel input image. This input image is first processed by a convolutional layer with a 4x4 kernel and a step size of 4. This operation acts as a patching step, analogous to tokenization in vision transformers (ViTs). The result is a set of low-resolution feature maps, reducing the input spatial dimensions to 128x128. These feature maps constitute the input for the hierarchical feature extraction steps.

[0049] "Stage-wise Hierarchical Design": ConvNeXt uses a hierarchical design divided into four stages, where each stage processes features at progressively coarser spatial resolutions while increasing channel depth. This allows the model to capture both fine details and high-level semantics.

[0050] ■ (stg1) Stage 1: The spatial resolution remains at 128x128 (feature maps - output images from the stem block) with a relatively small number of feature channels. This stage consists of several "ConvNeXt" blocks, designed to process shallow features at their original resolution.

[0051] ■ (stg2) Stage 2: The spatial resolution is halved to 64x64 using a subsampling layer (a 2x2 convolution with a step size of 2). The number of feature channels is doubled compared to stage 1, allowing for richer feature representations.

[0052] ■ (stg3) Stage 3: The spatial resolution is further reduced to 32x32 by another subsampling layer, and the channel depth is doubled again. This stage allows for the capture of mid-level features with a balance between spatial and channel granularity.

[0053] ■ (stg4) Stage 4: The spatial resolution is halved one last time to reach 16x16, and the channel depth is doubled again. This stage allows for the extraction of high-level abstract features suitable for classification tasks.

[0054] Final integration: After processing the input through hierarchical stages (each processing block (Stg1, ..., Stg4)), the model reduces the spatial dimensions of the final feature maps using global mean pooling, compressing each feature map into a single value. The pooled features then pass through a fully connected layer designed to predict survival outcomes. Instead of applying softmax activation for classification, the model produces a hazard score or a survival probability, depending on the specific survival analysis framework. For example, in a Cox proportional hazards framework, the output is a continuous hazard score used to estimate relative hazard, while for survival probability estimation, the output represents the probability of survival at specific time intervals.This adaptation ensures that the model is suited to survival analysis rather than classification.

[0055] According to this document, in this architecture of this tissue component, each ConvNeXt block is based on the foundations of ResNet-type blocks, while introducing several improvements for better efficiency and performance: Deep convolution: Replaces traditional convolutions, operating independently within each feature channel, which reduces computational complexity while preserving spatial information.

[0056] Point convolution (1x1): Re-projects the convolution output in the depth direction to the desired dimensionality, ensuring that representational power is maintained.

[0057] Layer Normalization (Pre-Norm): Applies normalization after convolution in the depth direction rather than BatchNorm, as inspired by Vision Transformers, a choice that improves training stability and convergence. GELU Activation: Replaces ReLU with GELU ("Gaussian Error Linear Unit"), a smoother activation function that improves optimization dynamics;

[0058] Residual connections: Incorporates "skip" connections for stable training and improved gradient flow;

[0059] Inverted bottleneck design: Extends feature dimensions into hidden layers, allowing the model to learn from complex feature interactions without significantly increasing computation.

[0060] The CompCC Cellular component is specifically designed for analyzing microscopic features by slicing the global image into smaller tiles. Figure 3 illustrates the architecture of the CompCC Cellular component. The input data for this component are the same WSI (IMG) images, which are then segmented into smaller tiles (e.g., 512x512) by a TSC tiling subcomponent. The cellular component captures morphological details at the cellular level and within the tumor microenvironment. A neural network using an attentional multiple learning (AB MIL) approach employs feature vectors representing the tiles from a ViT encoder (ViT E, also known as a feature extractor). The ViT encoder (ViT E) is used to encode each tile into feature vectors (WSI TiRe).The output data from this feature extractor (ViT E) are "aggregated" feature vectors representing the cellular information of each tile in the image.

[0061] Multiple Instance Learning (MIL) is a mechanism in which data is structured into bags containing multiple instances, with a single label assigned to the entire bag (i.e., all the features it contains). Attention-Based MIL (AB MIL) is an extension of the classical MIL mechanism that incorporates an attention mechanism to dynamically weight the contribution of individual instances. The principle of attention-based MIL is to focus on the most relevant instances through a trainable mechanism. This allows the model to assign importance to specific instances based on their relevance to the overall bag label, making it more flexible and easier to interpret.

[0062] Thus, within the framework of this technique, the WSI image is segmented, primarily due to its volumetric nature, which facilitates the implementation of ABMIL. By dividing the WSI into smaller, fixed-size tiles, the computational load becomes manageable, allowing the analysis to be focused on smaller elements (imagelets), either sequentially or in parallel. Unlike the "tissue component," tiling ensures that the model captures localized, fine-grained tissue features. The feature vectors from the encoder (Vit E) are provided to the neural network (AB MIL), which, in a simplified manner, groups and aggregates the contents of these feature vectors into "bags," which are then "weighted" via an attention mechanism.

[0063] In an example implementation, the ABMIL implementation includes modifications to the original model and can be described as follows.

[0064] Input for feature transformation: The raw instances in the bag are transformed into a more compact and expressive feature space. Each instance with an initial dimensionality passes through a linear layer followed by ReLU activation to map the input into a new dimensionality feature space. This step captures meaningful patterns in the data while reducing redundancy. Dual attention mechanism: This mechanism assigns importance to each instance by calculating attention scores through two different pathways: one pathway applies a "tanh" activation to emphasize content-focused features, while the other uses a "sigmoid" activation to emphasize relevance. These two outputs are combined per element, and the resulting values ​​are processed by a linear layer to calculate the raw attention scores.A softmax function then normalizes these scores, assigning probabilistic weights to instances based on their importance for the bag-level prediction.

[0065] Instance aggregation: The model uses calculated attention weights to aggregate instance features into a single bag-level representation. This aggregation is achieved by taking a weighted sum of instances, where the attention weights determine each instance's contribution. This step ensures that only the most relevant bag features are highlighted in the final representation.

[0066] “Prediction Head”: it takes the aggregated representation of the bag and passes it through a final linear layer to produce the prediction at the bag level of the patient’s risk score.

[0067] The model provides additional interpretive possibilities by returning attention weights along with the predictions. These weights indicate the relative importance of each instance in the bag, revealing which parts of the data contributed most to the model's decision. This allows for subsequent cross-checking, for example by a practitioner, helping to understand which regions of interest in a WSI image guided the scoring.

[0068] The initialization of the CompCC cellular component model is described. Specifically, prior to the end-to-end development phase, the ABMIL model is initialized and trained in a structured manner to ensure optimal learning. This process involves using a pre-trained feature extractor combined with a learning strategy designed for the ABMIL model. The pipeline begins by leveraging the pre-trained Vision Transformer (ViT) feature extractor, which, in one implementation, was trained using the DINO self-supervised learning framework. This ViT feature extractor processes histopathological image data, extracting high-level, meaningful features from the image tiles.

[0069] The initialized cellular component (including the ABMIL model) is then trained on these extracted features rather than raw image data using a meta-learner described below. During this phase, the ViT E encoder remains static, ensuring that its weights are not updated. The ABMIL model learns to classify the pathology slides by assigning attention scores to instances (image patches) and aggregating them into a bag-level representation. This initial training phase provides the ABMIL model with weights that reflect its ability to handle the features extracted by the ViT E encoder.

[0070] The use of a pre-trained feature extractor and an attention-based multi-instance learning (ABMIL) model for histopathological classification involves a pipeline to process histopathological image data, extract significant features, and train the ABMIL model to classify pathological slides. Consideration must be given to the dimensions and number of parameters required to refine all models within this pipeline. Indeed, the ViT E feature extractor uses a Vision transformer pre-trained with DINO. This model normally requires a powerful computer to run. The inventors decided to reduce the number of parameters to be trained so that the model can be adapted to the computing resources of an RTX® 30 or 40 series GPU.The inventors refined the ensemble weights (encoder and classifier) ​​using a low-level adaptation for large models (LoRA) technique. This technique efficiently trains large models by focusing on smaller trainable matrices, which are a low-level decomposition of the delta weight matrix learned during fine-tuning (typically in attention blocks). The original weight matrix of the pre-trained model is fixed, and only the smaller matrices are updated during training. This reduces the number of trainable parameters, thus decreasing memory and computation time. As previously mentioned, to refine the model, they start with ABMIL weights obtained through initial training on features extracted by the ViT E feature extractor.

[0071] The CompAG aggregation component merges features extracted by the tissue and cellular components to produce an overall risk score. The input data for this component are the aggregated feature vectors from the tissue component CompTS and the cellular component CompCC. A synthesis network, such as a multi-layer perceptron (MLP), is used to concatenate these feature vectors and learn the complex relationships between them. The goal is to generate an overall survival risk score for each patient. The output data of this component are the overall risk scores. The risk score (RS) calculated by the aggregation component is, for example, the probability of death for that patient within a predefined time range. It can also be the probability of survival for the patient or the probability of complications related to the patient's pathology.The neural network(s) of the CompAG aggregation component are trained to provide this data, notably through "meta-learning" (see below).

[0072] In other words, the system of the invention takes the form of a "fusion" pipeline comprising neural networks operating at different scales, each designed to extract features at a distinct level and granularity. The networks are specifically trained to detect the tissue and cellular morphology inherent to each WSI resulting from genomic alterations. The Cellular component, CompCC, captures morphological patterns that may exist at the cellular level and in the cellular microenvironment without distinguishing between cancer cells and other cells. The Tissue component, CompTS, focuses on the patient-associated tumor tissue and aims to capture tumor morphology.

[0073] As mentioned previously, to maximize the extraction of important features from WSI images, a specific learning mechanism is implemented. The goal of this learning mechanism, called "meta-learning" (or "Meta-Learner"), is to correctly configure the neural networks of the cellular and tissue components. The Meta-Learner uses a specific mechanism to train these two components on histopathological images while taking into account the inherent variability of real-world data.

[0074] Prior to implementing this mechanism, however, the amount of training data is increased, according to the present document, to obtain an augmented dataset (eWSI+). This augmentation aims to address artifacts and color heterogeneity caused by differences in pre-analytical processes from one laboratory to another, among other things.

[0075] Thus, to account for discoloration artifacts and color heterogeneity between different cohorts, which are inherent to the different and variable pre-analytical processes of stained H&E / HES slides, the preprocessing step considers the main "styles" of base images encountered under operational conditions. To account for these different styles (and this variability) of base images, the inventors determined it was important to artificially generate an augmented dataset that simulates the variability in staining, texture, and tissue thickness due to disparities in pre-analytical processes between histopathology laboratories. To do this, the inventors applied numerous data augmentation techniques to the original images, ranging from modifying the appearance of the stains (color distortions, saturation, etc.).) to the distortion of brightness and intensity. Secondly, the inventors generated an augmented dataset designed to reproduce disparities using a color shift intended to reproduce image quality and / or coloration likely to be observed from one scanner to another. Based on the original images, these augmentations are random. For example, the following modifications can be implemented:

[0076] gamma increase: adjusting gamma values ​​to simulate differences in color intensity and brightness between laboratories and scanners, which modifies the pixel intensity distribution while preserving relative contrast; "Color Jittering": introducing random changes in hue, saturation, brightness, and contrast to reproduce the variability of color protocols and scanner settings;

[0077] blend: mixing pairs of images with a weighted sum to create new samples, which may represent intermediate staining conditions or overlapping tissue appearances;

[0078] elastic transformations: application of random spatial distortions to simulate the variability of tissue texture and slide preparation, such as folding or stretching effects;

[0079] Gaussian noise injection: adding random noise to simulate imaging artifacts or scanner inconsistencies;

[0080] histogram equalization: modification of the histogram of image intensity values ​​to balance differences in scanner calibration;

[0081] task standardization: standardization of the appearance of the staining of the slides using reference models in order to ensure uniformity while preserving diagnostic characteristics;

[0082] rotation and flipping: random rotation and flipping of images to handle the variability of orientation in the dataset;

[0083] Random cutting and erasure: simulation of occlusions or missing tissue sections by masking random regions of the image.

[0084] These modifications produce training images that are added to the original images. In this sense, the initial dataset is augmented.

[0085] In addition to these "default" enhancements, the inventors also implemented color mixing from two (or more) blended images to generate new training samples. This method introduces variability in color profiles, allowing the models to become more robust to differences in staining protocols, scanner types, and other color domain changes encountered in histopathological images. This procedure mimics the variations in intensity, hue, and saturation of stains between different laboratories and scanners and aims to introduce synthetic diversity, thereby improving the models' ability to adapt to a variety of staining styles. Furthermore, this allows the model to focus on structural and cellular features rather than adapting to specific staining patterns.

[0086] Using this augmented dataset (eWSI+), training is performed on all networks of both components (cellular and tissue). Among the training parameters for these networks, we detail those that enable us to achieve an effective result using this technique, notably the performance metric and the loss function. Each subject "i" is represented by an image xi, and an observation period (until death or loss of follow-up) yi for the subject, as well as a survival indicator d(i), which is equal to 1 when the subject dies within the observation period. The objective of the model is to learn to associate morphological characteristics of the image (tissue or cellular) with the individual's survival time without death or relapse. Let u(xi) denote the prediction given by the model.As for the loss function, the random function generally depends on both time and a set of exogenous variables (outputs of the "deep learning" model, denoted x). Proportional randomness models separate the two effects, and we have:

[0087] h(y,x) = g(y)exp(u(x)),

[0088] where u(xi) is the network score for the image xi. Here, g(y) is considered a nuisance parameter, and we are interested in estimating the effect of the image represented by the prediction x given by the model. No information can be deduced from the time intervals during which no deaths occur, since the hazard function can be zero. However, we can reason conditionally on the set of ordered observed times of death t(l),...,t(k). For an observed time of death t(i), let R(t(i)) be the set of subjects at risk and d(i) the indicator of the event. For the observed times yl, ..., yn, the negative conditional Cox log-likelihood is: l(yl,...,yn) = SP = id(i)u(xi) + d(i) R(yi) exp(u( xi))

[0089] Since the set R(yi) can be very large, this loss function is calculated only for observations from a batch (of patients). To obtain a good approximation of this loss function, a "large" batch is used (typically 64 patients).

[0090] For loss function optimization, an "Once-Cycle" learning rate is used, with a maximum learning rate of 2e-4. This rate adjusts during training based on an "Once-Cycle" policy to improve convergence efficiency. Training is divided into two main phases: warm-up and cool-down. During the warm-up phase, the learning rate starts at a low initial value and gradually increases to the specified maximum learning rate of 2e-4. This gradual increase stabilizes the optimization process in its early stages and promotes efficient exploration of the parameter space. Subsequently, the cool-down phase gradually reduces the learning rate to a very low value.This step allows the optimizer to refine the model parameters through smaller, more precise updates, ensuring a smoother convergence to the optimal solution. Using the One-Cycle policy, combined with an Adam optimizer, enables faster convergence and robust model learning while maintaining computational efficiency.

[0091] The performance metric used to select the best models is the C-index. The function u(xi) can be considered a risk score; the C-index is calculated as follows:

[0092] if yi and yj are observed, the pair (i, j) is said to be concordant if u( xi) > u( xj) and yi < yj, otherwise it is discordant;

[0093] if yi is observed and yj is censored (patient lost to follow-up):

[0094] if yj < yi we cannot compare the time of the final event of subject j with yi and the pair is not taken into account.

[0095] if yj > yi, necessarily, the time of the event of the subject will be greater than yi and the pair is concordant if u(xi) > u(xj), otherwise it is discordant.

[0096] The index C is then:

[0097] number of concordant pairs

[0098]

[0099] number of pairs considered

[0100] Still related to the "meta-learner," a cross-validation and cross-testing (CV-CT) technique is implemented. To improve the model's performance and assess its robustness, a cross-validation technique is employed. This technique creates, at each iteration, a cohort of synthetic tests that are not seen by the model at each training fold. This cross-validation and testing approach aims to identify the model that minimizes the loss function. The test set is used to evaluate the model's performance. Instead of maintaining a complete separation of the test set, a modified training set is iteratively introduced for each fold.At each iteration, the training process involves combining the original training data with a portion of the initially separate test data, while the remaining test data is reserved for evaluating network performance, thus enabling testing on the entire population. In the case of a five-fold CV-CT, this technique includes the following steps:

[0101] 1. The dataset is divided into five folds using the stratified group "k-fold" technique, ensuring that the same patients remain together in the same fold and that patients whose data are censored are distributed equally between the different folds;

[0102] 2. Each of the five folds is then used as a test set, while the four remaining folds collectively constitute the training set;

[0103] 3. In the training set, we train four models using a cross-validation approach; Each of the four folds is used as a validation set, the other folds serving as training sets;

[0104] 4. An ensemble model combines the risk scores by averaging the scores obtained from the four models. Finally, based on the above, it remains to determine which of the models created from the provided data is the most relevant. An ensemble approach is used. More specifically, an ensemble approach is used to aggregate the scores of each model for CV-CT in both components (tissue / cellular). It is appropriate for the Cox proportional hazards (CP) model because it leverages the additive nature of the CP assumption, ensuring that the ensemble averaging process preserves the fundamental structure of the model. The CP framework models the hazard function as the product of a base hazard and an exponential term based on the covariates.

[0105] h(y,x) = g(y)exp(u( x)),

[0106] where u(x) represents the log-risk score for the individual with the WSI image x. This structure allows for the efficient combination of several models with different parameters. A model weight can optionally be assigned when one model performs better than another.

[0107] It is also possible to perform a set-based combination of models. When averaging several hazard functions from different models within the set, the resulting hazard function retains the proportional hazards property. Mathematically, the harmonic mean of the hazard functions leads to a combined model in which the base hazard is averaged across the models, and the logarithm of the hazard u(x) is averaged over the set. This ensures that the assumption of proportional hazards remains valid after calculating the set average. The set-based approach is particularly beneficial when the models are trained using stochastic methods, such as mini-batch stochastic gradient descent (SGD) neural networks.The variability in training due to random initialization, data augmentation, and sampling can lead to models that capture different aspects of the data distribution. Averaging these models mitigates overfitting, reduces variance, and improves the generalizability of the predictions. This is important in survival analysis, where censored data and small sample sizes can make individual models susceptible to overfitting.

[0108] Within the framework of validation and cross-testing (CV-CT), the ensemble method also incorporates predictions from models trained on different subsets of data. This ensures greater robustness by capturing diverse models and reducing dependence on specific training folds. The ability to combine predictions while adhering to the CP framework makes the ensemble approach an ideal solution for improving the accuracy and consistency of survival predictions in this context.

[0109] Figure 4 illustrates the process for determining a risk score associated with a pathology, the process comprising the following steps:

[0110] obtaining (S01) at least one representative image of a tissue section of a tumor of a patient's tissue;

[0111] determination (S02), by means of a tissue component (CompTS) to which said at least one representative image of a tissue section is provided, of at least one first vector (v1) representative of tissue characteristics of the tumor; determination (S03), by means of a cellular component (CompCC) to which said at least one representative image of a tissue section is provided, of at least one second vector (v2) representative of cellular characteristics of the tumor; calculation (S04), by means of an aggregation component (CompAG), as a function of said at least one first vector (v1) representative of tissue characteristics of the tumor and said at least one second vector (v2) representative of cellular characteristics of the tumor, of the risk score (RS) associated with the pathology related to the tumor.

[0112] The method is implemented by an electronic selection device capable of performing all or part of the processing as described above. An electronic selection device comprises a first electronic module including memory, a processing unit equipped, for example, with a microprocessor, and controlled by a computer program. In at least one embodiment, the present technique is implemented as a set of programs installed partially or entirely on the electronic selection device. In at least one other embodiment, the present technique is implemented as a dedicated component capable of processing data from the processing units and installed partially or entirely on the selection device.Furthermore, the system also includes communication means, such as network components (Wi-Fi, 3G / 4G / 5G, wired, RFID / NFC, Bluetooth, BLE, LPWAN, VLC, etc.), which enable the system to receive data from entities connected to one or more communication networks and to transmit processed data to such entities. Such a determination system also includes parallelizable computing means (such as graphics processing units), which are used to perform risk scoring calculations, as described above.

[0113] In preferred embodiments, the computer system or device comprises one or more processors (which may belong to the same computer or to different computers) and one or more memories (magnetic hard drive, optical disc, electronic memory, or any computer-readable storage medium) in which a computer program product is stored, in the form of a set of program code instructions to be executed in order to implement all or part of the steps of the control method. Alternatively, or in combination, the computer system, device, or module may comprise one or more programmable logic devices (FPGAs, PLDs, etc.), and / or one or more specialized integrated circuits (ASICs), etc., adapted to implement all or part of said steps of the control method.In other words, the computer system comprises a set of means configured by software (specific computer program product) and / or by hardware (processor, FPGA, PLD, ASIC, etc.) to implement the steps of the selection method.

Claims

DEMANDS 1. A method for determining a risk score (RS) associated with a pathology, implemented by an electronic device comprising at least one processing unit and at least one memory, the method comprising the following steps: obtaining (S01) at least one overall image representative of a tissue section of a tumor of a patient's tissue; determination (S02), via a tissue component (CompTS) to which said at least one global image representative of a tissue section is provided, of at least one first vector (v1) representative of tissue characteristics of the tumor; determination (S03), via a cellular component (CompCC) to which said at least one global image representative of a tissue section cut into a set of tiles is provided, of at least one second vector (v2) representative of cellular characteristics of the tumor; calculation (S04), by an aggregation component (CompAG), as a function of said at least a first vector (v1) representative of tissue characteristics of the tumor and said at least a second vector (v2) representative of cellular characteristics of the tumor, of the risk score (RS) associated with the tumor-related pathology.

2. A determination method according to claim 1, characterized in that the tissue component (CompTS) implements a convolutional neural network (CNN), more particularly a convolutional neural network comprising at least one block “ConvNext”.

3. Method of determination according to claim 2, characterized in that the convolutional neural network (CNN) comprises at least four processing blocks (Stg1,... Stg4), each processing block (Stg1,... Stg4) comprising at least one "ConvNext" block and each block performing a reduction of the spatial resolution of image data which it receives.

4. Method of determination according to claim 2, characterized in that the convolutional neural network (CNN) further comprises at least one fully connected layer delivering said at least one first vector (v1) representative of tumor tissue characteristics.

5. Method for determining according to claim 1, characterized in that the determination step (S03), by the cellular component (CompCC), of said at least a second vector (v2) representative of cellular characteristics of the tumor comprises: a step of cutting said at least one representative image of the tissue section (IMG) into a set of tiles of predetermined size; for at least some of the tiles in the tile set, implementation of a neural network using an attentional multiple learning approach (ABMIL), delivering a plurality of feature vectors; from the plurality of characteristic vectors, calculation of said at least a second vector (v2); 6. A method for determining the risk score (RS) associated with the tumor-related pathology according to claim 1, characterized in that the calculation step (S04), by an aggregation component (CompAG), based on said a multi-layer perceptron-type synthetic neural network, implements a risk score (RS) associated with the tumor-related pathology based on at least a first vector (v1) representative of tissue characteristics of the tumor and said at least a second vector (v2) representative of cellular characteristics of the tumor.

7. A method for determining the risk score (RS) calculated by the aggregation component is representative of the probability of death of said patient within a predefined time range.

8. Method of determination according to claim 1, characterized in that the pathology associated with the risk score is melanoma.

9. Device for determining a risk score associated with a pathology, said determination device comprising at least one computing unit and at least one memory, the device being configured to implement the following steps: obtaining at least one overall representative image of a tissue section of a tumor from a patient's tissue; determination, by means of a tissue component to which said at least one global image representative of a tissue section is provided, of at least one first vector representative of tissue characteristics of the tumor; determination, by means of a cellular component to which said at least one global image representative of a tissue section cut into a set of tiles is provided, of at least one second vector representative of cellular characteristics of the tumor; calculation, by an aggregation component, as a function of said at least a first representative vector of tissue characteristics of the tumor and said at least a second representative vector of cellular characteristics of the tumor, of the risk score associated with the pathology related to the tumor.

10. Computer program product, characterized in that it includes program code instructions for implementing the steps of the prediction process according to at least one of claims 1 to 8.