Image Analysis System and Method Using a Neural Network Adapted to Continuous Learning

Möbius transformation and weighted distillation enhance incremental learning in medical imaging by reducing forgetting and improving accuracy on new classes, addressing the challenge of limited data availability and storage constraints.

JP2025524547APending Publication Date: 2025-07-30UCB BIOPHARMA SPRL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577255
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-11
Filing Date
2023-07-06
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing deep learning models in medical imaging struggle with catastrophic forgetting when incrementally learning new classes due to limited data availability and storage constraints, necessitating a robust method for continuous adaptation without memorizing past classes.

Method used

The method employs Möbius transformation for incremental learning, combining it with weighted cross-distillation to expand sample diversity and maintain past knowledge, using a limited number of past task samples, and optimizing new class learning with weighted logits.

Benefits of technology

This approach effectively reduces catastrophic forgetting by maintaining past class performance while improving accuracy on new classes, enhancing model robustness and generalization in incremental learning scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524547000001_ABST
    Figure 2025524547000001_ABST
Patent Text Reader

Abstract

The present invention provides a novel approach to incremental neural network training, in which retaining a limited number of examples of old classes serves to mitigate forgetting by using a strategy of incremental time expansion with Möbius transformation and weighted distillation instead of large-scale data storage to correct for the effects of evolving class imbalance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system and method for training and generating a deep neural network for temporally spaced arrival of non-related information using an extension based on fractional linear transformation as part of a process.

Background Art

[0002] Deep learning-based methods have become common in medical imaging research. In a realistic scenario, clinical imaging systems often do not initially have access to all the required data, and the data arrives incrementally over time, in fragmented form across multiple devices and various centers. This problem is prominent in healthcare systems in low- and middle-income countries (LMICs) where the data acquisition and quality assurance infrastructure may not be as developed (Becker et al., Tropical Medicine & International Health 21.3 (2016), pp. 294 - 311). In cases regarding the accessibility of such variable data, in order to maintain clinical significance throughout their retention period under evolving requirements and reliably support diagnostic efforts, machine learning algorithms need to be robust to the temporal adaptation to new data distributions and generalizable to new classes of data. This requirement for continuous adaptation in deep networks for clinical imaging implies the need to ensure that model parameters maintain relevance for both old and new tasks in an incremental data regime. This has to be done without memorizing a large number of examples from past classes for subsequent learning schedules due to constraints on the long-term storage of clinical data regarding memory, legal, and privacy issues (see, for example, FDA et al., Proposed regulatory framework for modifications to artificial intelligence / machine learning (AI / ML)-based software as a medical device (SaMD)-discussion paper. 2019). Therefore, the ideal co-training conditions of optimizing the model using all the data sets that have ever been used in each incremental retraining are difficult in medical imaging.

[0003] The adaptation of existing models to learn new classes has been attempted by transfer learning (Ravishankar et al., Deep Learning and Data Labeling for Medical Applications, pp. 188 - 196 (2016)). Transfer learning is useful in that a prior learning episode helps to enhance future task learning, but it has been found that the way of balancing the knowledge of old tasks and new tasks in the ultimately available model is inefficient. The findings show a decline or catastrophic forgetting in past performance (Goodfellow et al., arXiv:1312.6211 (2013)), which is because the previously learned information is lost, causing a high validation loss for past data. Recent research has pursued alleviating forgetting in deep networks using parameter expansion (Rusu et al., arXiv:1606.04671 (2016)), instance replay (Li et al., IEEE Transactions on Pattern Analysis and Machine Intelligence 40(12), pp. 2935 - 2947 (2017)), generative rehearsal (Kemker et al., arXiv:1711.10563 (2017)), and weight regularization (Kirkpatrick et al., Proceedings of the national academy of sciences, pp. 3521 - 3526 (2017)). Knowledge distillation, where the representations learned by one model are transferred to another model, is often used in model compression (Hinton et al., NIPS 2014 Deep Learning Workshop (2014)). This is used in incremental learning because the representations from one learning session help to regularize future sessions, and the logits of old tasks can regularize the learning for new data.Such methods include learning without forgetting (LwF) with distillation and cross-entropy objective functions, iCaRL that incrementally learns representations, learning without memorizing (LwM) where distillation and class activation are continuously performed, and progressive retrospection (PDR) that uses distillation from both old and new models.

[0004] In medical imaging, data availability is often not immediate, and models that incrementally learn over time without affecting past performance, such as pixel regularization for MRI segmentation, modeling Alzheimer's progression, weight integration and distillation, hierarchical continual learning, etc., are being studied. Data augmentation is widely used in machine learning, but there has been relatively little research on runtime augmentation for examples retained in incremental learning.

[0005] The extended approach has been used as a normal process during the training of deep neural networks. This is also the case for Möbius extensions as first shown in arXiv:2002.02917 by SHARON ZHOU et al. The utilization of such extension methods has been difficult and insufficient in related literature such as Fei et al., IEEE / CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, June 20, 2021, pages 5867 - 5876. Different from the training of normal deep neural networks, the adoption of extension methods in incremental and continuous learning is technically difficult, especially in a regime where data availability is limited (a scenario specifically targeted by this disclosure and not considered in the prior art). Incorporating fractional linear transformation into the training process when data sets arrive at incremental time intervals is a non - trivial inclusion from the perspective of the catastrophic forgetting phenomenon emphasized in the prior art, and thus an approach that has not been explored so far. The prior art in this field relies on simple incremental distillation techniques or geometric transformations for the stored examples rather than mathematical transformations of the input space.

Prior Art Documents

Non - Patent Documents

[0006]

Non - Patent Document 1

Non - Patent Document 2

Non - Patent Document 3

[0007] The present disclosure relates to the design of neural networks adapted to learning for the temporally - spaced arrival of relevant data in a process known as "continuous learning", "lifelong learning", or "incremental learning". The present disclosure provides a system and method for training and generalizing deep neural networks for the temporally - spaced arrival of irrelevant information, using an extension based on fractional - linear transformation as part of the process.

Means for Solving the Problem

[0008] The present invention provides a novel approach for incremental training of a neural network that does not memorize a large number of samples in image analysis. This approach is illustrated using a dataset of colorectal cancer images for the purpose of demonstrating proof of concept. This is achieved by expanding sample diversity through a novel online expansion for a limited number of samples of past tasks, while training on new class data for the available model and performing weighted cross-distillation on the logits of past classes. The main differences from existing approaches are: a) the concept of an incremental time data augmentation strategy using Möbius transformation, b) weighted cross-distillation for continuous learning of new classes, and c) online adaptation of Möbius expansion in incremental learning tasks.

[0009] The present disclosure provides a method of incremental learning that relies on Möbius transformation and its interpretation as a composition of basic operations such as translation, rotation, etc., each of which forms the basis of a data augmentation method at the sample level. This method further uses a distillation approach where class-specific accuracy is used for the distribution of importance over the entire vector of past logits. The combination of the two steps involves performing online expansion using Möbius transformation on a few samples from old classes, which improves the representation of previously seen classes as the model is optimized for new classes. During optimization for new classes, the overall representation of the old task is weighted to reflect it, and the summed logits of the old classes, in addition to cross-entropy optimization that enables new task learning considering both the knowledge of old classes and the sample information of new classes, allows the model to have a snapshot of past learning and prevents catastrophic perturbations to the parameter space.

[0010] The present invention will be described below with reference to the following drawings.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

[0012] Problems to be Addressed The prior art does not show the integration of mathematical transformations of input data as a basis for effective knowledge transfer under conditions where data availability during incremental learning is limited (in itself, this training methodology and the associated system / method are clearly different from such standard processes of machine learning in that the standard process of learning from data is a static process that does not take into account the dynamic time-interval arrival of new information that incremental learning is attempting to achieve).

[0013] In an incremental learning scenario, particularly in situations where data availability is restricted for clinical / pre-clinical tissue structure data, the technical problem addressed by the present disclosure to achieve achievable incremental learning performance is solved by various proposed methodological approaches, specifically, by the use of Möbius transformation on the input space. The use of the distillation process itself was known from the prior art, but such a process alone does not enable effective generalization for the incremental regime when the data availability of imaging data (such as tissue structure data) is considered restricted. Such a function is obtained by the integration of a mathematical transformation of the input space at incremental times, as shown by the present disclosure. The present disclosure realizes for the first time the use of such a transformation in an incremental learning process and shows the resulting effect on the input space of the data set.

[0014] The overall pipeline of continuous learning provided by the present disclosure addresses a two - element technical problem of dynamically learning with respect to the time - spaced arrival of data sets and enabling such incremental learning for a relatively small number of cases representing such dynamic arrivals of data classes in a clinical / pre - clinical tissue structure context (few - shot learning). In the prior art, the use of mathematical transformations on the data input space has only been explored as a data augmentation strategy in static machine learning methods. The present disclosure provides a way to integrate such transformations into a dynamic incremental learning context as a sub - division of the overall continuous learning procedure centered around distillation. The present disclosure provides an approach that has not been discussed or proposed in the prior art by enabling such a continuous learning procedure in cases where data availability is restricted or in situations representing the complexity corresponding to tissue structure data sets.

[0015] The model needs to be trained in an M - stage manner, where each stage is \(X_t=\{X_{t,i}\}\) KtConsider a classification task with classes for \(i = 1, t\in[1,M]\), where each \(X\) is a class containing samples \(x_t\in X_t\), and \(K_t\) is the number of classes at each stage \(t\). The classifier learning at stage \(t - 1\) should not show a significant decline in the inference ability for the validation set instances from the \((t - 1)\) - th or earlier stages after being incrementally optimized for the classes at the \(t\) - th stage. Here, the inventors have designed an incremental learning experiment with 4 classes in the initial training stage and 4 classes in the incremental stage (\(M = 2,K_1=K_2 = 4\)).

[0016] This investigation is modeled as the continuous class learning task described above. In the initial training stage, a proportion of classes are learned as "base classes". Next, the remaining classes are learned as "incremental classes" in subsequent learning stages, resulting in a multi - stage learning system over a certain time interval. The former is used to optimize for the initial task (Task 1), and the latter helps in training the model trained on the base classes for the incremental task (Task 2), thus simulating a continuous learning scenario.

[0017] Möbius extension Many sample - level data augmentation methods during training belong to the set of affine transformations, which include groups of mappings such as rotation, scaling, translation, and inversion. Such operations can be modeled as bijective mappings on the complex plane as \(z\rightarrow az + b\), where the variables \(z\), parameters \(a,b\in\mathbb{C}\) (the set of complex numbers).

[0018] The generalization of this mapping takes into account the presence of a non-zero imaginary part of the complex number in the transformation and the fact that the affine mapping is performed on the Argand plane (Oezdemir et al., Communications in Nonlinear Science and Numerical Simulation, 16(12):4698 - 4703, 2011). This expands the upper set of possible image transformations while maintaining valid labels. The denominator of the linear transformation z → az + b can be assumed to be 1. This can also be obtained by treating the denominator as a complex number cz + d and making the real part of this complex quantity 1 and the imaginary part 0. This provides a clue in the next stage of abstraction when introducing a denominator (c, z ≠ 0) with non-zero real and imaginary components. This creates a group of transformations in the set of complex numbers. f(z)=(az + b) / (cz + d) (1) Here, a, b, c, d ∈ C and ad - bc ≠ 0 are the invertibility conditions.

[0019] This includes a superset of basic mappings including inverse transformation, translation, rotation, and inversion, called a Möbius transformation

[19] when \(z\in\mathbb{C}\), \(f(z)\) is not constant, and \(cz + d\neq0\). The point \(z\) is mapped from one complex plane to another using the parameters \(a\), \(b\), \(c\), and \(d\). Since all real numbers can be in the form \(x + iy\) where \(x\in\mathbb{R}\) and \(y = 0\), this can proceed without an explicit imaginary part defined for the complex entity \(z\). This enables defining points on the image to estimate \(a\), \(b\), \(c\), and \(d\). Three points are randomly selected on the image space using different combinations, allowing different outputs in the conclusion of the mapping operation while maintaining label information. This enables expanding the sample diversity for each input in the available dataset using a much larger set of possible modifications for a particular class compared to existing sample-level methods. Using the two-dimensional transformed appearance, the Möbius transformation improves model generalization and robustness against noise and dataset shift. Assuming three points as \(z_1\), \(z_2\), \(z_3\) in the initial plane and \(w_1\), \(w_2\), \(w_3\) in the target plane, considering the preservation of the cross-ratio

[19] , [Number] which gives, where [Number] is.

[0020] The reduced form of the transformation function can be expressed as follows. [Number]

[0021] Subsequently, the values of the coefficients \(a\), \(b\), \(c\), and \(d\) for the selected points \((z_1,z_2,z_3)\) and \((w_1,w_2,w_3)\) can be obtained as follows through substitution in equations (1), (3), and (4). a = w1w2z1 - w1w3z1 - w1w2z2 + w2w3z2 + w1w3z3 - w2w3z3 (5a) b = w1w3z1z2 - w2w3z1z2 - w1w2z1z3 + w2w3z1z3 + w1w2z2z3 - w1w3z2z3 (5b) c = w2z1 - w3z1 - w1z2 + w3z2 + w1z3 - w2z3 (5c) d = w1z1z2 - w2z1z2 - w1z1z3 + w3z1z3 + w2z2z3 - w3z2z3 (5d)

[0022] Based on Liouville's theorem (Liouville, J., Extension au cas des trois dimensions de la question du trace geographique. Note VI, pp. 609 - 617, 1850), a Möbius transformation can be represented as a composition of translations, orthogonal transformations, and inversions, which encompasses a superset of some common extension operations in deep learning. This helps to design a framework of an algorithm for real - time generation of Möbius transformations using the values of a, b, c, d from (5a, 5b, 5c, 5d), and forms a subspace of the composition of basic transformations from a superset of generalized Möbius transformations. A finite number of Möbius samples can be obtained, and the number of samples is bounded by a cut - off randomly assigned at runtime within [1, R], where R is the maximum number of samples allowed by the memory constraint. In the following exemplary embodiment, R was set to 250 based on the available RAM environment in this exemplary embodiment.

[0023] Weighted distillation The representation learned by the model can also be considered to represent the "dark knowledge" (Hinton et al., NIPS 2014 Deep Learning Workshop (2014)) about the model-data dynamics in a compact vectorized form. Since the learning of a heavier model is distilled into an essential and compact representation that can be used in other tasks, this process is called knowledge distillation. The method of the present disclosure uses the vectors as the "memory" of past class learning and regularizes the incremental training. Based on the initial learning, the class-averaged logits are retained for each class by performing a save to store the validation logits at the end of the training schedule of the initial (task 1) training. Next, weighted logits are calculated by applying a weighting factor to the logits of individual classes, and the weights are the reciprocals of the class-specific validation accuracies. This enables the distilled logits to reflect the bias regarding the classes in proportion to the difficulty of the model learning. The training of the initial classes employs cross-entropy loss. The probability vector of the initial task is p = softmax(z) ∈ 1, where z is a set of logits. The objective function in the initial training stage is as follows.

Number

[0024] Here, pi is the predicted probability score vector for each class in the new task, and yi is the related ground truth in the form of one-hot encoding. In the next session, a distillation term is added to the objective function to enable the representation of past knowledge in the learning process (y' is the final layer class score for the new task class before the softmax step).

Number

[0025] The logits and predictions are scaled using the temperature term T in the softening process. Softening using the temperature hyperparameter helps reduce the imbalance between the class label with the highest confidence score in the probability vector and other class labels, and also helps more appropriately reflect the relationships between classes in the representation learning stage. Considering the overall logit vector for the old classes and weighting it as zold, the class-specific logits are weighted, and the sum of the class-weighted logits is obtained as follows. [Number]

[0026] The logits zi from individual classes, i ∈ [1, K1], are calculated, for example, by averaging the pre-softmax probability values (after sigmoid activation) from each of the K1 classes. The weights (u1, u2, …, uk1) are calculated as the reciprocal of the class-specific accuracy for the validation set of the initial classes. This idea is to boost the logits from classes that are essentially difficult for the model to learn (the lower the class-specific accuracy, the higher the class weight). This reduces the imbalance between classes in the contribution to the overall session representation vector that is saved as the imprint of the stage 1 learning. Overall, the net incremental objective function for learning beyond the initial session is (γ = 0.5): L = γL crossent +(1 - γ)L distillation (9) That is.

[0027] Method for continuously training a neural network The present disclosure provides a novel method of continuous machine learning using the Möbius transformation for online extension. Subsequently, the present disclosure shows the values of the generalized Möbius transformation for performing the extension of the case set in the distillation-based incremental learning scenario and introduces a new concept of incremental extension for the retained cases. As shown in the examples, this method has been verified against a real-world data set of histological images of liver cancer. The present disclosure provides a computer-implemented method for improving the classification performance of a neural network module adapted to learning on temporally spaced inputs of imaging data, the method comprising: a) receiving an imaging data set, the set comprising data about a set of images having assigned classification labels; b) training a neural network module on the imaging data set from step (a) and storing the resulting neural network module; c) applying Mobius data augmentation by means of a Mobius transformation to one or more images from the data set of step (a) and storing the resulting transformed images; d) receiving a new imaging data set, the set comprising data about a set of images having assigned classification labels; e) training a neural network module (100) on the combination of imaging data obtained in steps (c) and (d) and storing the resulting neural network module comprising.

[0028] A neural network module refers to an instruction set that describes the architecture of such a neural network and a related data file containing information about the nodes of such a neural network.

[0029] In one embodiment, the additional imaging data set in step (d) comprises images of the same class that are present in the original data set of step (a). This makes it possible to improve the classification accuracy of the neural network for sets of the same class.

[0030] In one embodiment, to achieve continuous learning of the neural network, the additional image data set in step (d) includes not only images of classes that exist in the original data set of step (a), but also images belonging to new classes (incremental classes). In one embodiment, the new image data set includes one image of a new class data set that has not yet been presented to the neural network.

[0031] As shown by way of example, the present method can be applied to medical imaging data. Specifically, as a proof of concept, the present method has been applied to a set of tissue structure images. Such a method can be similarly applied to three-dimensional images such as MRI images and CT images. One approach to adapting the present method for use with three-dimensional images is to convert them into two-dimensional space by slicing those images.

[0032] If the network has already been trained on a set of imaging data, the method may repeat steps (c - e), and subsequent new images received at a point in time after step (d) are added to the extended set of step (c). In such a case, the present disclosure provides a computer-implemented method for improving the classification performance of a neural network module pre-trained on a set of imaging data, the neural network module being adapted to learn on temporally spaced inputs of imaging data, the method comprising: a) applying data augmentation by Mobius transformation to one or more imaging data from the data set already used to train the neural network module (100) and storing the resulting transformed imaging data, wherein the imaging data has a classification label assigned to each image; b) receiving a new imaging data set, the set including data for a set of images having assigned classification labels; c) updating the neural network module by training the neural network (100) on the combination of the imaging data obtained in step (a) and step (b) and storing the resulting neural network module; Includes:

[0033] In one embodiment, the additional image dataset in step (b) includes images of the same classes present in the original dataset in step (a), thereby improving the classification accuracy of the neural network for the same set of classes.

[0034] In one embodiment, to achieve continued training of the neural network, the additional image data set in step (b) includes images of classes present in the original data set used to generate the imaging data in step (a) as well as images belonging to a new class (an incremental class). In one embodiment, the new image data set includes one image of a new class data set that has not yet been presented to the neural network.

[0035] Such methods may be implemented as part of an image analysis system that is adapted for continuous learning and improvement of classification performance at subsequent time points.

[0036] Image Analysis System It is clear that the above method can be implemented on an image analysis system, such as a medical image analysis system, that receives classified images and incrementally trains a neural network using new and previous images processed by applying a Möbius transform.

[0037] Implementation of the method The present disclosure is also applicable to computer programs, particularly those adapted to implement the systems and methods of the present invention, especially computer programs on or in a carrier. The present invention further provides a computer program including code means for performing the steps of the methods described herein, wherein the execution of the computer program is performed on a computer. The present invention further provides a non-transitory computer-readable medium having stored thereon executable instructions that, when executed by a computer, cause the computer to perform a method for incremental training of a deep learning model as described herein. The present invention further provides a computer program including code means for elements of the systems disclosed herein, wherein the execution of the computer program is performed on a computer.

[0038] The computer program may be in the form of source code, object code, or intermediate source code. The program may be in a partially compiled form or in any other form suitable for use in implementing the methods according to the present invention and variations thereof. Such a program may have designs of different architectures. The program code for implementing the functions of the method or system according to the present invention may be subdivided into one or more subroutines or sub-components. There are many different ways of distributing functions among these subroutines, which will be known to those skilled in the art. The subroutines may be stored together in one executable file to form a self-contained program. Alternatively, one or more or all of the subroutines may be stored in at least one external library file and may be statically or dynamically linked, for example, to the main program at runtime. The main program includes at least one call to at least one of the subroutines. The subroutines may also call each other.

[0039] The present invention further provides a computer program product including computer-executable instructions for performing the steps of the method described herein or a variation thereof described herein. These instructions may be subdivided into sub-routines and / or stored in one or more files that may be linked statically or dynamically. Another embodiment of the computer program product includes computer-executable instructions corresponding to each means of at least one of the systems and / or products described herein. These instructions may be subdivided into sub-routines and / or stored in one or more files.

[0040] Examples of computer-readable media suitable for storing computer program instructions and data include, by way of illustration, all forms of non-volatile memory, media, and memory devices including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory may be supplemented or incorporated by dedicated logic circuitry.

[0041] It should be noted that the above-described embodiments are illustrative rather than limiting of the present invention, and those skilled in the art can design many alternative embodiments without departing from the scope of the claims. In the claims, any reference signs within parentheses shall not be construed as limiting the claims. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those recited in the claims.

[0042] The article "one" preceding an element does not exclude the existence of a plurality of such elements. The present invention may be implemented by hardware comprising several distinct elements and preferably by a computer programmed accordingly. In system claims enumerating several elements, some of these elements (subsystems) may be embodied by exactly the same item of hardware. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used.

[0043] "Exemplary embodiments" Anonymized colorectal cancer HE-stained tissue slides at 20x magnification were obtained using an Aperio ScanScope scanner. These are digitized and anonymized images of formalin-fixed paraffin-embedded human colorectal adenocarcinoma, which are publicly available through the pathology archive of the University Medical Center Mannheim (Kather et al., In Scientific reports, 2016). These slides were manually annotated and contain tessellated continuous tissue regions. These were converted into RGB patches of 150×150×3. In total, 5000 images were obtained for different tissue classes. In this study, eight classes with 625 samples each were, respectively, 1. Tumor epithelium, 2. Simple stroma (tumor stroma, peritumoral stroma, and smooth muscle homogeneous), 3. Complex stroma (single tumor cells and immune cells), 4. Debris (necrosis, hemorrhage, and mucus), 5. Immune cells (aggregates of immune cells and submucosal lymphoid follicles), 6. Normal mucosal gland, 7. Adipose tissue, 8. Background. The base classes include tumor epithelium (TE), simple stroma (SS), immune cells (IC), and adipose tissue (AT). The incrementally learned classes include complex stroma (CS), debris (De), normal mucosal gland (NMG), and background (BG).

[0044] The experiment is divided into two consecutive tasks labeled as Task 1 and Task 2. The initial task is advanced using a standard cross-entropy objective function, and Task 2, i.e., the incremental task, utilizes a joint loss with a cross-entropy term and a distillation loss. A ResNet-50 feature extractor is utilized, the layers following the last residual block are removed, a fully-connected (FC) layer of 512 units is added to the last residual block, and then an FC layer of 4 units (the number of classes) and a loss head are added. The pre-softmax layer generates probability scores by means of a sigmoid operation. An 80:20 split is used for the training:test split of the dataset. The input images are resized to 224×224, and a batch size of 50 is used along with a learning rate of 0.001 and adaptive moment optimization (Adam)

[26] . The model for Task 1 is trained for 150 epochs for the (N, label) set for all N frames. In Task 2, the model is trained for 150 epochs for the (N’, label, logit) tuple, where N’ has a Möbius-transformed version of the old samples selectively retained in addition to the new class data. Note that data augmentation during training is not performed except for the samples retained in the incremental training. This is in contrast to most efforts in machine learning in clinical imaging, but the aim of the inventors is to analyze the specific impact of the Möbius transformation on the performance of incremental learning together with distillation and the like. Therefore, boosting the accuracy of the base model is not an aim in this investigation. After performing a grid search with T∈[1,5], it was set to T = 4.0. Two 32GB Nvidia V100 GPUs and 512MB of RAM were used for a ResNet 50-based model with approximately 24.8 million parameters, and the average training time per epoch in both tasks was 102s. The Möbius extension module and the deep model are coded in Python3.7.1 and TensorFlow2.0, respectively.

[0045] Table 1. Accuracy (%) for the classes of Task 1 / Stage 1 after Task 1 has been trained and after Task 2 has been incrementally added at Stage 2. The difference in accuracy in the validation set of the classes of Task 1 indicates that they are being forgotten due to the addition of Task 2. [Table 1]

[0046] In the incremental task (Task 2), data from four classes added incrementally are used. The Möbius transformation for extension is applied exclusively to the retained instances from the classes of Task 1. Based on memory constraints, the top 20 retained instances sorted by the magnitude of the class confidence scores after the validation set when passing through the trained model after Task 1 are selected. Naturally, as shown theoretically in [2], performance can be improved by including more samples, and full co-training is the upper bound for incremental performance. Since local storage conditions limit the available memory buffer, it is necessary to minimize memory usage as in some clinical imaging workflows around the world. Therefore, the retention set is limited to 20 instances. The reduction of forgetting (Table 1) is remarkable with the weighted extension method having a ΔACC of 1.92 (the difference in the overall accuracy for the validation set of Task 1 before and after training of Task 2). In Table 1, the method using both weighted distillation and Möbius extension is labeled as "ours (MT+wKD)" and "ours (MT+KD)" (when using unweighted distillation). "ours (MT+FT)" is the method where fine-tuning is combined with Möbius extension. For the incremental task after training Task 2 and for the overall accuracy of all classes, a significant gain is seen with the method using the Möbius operation on the retained instances for the classes of the initial task before scattering batches of incremental classes in both the distillation approach and the fine-tuning approach. Overall, a clear advantage is seen when using distillation compared to the case of using only fine-tuning. The best results can be confirmed in the case of using a combination of distillation and Möbius extension before incremental optimization. This clearly shows the significance of extending the retained old samples. This is different from most distillation-based methods that retain some old samples without incremental extension for the retained samples and use data augmentation only in the initial session and for new incremental data.

[0047] The baseline from the literature is used with the ResNet-50 backbone and the original incremental training configuration shaped to fit the incremental aspects of the two tasks of this investigation. Figure 4 shows the performance of the final model in Task 2. Conventionally, it could be expected to have approximately equal accuracy among methods, but a slight difference is seen in the prediction of accuracy within the same class of Task 2. The effect of forward transfer of training for Task 1 combined with distillation-based regularization is more optimal when using the intermediate Möbius extension step for old examples that creates diverse sample sets for incremental training. The model with distillation shows generally better performance for new tasks due to distillation-induced regularization for parameter shift, unlike optimization without regularization in fine-tuning (FT). The Möbius transformation-based incremental augmentation (MT) is also compared with other augmentation ideas such as cutout (DeVries et al., arXiv preprint arXiv:1708.04552, 2017), Adatransform (Tang et al., Proceedings of the IEEE International Conference on Computer Vision, pp. 2998 - 3006, 2019), AutoAugment (Cubuk et al., In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 113 - 123, 2019), population-based augmentation (PBA) (Ho et al., arXiv preprint arXiv:1905.05393, 2019), RandAugment (Cubuk et al., arXiv preprint arXiv:1909.13719, 2019), 20° step rotation, and 10px window translation. This comparison of Δacc values shows that it outperforms some sample-level methods in reducing forgetting by augmenting old task samples before incremental training.Future research could focus on investigating the effectiveness of the Möbius transformation for other tasks such as segmentation compared to the generation expansion method, and exploring Möbius extensions combined with simultaneous methods in the literature.

Claims

1. A method implemented by a computer for improving the classification performance of a neural network module pre-trained on a set of imaging data, wherein the neural network module is adapted to learn on temporally spaced inputs of imaging data, the method comprising: a) applying Mobius data augmentation by means of a Mobius transformation to one or more imaging data from a data set already used for training a neural network module (100), and storing the resulting transformed imaging data, wherein the imaging data has a classification label assigned to each image; b) receiving a new imaging data set, the set comprising data about a set of images having assigned classification labels; c) updating the neural network module by training the neural network (100) on the combination of the imaging data obtained in steps (a) and (b), and storing the resulting neural network module. A method implemented by a computer, comprising the steps above.

2. The method according to claim 1, wherein steps (a) to (c) are repeated and subsequent new images from step (b) are added to the augmented set of step (a).

3. The method according to claim 1, wherein the additional image data set in step (b) comprises images of the same class as those already used for training the neural network module.

4. The method according to claim 1, wherein the additional image data set in step (b) comprises images belonging not only to the classes (base classes) of images present in the original data set used for training the neural network module, but also to new classes (incremental classes).

5. The method according to claim 1, wherein the imaging data set is an imaging data set of medical images.

6. The method according to claim 5, wherein the medical images are two-dimensional medical images.

7. The method according to claim 5, wherein the medical image is a tissue structure image.

8. An image acquisition device, A machine-readable medium for storing a neural network module, One or more processors configured to perform the steps of the method according to claim 1 An image analysis system comprising the same.

9. The image analysis system according to claim 8, further comprising a user interface configured to enable a user to classify an image generated by the image acquisition device.

10. A machine-readable medium that, when executed by a processor, includes instructions that cause the processor to perform the steps of the method according to claim 1.