Patient-specific artificial neural network training system and method

The patient-specific training system enhances neural network diagnostic accuracy by separating data based on confidence and training with reliable patient-specific images, addressing the challenge of insufficient data in medical diagnostics.

JP2025527413APending Publication Date: 2025-08-22SIEMENS HEALTHINEERS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504228
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-27
Filing Date
2023-07-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Existing artificial neural networks struggle to accurately diagnose patient-specific conditions due to insufficient or rare data, leading to misclassification and reduced diagnostic accuracy, particularly in medical applications like white blood cell differentiation.

Method used

A patient-specific training system that separates image data into reliable and unreliable sets based on classification confidence, training the neural network only with the reliable set, and refining the classification using patient-specific features.

Benefits of technology

Improves diagnostic accuracy by training the neural network with patient-specific data, enhancing its ability to classify images with high confidence and adapt to individual patient characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527413000001_ABST
    Figure 2025527413000001_ABST
Patent Text Reader

Abstract

The present invention relates to a training system and a training method for training an artificial neural network with patient-specific characteristics, particularly for medical applications. The patient-specific artificial neural network training system includes an input interface configured to receive image data from a patient, a computing device further including a classification module, a separation module, and a training module, and an output interface configured to output a diagnostic signal, wherein the classification module is configured to take the received image data as input and generate a first classification signal for each image of the image data, the separation module is configured to separate the received image data into at least a first data set and a second data set based on the first classification signal and according to a reliability criterion, and the training module is configured to train the artificial neural network by using only the first data set or a part thereof as a training data set.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a training system and method for training artificial neural networks, preferably for medical applications, taking into account patient-specific characteristics. In particular, the present invention relates to improving the performance of artificial neural networks as classifiers and tools for aided diagnosis, as well as increasing their adaptability to each patient's characteristics.

[0002] By way of example, while the present invention applies primarily to the fields of hematology and immunology, and in particular to the differentiation of white blood cells, the principles of the present invention have a much broader scope. [Background technology]

[0003] Medical diagnostics is an area where the introduction of computer-aided applications, particularly artificial intelligence and machine learning, has opened up vast possibilities. In particular, computerized clinical decision support systems have continued to evolve since their introduction 40 years ago. In recent years, several countries have progressed to HIMSS Stage 6, and the use of electronic health records (EHRs) has become the rule rather than the exception.

[0004] Artificial neural networks are one of the most promising methods for supporting medical decision-making: in samples where large amounts of training data are available, artificially intelligent entities compete with unaided human decision-making.

[0005] Even more problematic are cases where there is insufficient data: similarly, in rare cases where the training data is only sparsely covered, the number of outliers may be too small for the neural network to identify them and extract meaningful information from them.

[0006] In medicine and biology, this can occur not only with rare disease markers, but also with the diagnosis of a particular patient. Data on a particular patient is rare compared to the average data collected from a patient population. Deviations from the average are expected and, in some cases, may interfere with diagnosis. For example, a patient may have low blood iron levels or low blood pressure, but this may not indicate disease. In a white blood cell classification, a patient may exhibit abnormal basophils, for example, with below-average granulation, but again, this is a feature rather than a marker for disease.

[0007] Prior to diagnosing a patient, experts seek to gather as much information about the patient as possible in order to adapt to the patient's unique characteristics to the extent possible and arrive at a correct diagnosis. Incorporating this patient-specific knowledge into an artificial neural network is challenging, yet represents a key challenge for improving computer-aided diagnostic systems. Summary of the Invention [Problem to be solved by the invention]

[0008] It is an object of the present invention to provide a training system and method for training artificial neural networks used as tools for aided diagnosis in medical applications, with an emphasis on adapting the artificial neural network to patient characteristics in order to improve diagnostic accuracy. [Means for solving the problem]

[0009] The object of the invention is achieved by a device with the features disclosed in claim 1 and by a method with the features detailed in claim 14. Preferred embodiments of the invention with advantageous features are set out in the dependent claims.

[0010] A first aspect of the present invention provides a patient-specific artificial neural network training system, comprising: an input interface configured to receive image data from a patient; a computing device further including a classification module, a separation module, and a training module; and an output interface configured to output a diagnostic signal, wherein the classification module is configured to take the received image data as input and generate a first classification signal for each image of the acquired image data, the separation module is configured to separate the received image data into at least a first data set and a second data set based on the first classification signal and according to a first reliability criterion, and the training module is configured to train the artificial neural network by using only the first data set, or a portion thereof, as a training data set.

[0011] The input and / or output interfaces are broadly understood as entities that can acquire and send image data for further processing. Accordingly, each of them can include at least a central processing unit (CPU), at least one graphics processing unit (GPU), at least one field programmable gate array (FPGA), at least one application-specific integrated circuit (ASIC), and / or any combination thereof. Each of them can further include working memory operatively coupled to the at least one CPU and / or non-transitory memory operatively coupled to the at least one CPU and / or working memory. Each of them can be implemented in part and / or in whole in a local device and / or a remote system, such as a cloud computing platform.

[0012] The input and / or output interfaces may be implemented in hardware and / or software, cabled and / or wireless, and any combination thereof. Any of the interfaces may include interfaces to an intranet or the internet, cloud computing services, remote servers, and / or the like.

[0013] Image data broadly refers to any visual information that can be used for further processing and analysis in medical applications. In the context of blood cell classification, image data may consist of images of artificially stained blood samples (or: stained blood smears) using one of the staining methods used in blood cell classification. These image data may also include selected sub-image portions or crops of larger images (with the application of segmentation and / or inclusion of the object to be classified together with the surrounding background). In the following, image data, image dataset, or image will be used interchangeably to refer to information in image format.

[0014] Similarly, in this specification, a training dataset primarily refers to training images, and hereafter both terms will be used interchangeably.

[0015] The artificial neural network referred to in this specification is always understood as a computerized entity that can implement various data analysis methods broadly referred to under the terms artificial intelligence, machine learning, deep learning, or computer learning. The artificial neural network can be a generative adversarial network (GAN), a convolutional neural network (CNN), or any other neural network. The artificial neural network to be trained is also referred to as an accurate classifier or a patient-specific artificial neural network.

[0016] A computing device is broadly understood as an entity capable of processing data. The computing device may be realized as any device including or consisting of at least a central processing unit (CPU), at least one graphics processing unit (GPU), at least one field programmable gate array (FPGA), at least one application-specific integrated circuit (ASIC), and / or any combination thereof. The computing device may further include a working memory operatively coupled to the at least one CPU and / or a non-transitory memory operatively coupled to the at least one CPU and / or the working memory. The computing device may execute software, applications, or algorithms with different capabilities for data processing. The computing device may be implemented in part and / or in whole on a local device and / or a remote system such as a cloud computing platform.

[0017] The diagnostic signal may include an indication that a particular disease or condition may be present or is present in the patient. The diagnostic signal may include a list of possible or probable diseases associated with the classified image after analysis by the patient-specific artificial neural network.

[0018] A second aspect of the present invention provides a computer-implemented patient-specific training method, preferably for use with the patient-specific training system of the first aspect, comprising: (a) acquiring image data from a patient; (b) generating a first classification signal for each image of the acquired image data; (c) separating the image data into at least a first data set and a second data set based on the first classification signal and according to a first reliability criterion; (d) initializing an artificial neural network; and (e) training the artificial neural network with all or a portion of the first data set.

[0019] In particular, the method according to the second aspect of the invention may be carried out by an imaging device according to the first aspect of the invention, and therefore features and advantages disclosed herein in relation to the imaging device are also disclosed for this method, and vice versa.

[0020] According to a third aspect, the present invention provides a computer program product comprising executable program code configured to, when executed, perform a method according to the second aspect of the invention.

[0021] According to a fourth aspect, the present invention provides a non-transitory computer-readable data storage medium comprising executable program code configured to, when executed, perform a method according to the second aspect of the present invention.

[0022] The non-transitory computer-readable data storage medium may comprise or consist of any kind of computer memory, in particular semiconductor memory such as solid-state memory. The data storage medium may also comprise or consist of a CD, DVD, Blu-ray disc, USB memory stick, etc.

[0023] According to a fifth aspect, the present invention provides a data stream comprising, or configured to generate, executable program code that is configured to perform a method according to the second aspect of the invention when executed.

[0024] One of the main ideas underlying the present invention is to provide a training system configured to train an artificial intelligence neural network for classification of medical image data that can incorporate patient-specific features. This is achieved by processing a set of images from a patient with a classification module (or trained classifier) ​​and, based on its output, dividing the images into two groups. The first group contains images that have been identified with high confidence, and the second group contains images that have not been identified with sufficient confidence. A training image dataset is generated from the images in the first group, which is used to train a patient-specific artificial neural network. The patient-specific artificial neural network is a precise classifier that contains more information about the patient than the trained classifier.

[0025] The device described above provides a simple implementation of a method for training an artificial neural network for the classification of medical images, which can then be used as a diagnostic aid. In one step, medical images, often having undergone a preliminary staining process, are acquired. In another step, a classification signal is acquired for each image, indicating, for example, the likelihood that the image belongs to a different class of the classifier. Based on this classification signal and quantitative labeling criteria, the images can be separated into a first and a second data set. The (untrained) artificial neural network is initialized and trained with a portion of the first data set that contains images that can be more reliably identified.

[0026] An advantage of the present invention is that by training an artificial neural network with the most suggestive cells to perform refined classification, patient-specific features are acquired and learned by the artificial neural network. Thus, the classification and diagnostic accuracy of the patient-specific artificial neural network is higher than that achievable by an artificial neural network trained on data from many patients. This process of acquiring patient-specific features mimics the way humans incorporate patient-specific characteristics into each diagnosis.

[0027] Another advantage of the present invention is that the performance of the patient-specific artificial neural network can be easily tested by inputting images from a second data set into the trained artificial neural network. In performing this test, the quality of the refined classifier can be assessed by comparing the classification signal of the refined classifier to that achieved by the trained classifier. Successful training should increase the confidence measure for the images from the second data set, and in some cases, may even result in reliable identification and labeling of all or part of the images from the second data set.

[0028] Advantageous embodiments and further developments result from the dependent claims as well as from the description of various preferred embodiments shown in the accompanying drawings.

[0029] According to some of the embodiments, refinements or variations of the embodiments, the classification module comprises a trained artificial neural network structure.

[0030] According to some of the embodiments, refinements or variations thereof, the classification signal indicates, for each image, probability values ​​associated with a first set of classes. In other words, the artificial neural network can post-process the probability distributions associated with the images to assign class labels to the images.

[0031] According to some embodiments, refinements, or variations, a first data set includes all images above a threshold probability value, and a second data set includes all images below the same threshold probability value. The first data set includes highly suggestive cells of different classes from a first set of classes. Introducing an absolute probability threshold is a simple way to separate this first data set from the less conclusive cases that define the second data set. It is also possible, and in some applications advantageous, to divide the images into three or more data sets, for example by setting different probability thresholds.

[0032] According to some embodiments, refinements, or variations of the embodiments, the first data set includes all images above a relative threshold probability value between the classes, and the second data set includes all images below the same relative threshold probability value between the classes. In some cases, it is advantageous to define the reliability of the classification based on a relative probability measure. This reliability criterion emphasizes the discriminability of one class from another.

[0033] According to some embodiments, refinements, or variations, the training module further includes a selection unit configured to determine a second set of classes and assemble a training dataset based on information from the first dataset and / or the second dataset. While the first set of classes is determined by the trained artificial neural network of the classification unit, the second set of classes can be determined prior to training the artificial neural network. This depends on the results of the classification performed by the classification unit. For example, if the representatives of the first dataset significantly define only a subset of classes, it may be preferable to train the artificial neural network only on these well-represented classes. On the other hand, the second dataset may indicate that most ambiguous cases result from confusion between two classes. In this case, it may be advantageous to focus the refinement of the classification on these two classes, as this is most likely to improve the classification. Furthermore, the selection of a particular training dataset (e.g., the proportion of images from each selected class) is a parameter that is adjusted depending on the class distribution of the first dataset. In certain preferred embodiments, the second set of classes may also be considered a subset of the first set of classes. In certain other embodiments, the second set of classes may also include classes that are not present in the first set of classes.

[0034] According to some of the embodiments, refinements or variations thereof, the selection unit further includes a threshold discrimination subunit configured to select one of the first set of classes as one of the second set of classes on the condition that the number of images classified into one of the first set of classes is above a certain threshold. In other words, the formation of the second set of classes can be based on a quantitative criterion, in which only statistically significant classes are selected by setting a low threshold of representativeness per class.

[0035] According to some embodiments, refinements, or variations of the embodiments, the selection unit further includes a sampling subunit configured to sample a subset of data from the first dataset and / or augment the first dataset so that each of the second set of classes has a statistically significant number of images in the training dataset. Even with the introduction of a threshold, some classes may be underrepresented relative to other selected classes. This may result in unwanted bias in the training of the artificial neural network. To avoid this possible bias, it is advantageous to balance the amount of representatives per class. Typically, the sampling subunit will randomly select representatives of overrepresented classes to balance the distribution of training images according to the second set of classes. The sampling subunit can also generate additional training data for underrepresented or rare classes by applying augmentation techniques, such as rotating, mirroring, flipping, scaling, and / or resampling the image data. In general, increasing the amount of eligible training data through augmentation improves the performance of the artificial neural network.

[0036] According to some embodiments, refinements, or variations of the embodiments, the selection unit further includes an artificial intelligence subunit pre-trained to recognize a specific diagnosis and configured to select training images based on a second reliability criterion. If a specific disease or abnormality is the subject of the analysis of the patient images, the second set of classes will be associated with markers of the specific disease or abnormality. In this case, using an artificial neural network pre-trained to identify the specific disease or abnormality allows for a more efficient selection of the training data set. Depending on the available data, the artificial neural network can be trained with general data or patient-specific data. The reliability criterion can be a probability distribution among the various classes on which the artificial neural network is pre-trained.

[0037] According to some of the embodiments, refinements or variations of the embodiments, the training module is further configured to generate a second classification signal associated with each image of the training dataset, wherein the second classification signal indicates, for each image of the training dataset, a probability value associated with a second set of classes.

[0038] According to some of the embodiments, refinements or variants of the embodiments, the selection unit further comprises a training scenario subunit configured to select an artificial neural network of the training module, which can be based on the artificial neural network of the classification module or a different model, the training scenario subunit being able to process different models, the performance of which can be subsequently evaluated and compared.

[0039] According to some of the embodiments, refinements or variations thereof, the training scenario subunit is further configured to add patient-specific non-image data to a training dataset used to train the artificial neural network of the training module, in this way the training dataset of the artificial neural network is augmented with patient-based metadata.

[0040] According to some of the embodiments, refinements, or variations, the patient-specific non-image data includes information about at least a previously diagnosed disease. The patient's existing diagnoses determine the type of algorithm to be used, particularly an algorithm pre-trained to identify a specific diagnosed disease. The artificial neural network can also be trained with general data or data from the patient itself, depending on the availability of data.

[0041] According to some embodiments, refinements, or variations of the embodiments, the patient-specific training system further includes a quality control module configured to input image data from a second data set into the trained artificial neural network of the training module and compare, for each image, the second classification signal with the first classification signal. To test the quality of the precise classification, the artificial neural network of the training module can be used with the second data set, i.e., a data set that is considered non-deterministic in the classification performed by the classification module. The probability distribution obtained by the trained artificial neural network can then be compared with the probability distribution of the patient-specific artificial neural network. Training is successful if the images of the second data set are better identified, i.e., if their probability distributions lead to safe identification or at least improved identification.

[0042] According to some embodiments, refinements, or variations of the embodiments, the patient-specific training system further includes a validation module configured to input image data from the first data set into a trained artificial neural network of the training module and compare, for each image, the second classification signal with the first classification signal. The validation step is defined by inputting a first data set not used for training into the trained artificial neural network. The accurate classifier preferably correctly labels images with a high level of confidence.

[0043] Although certain functions are described herein above and below as being performed by modules, it will be understood that this does not necessarily mean that such modules are provided as separate entities from one another. When one or more modules are provided as software, these modules are implemented by program code sections or snippets that can be distinct from one another, but can also be interwoven or embedded within one another.

[0044] Similarly, when one or more modules are provided as hardware, the functionality of one or more modules may be provided by the same hardware component, or the functionality of multiple modules may be distributed across multiple hardware components, which do not necessarily correspond to modules. Alternatively or additionally, one or more of such modules may be associated with a more specific computer structure, such as implemented using FPGA technology. Thus, any apparatus, system, method, etc. that exhibits all of the structure and functionality attributed to a particular module shall be understood to include or implement that module. In particular, all modules may also be implemented by program code executed by a computing device (e.g., a server or cloud computing platform).

[0045] According to some of the embodiments, the method of the second aspect further comprises: validating the artificial neural network on a portion of the image data of the first data set, wherein a second classification signal is generated for each image and compared to the first classification signal for the same image; and evaluating the performance of the artificial neural network on the image data of the second data set, wherein a second classification signal is generated for each image and compared to the first classification signal for the same image.

[0046] According to some embodiments, the patient-specific training system is part of a computer-aided clinical diagnosis system embedded in a diagnostic decision support system (DDSS) that assists human decision making in clinical workflow.

[0047] The above embodiments and implementations can be combined with each other as necessary, as long as it is reasonable.

[0048] Further scope of applicability of the present method and apparatus will become apparent from the following drawings, detailed description, and claims. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of example only, since various modifications and improvements within the spirit and scope of the invention will become apparent to those skilled in the art.

[0049] Aspects of the present disclosure will be better understood with reference to the following drawings. The components in the figures are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. In the drawings, like reference numerals are used to refer to parts in different figures that correspond to the same elements. [Brief explanation of the drawings]

[0050] [Figure 1] FIG. 1 is a schematic diagram of a patient-specific training system, according to one embodiment of the present invention. [Figure 2]1 is a schematic diagram of the structure of a selection unit according to some embodiments of the present invention. [Figure 3] FIG. 1 is a block diagram illustrating an exemplary embodiment of a patient-specific training method. [Figure 4] FIG. 2 illustrates an example of possible image processing according to some embodiments of the present invention. [Figure 5] FIG. 10 is a schematic block diagram illustrating a computer program product according to an embodiment of the third aspect of the present invention. [Figure 6] FIG. 10 is a schematic block diagram illustrating a non-transitory computer-readable data storage medium according to an embodiment of the fourth aspect of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The drawings may not be to scale, and certain components may be shown in generalized or schematic form for clarity and simplicity. In some instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the concepts of the present invention. Similarly, numbering of method steps is for ease of description of each. They do not necessarily imply a particular order of steps. In particular, multiple steps may be performed simultaneously.

[0052] The detailed description includes specific details for the purpose of providing a thorough understanding of the present invention, although it will be apparent to one skilled in the art that the present invention may be practiced without these specific details.

[0053] Figure 1 is a schematic diagram of a patient-specific training system 100, according to one embodiment of the present invention. For illustrative purposes, the training system of Figure 1 is designed to train an artificial neural network for white blood cell classification, although it should be understood that this particular example does not limit the scope of the present invention.

[0054] The training system 100 includes an input interface 10, a computing device 20, and an output interface 30. The input interface 10 is broadly understood as any entity that can acquire image data D1 and send it for further processing.

[0055] A computing device 20 may be broadly understood as any entity capable of processing data. The computing device 20 shown in FIG. 1 includes a classification module 2, a separation module 3, and a training module 4. The classification module 2 is configured to receive image data D1 as input and classify it according to some classification criteria. In its simplest version, the classification module 2 includes a trained artificial neural network. In the particular application depicted in FIG. 1, the artificial neural network is pre-trained to classify different types of white blood cells, e.g., basophils C1, monocytes C2, lymphocytes C3, blasts C4, and myeloid cells C5. The classification module 2 outputs a first classification signal that assigns to each image D1 a probability distribution that measures the likelihood that the image D1 belongs to each of the classes C1-C5.

[0056] The separation module 3 is configured to interpret the probability distribution for each image D1 and separate the images D1 into at least a first data set D2 and a second data set D3. The first data set D2 contains images D1 that can be reliably assigned to one of the classes C1 to C5, while the second data set D3 contains images D1 whose identification is considered uncertain.

[0057] According to some embodiments of the present invention, the criterion for confidently labeling an image is determined by a probability threshold. For example, an image is labeled as a class if the associated probability is greater than 70%. This probability threshold can be predetermined. For example, it can be considered part of an existing classification protocol, or it can be based on an obtained probability distribution if it can be determined inductively. With this absolute probability threshold criterion, a first image with a probability distribution of (75, 5, 2, 10, 8) is retained in the first dataset D2 and classified as a basophil (75% probability), while a second image with a probability distribution of (60, 20, 5, 5, 10) is retained in the second database D3.

[0058] According to some embodiments of the present invention, relative probability may be preferred. For example, if the probability difference between the highest probability and the next highest probability is greater than 30%, the image is labeled as positively identified. Returning to the previous example, under this relative probability criterion, a first image with a probability distribution of (75, 5, 2, 10, 8) and a second image with a probability distribution of (60, 20, 5, 5, 10) would both be retained in the first dataset D2. Both images would then be classified as basophils (with probabilities of 75% and 60%, respectively).

[0059] The training module 4 is configured to input the first data set D2 or only a part of it to train the artificial neural network 41. The training module 4 comprises or consists of the artificial neural network 41 to be trained, a selection unit 42, a quality control module 43 and a validation module 44.

[0060] The artificial neural network 41 can be, for example, a generative adversarial neural network (GAN) or a convolutional neural network (CNN). Figure 1 shows, by way of example, an artificial neural network 41 that includes three classes C1' to C3'. As mentioned above, the classes C1' to C3' are preferably a subset of the classes C1 to C5 of the classification module 2.

[0061] The selection unit 42 is configured to generate a training data set D4 from the first data set D2 and, in some embodiments, to select an artificial neural network 41 to train and classes C1'-C3' for classification. To perform these functions, the selection unit 42 may include a threshold discrimination subunit 421, a sampling subunit 422, an artificial intelligence subunit 423, and / or a training scenario subunit 424. The structure of the selection unit 42 is discussed with respect to FIG. 2.

[0062] After training the artificial neural network 41, a quality control module 43 and a validation module 44 can be employed. The quality control module 43 is configured to input images of a second data set D3 that were not used for training into the trained patient-specific artificial neural network 41. To evaluate the performance of the patient-specific artificial neural network 41 relative to the trained classifier of the classification module 2, the quality control module 43 compares a first classification signal and a second classification signal for each image of the second data set D3. As mentioned above, these classification signals are confidence measures and, in some embodiments, will be realized as probability distributions over all available classes of the classifier.

[0063] The quality control module 43 compares these probability distributions. In this way, the extent to which the artificial neural network 41 has learned the patient-specific features can be quantitatively determined. The probability distribution for the images in the second data set D3 from the second classification signal should be better than the probability distribution of the first classification signal. Ideally, this improvement will be good enough that images in the second data set D3 that could not be identified reliably by the separation module 3 can be reclassified and identified reliably. This reclassification of images D1 can also be performed by the separation module 3 using the same probability thresholds employed in processing or modifying the first classification signal.

[0064] The validation module 44 is preferably configured to input images of the first data set D2 that were not employed in training, e.g. to avoid overfitting, into the trained patient-specific artificial neural network 41. To validate the artificial neural network 41, the first classification signal is compared for each image against the corresponding second classification signal. Since these images have already been reliably classified by the separation module 3 based on the first classification signal, it is expected that the second classification signal will only confirm the classification, possibly with a higher degree of confidence.

[0065] The output interface 30 is broadly understood as any entity capable of generating a diagnostic signal based on information obtained from the training module 4. It may therefore include at least a central processing unit (CPU), at least one graphics processing unit (GPU), at least one field programmable gate array (FPGA), at least one application specific integrated circuit (ASIC), and / or any combination thereof. It may further include working memory operatively coupled to the at least one CPU and / or non-transitory memory operatively coupled to the at least one CPU and / or the working memory. This may be implemented in part and / or in whole in a local device and / or a remote system such as a cloud computing platform.

[0066] When the training system 100 is implemented as part of a computer-aided clinical diagnosis embedded in a diagnostic decision support system (DDSS), the output interface 30 can present the diagnostic signal on a user interface, such as a website, application, software, or front-end interface of an electronic health record (EHR).

[0067] Any of the input interface 10 and / or output interface 30 may be implemented in hardware and / or software, cabled and / or wireless, and any combination thereof. Either of the interfaces 10, 30 may include an interface to an intranet or the internet, a cloud computing service, a remote server, and / or the like.

[0068] Figure 2 is a schematic diagram of the structure of a selection unit 42 according to some embodiments of the present invention. The selection unit 42 is configured to determine, based on information from a first data set D2 and a second data set D3, a training data set D4 from the first data set D2 to be used for training an artificial neural network (41 in Figure 1). Furthermore, the selection unit 42 may determine the type of artificial neural network 41 to be used, the training algorithm, as well as the selection of classes (C1'-C3' in Figure 1) that the artificial neural network 41 will use to classify images.

[0069] FIG. 2 shows a possible configuration of the selection unit 42, which includes a threshold discrimination subunit 421, a sampling subunit 422, an artificial intelligence subunit 423, and a training scenario subunit 424.

[0070] The threshold discrimination subunit 421 is configured to select a particular class from the first set of classes C1 to C5 as one of the second set of classes (C1' to C3' in FIG. 1) provided that the number of images classified into that particular class from the first set of classes (C1 to C5 in FIG. 1) exceeds a certain threshold. The data reduction when passing from the received images D1 to the first data set D2 means that it may be statistically significant for the second set of classes C1' to C3' to be smaller than the first set of classes C1 to C5. The threshold applied by the threshold discrimination subunit 421 is a measure of the statistical significance required for training multiple classes from the second set of classes C1' to C3'. A lower level option is to simply keep a fixed number of classes (two, three or more) from the first set of classes C1 to C5 when training the artificial neural network 41. Preferably, these classes are the classes with the most representation for training.

[0071] The sampling subunit 422 is configured to sample a subset of data from the first data set D2 and / or augment the first data set D2 so that each of the second set of classes C1'-C3' has a statistically significant number of representatives. The sampling subunit 422, in conjunction with the threshold discrimination subunit 421, can avoid imbalance among the second set of classes C1'-C3' selected for training the artificial neural network 41. It is important to maintain a roughly equal number of training representatives for each class, or at least to avoid under- or over-representation of certain classes. The sampling subunit 422 can perform random selection, pruning representatives from the highest-ranking classes, or it can simply hierarchically prune representatives of classes with lower discrimination confidence. Depending on the capabilities of the data, image augmentation may also be used to balance the training data set D4. An added benefit is that the overall number of representatives increases, potentially improving the classification quality of the artificial neural network 41.

[0072] The artificial intelligence subunit 423 is configured to select the second set of classes C1'-C3' based on a particular diagnosis. The artificial intelligence subunit 423 includes an artificial neural network (not shown) pre-trained with data to recognize at least a particular disease. In some embodiments, since the training system 100 is intended to be part of a computer-aided diagnosis system, using a trained artificial neural network to determine the second set of classes C1'-C3' may be advantageous. For example, one may wish to use the training system 100 to determine whether a patient previously diagnosed with a blood cancer will relapse. The artificial intelligence subunit 423 can be trained with a blood sample taken when the patient was first diagnosed, and this trained artificial neural network is used to determine the classes C1'-C3' for the artificial neural network 41. In this case, the amount of data available for training is obviously a limiting factor, but the patient's blood sample could also be supplemented with blood samples from other patients diagnosed with the same type of blood cancer.

[0073] The training scenario subunit 424 is configured to select a model and algorithm for the artificial neural network 41. One advantage of this operation is that it allows for training of a variety of artificial neural network structures, since it is not always clear which artificial neural network 41 will perform better for a given application. The training scenario subunit 424, in combination with the quality control module 43, can then determine the artificial neural network 41 that is best suited to the application at hand.

[0074] The training scenario subunit 424 is further configured to add patient-specific non-image data to the data used to train the artificial neural networks of the training module 4. The use of metadata in some artificial neural networks 41 can certainly improve their performance, for example if the metadata contains information about the patient's past condition or disease.

[0075] All of these subunits shown in Figure 2 can operate simultaneously as long as their respective operations do not cause contradictions or conflicts.

[0076] FIG. 3 is a block diagram illustrating an exemplary embodiment of a training method M, preferably applied in conjunction with the training system 100 of FIG. 1. Step M1 involves acquiring a patient image dataset D1, which in most preferred embodiments of white blood cell classification has previously undergone a staining and segmentation process. Step M2 involves generating a first classification signal for each acquired image D1. In most preferred embodiments, this is achieved by inputting the images D1 to a trained artificial neural network, which outputs a probability distribution for each image as the first classification signal. Step M3 involves dividing the image dataset D1 into at least a first dataset D2 and a second dataset D3 based on the first classification signal. This division is performed using several quantitative criteria, as discussed above in the description of FIG. 1. After initializing the artificial neural network 41 in step M4, step M5 involves training it on the first dataset D2 or a portion thereof as a training dataset D4. The method for defining the subset of the first dataset D2 is discussed in detail, for example, in the description of FIGS. 1 and 2.

[0077] After the artificial neural network 41 has been trained, it can be validated in another step M6. In this step M6, a portion of the images of the first data set D1 that were not used in the training data set D4 is taken. A second classification signal is then generated for each image and compared with the first classification signal for the same images. In the most preferred embodiment, this corresponds to a comparison of two probability distributions. In another step M7, the performance of the artificial neural network 41 is evaluated using the images of the second data set D3. A second classification signal is then generated for each image and compared with the first classification signal for the same images. In the most preferred embodiment, this corresponds to a comparison of two probability distributions.

[0078] Figure 4 illustrates an example of possible image processing according to some embodiments of the present invention. Figure 4A shows six stained images N1-N6 as an example of image data D1 received by input interface 10. These images show a normal monocyte N1, a dysplastic monocyte N2, a post-dysplastic myelocyte N3, a normal segmented neutrophil N4, and two dysplastic myelocytes N5 and N6. These cells are from a patient with myelodysplastic syndrome (MDS), a type of blood cancer in which myelocytes, the precursor cells to neutrophils, do not mature normally.

[0079] The images N1-N6 are processed by a classification module 2, which generates a first classification signal, preferably in the form of a probability distribution, for each of the images N1-N6 according to a set of classes (C1-C5 in FIG. 1), preferably using a trained artificial neural network. For simplicity, the classes can be limited to two types of white blood cells (e.g., monocytes C2 and myeloid cells C5). In this case, if the goal is to determine whether a patient has MDS, pre-training the trained artificial neural network with samples diagnosed with MDS is preferred.

[0080] FIG. 4B illustrates the segmentation of images by the separation module 3. The probability distribution for each image N1-N6 becomes a label after applying a confidence criterion. As mentioned above, according to some preferred embodiments, a good criterion is the application of a probability threshold. In the example of FIG. 4B, after applying this threshold, images N4, N5, and N6 are characterized as myeloid cells C5, while image N1 is identified as monocyte C2. Therefore, the first dataset D2 includes images N1, N4, N5, and N6. In contrast, images N2 and N3 are non-deterministic and belong to the second dataset D3.

[0081] According to the principles of the present invention, this classification can be refined by training an artificial neural network 41 (see FIG. 1) using images N1, N4, N5, and N6 from the first dataset D2 or a subset thereof. FIG. 4B shows an example of a training dataset D4 including images N1, N4, and N6. This may be the result of applying the threshold determination subunit 421 or the sampling subunit 422 of the selection unit 42. In the first case, N5 is below a certain probability threshold. In other words, N5 is one of the less certain images in the first dataset D2 and is therefore excluded from the training dataset D4. In the second case, the sampling subunit 422 is used to reduce the ratio of myeloid cells to monocytes from 3:1 to 2:1. The exclusion of N5 may then be a random process. Alternatively, N5 may be specifically selected, for example, because it is the least certain of the myeloid cells.

[0082] Figure 4C shows the stage after the artificial neural network 41 has been trained on the training data set D4 shown in Figure 4B. Figure 4C shows the prediction of successful refined classification when images N2 and N3 from the second data set D3 shown in Figure 4B are input to the artificial neural network 41. This operation is managed by the quality control module 43, which generates a second classification signal, preferably in the form of a probability distribution for images N2 and N3, indicating the probability that images N2 and N3 are myeloid cells C5 or monocytes C2. The separation unit 3 determines whether the images can be identified (i.e., can be reliably determined according to a selected reliability criterion), as described above. In the example of Figure 4C, N3 is unexpectedly identified as myeloid cell C5, while N2 is labeled as monocyte N2.

[0083] Image N2 may not be identifiable by a trained artificial neural network trained to identify generic monocytes. Instead, refined training adds patient-specific features learned by the artificial neural network 41. As a result, if the patient's white blood cells were used for training, N2 would be found to resemble the patient's monocytes and can be accurately labeled. The end result is a trained artificial neural network 41 that is customized for each patient and can provide more reliable information for aiding in the diagnosis of MDS, in this particular example.

[0084] Figure 5 is a schematic block diagram illustrating a computer program product 300 according to an embodiment of the third aspect of the present invention. The computer program product 300 comprises executable program code 350 that, when executed, is configured to perform a method according to any embodiment of the second aspect of the present invention, particularly as described with reference to the preceding figures.

[0085] 6 is a schematic block diagram illustrating a non-transitory computer-readable data storage medium 400 according to an embodiment of the fourth aspect of the present invention. The data storage medium 400 includes executable program code 450 that, when executed, is configured to perform a method according to any embodiment of the second aspect of the present invention, particularly as described with reference to the preceding figures.

[0086] The non-transitory computer-readable data storage medium may comprise or consist of any kind of computer memory, in particular semiconductor memory such as solid-state memory. The data storage medium may also comprise or consist of a CD, DVD, Blu-ray disc, USB memory stick, etc.

[0087] The above description of the disclosed embodiments is merely illustrative of possible implementations, and is intended to enable any person skilled in the art to make or use the present invention. Various modifications and improvements to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the present disclosure. Thus, the present invention is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Accordingly, the present invention is not limited except as by the appended claims. [Explanation of symbols]

[0088] 2. Classification Module 3 Separation Module 4 Training Modules 10 Input Interface 20. Computer Devices 30 Output Interfaces 41 Artificial Neural Networks 42 Elective Units 43 Quality Control Module 44 Verification Module 100 Patient-Specific Artificial Neural Network Training Systems 300 Computer Program Products 350 executable program code 400 Non-transitory computer-readable data storage medium 421 Threshold Discrimination Subunit 422 Sampling Subunit 423 Artificial Intelligence Subunit 424 Training Scenario Subunit 450 executable program code D1 Image data D2 First Dataset D3 Second Dataset D4 training dataset N1 normal monocytes N2 dysplastic monocytes N3 postdysplastic myelocytes N4 Normal segmented neutrophils N5 Dysplastic myelocytes N6 Dysplastic myelocytes C1~C5 Classes in the first group C1'~C3' Second group classes M Computer-Implemented Patient-Specific Training Method M1~M7 Method steps

Claims

1. A patient-specific artificial neural network training system (100) comprising: an input interface (10) configured to receive image data (D1) from a patient; a computing device (20) further including a classification module (2), a separation module (3), and a training module (4); an output interface (30) configured to output a diagnostic signal; Including, The classification module (2) is configured to take the received image data (D1) as input and to generate a first classification signal for each image of the image data (D1); the separation module (3) is configured to separate the received image data (D1) into at least a first data set (D2) and a second data set (D3) based on the first classification signal and according to a first reliability criterion; The patient-specific artificial neural network training system, wherein the training module (4) is configured to train the artificial neural network (41) by using the first data set (D2) or only a part of it as a training data set (D4).

2. 10. The patient-specific training system of claim 1, wherein the classification module comprises a trained artificial neural network structure.

3. A patient-specific training system (100) according to claim 1 or 2, wherein the first classification signal indicates, for each image, a probability value associated with a first set of classes (C1-C5).

4. 4. A patient-specific training system (100) according to any one of claims 1 to 3, wherein the first data set (D2) comprises all images above a threshold probability value and the second data set (D3) comprises all images below the same threshold probability value.

5. 4. A patient-specific training system (100) according to any one of claims 1 to 3, wherein the first data set (D2) comprises all images above a relative threshold probability value between the classes (C1 to C5) and the second data set (D3) comprises all images below the same relative threshold probability value between the classes (C1 to C5).

6. 6. The patient-specific training system according to claim 1, wherein the training module further comprises a selection unit configured to determine a second set of classes and to assemble a training dataset based on information from the first dataset and / or the second dataset.

7. 7. The patient-specific training system according to claim 1, wherein the selection unit further comprises a threshold discrimination subunit configured to select one of the first set of classes as one of the second set of classes, provided that the number of images classified into one of the first set of classes exceeds a certain threshold.

8. 8. The patient-specific training system according to claim 1, wherein the selection unit further comprises a sampling subunit configured to sample a subset of data from the first data set and / or to extend the first data set such that each of the second set of classes has a statistically significant number of images in the training data set.

9. 9. The patient-specific training system (100) of claim 1, wherein the selection unit (42) further comprises an artificial intelligence subunit (423) pre-trained to recognize a particular diagnosis and configured to select the training data set (D4) based on a second reliability criterion.

10. 10. The patient-specific training system of claim 1, wherein the training module is further configured to generate a second classification signal associated with each image of the training data set, the second classification signal indicating, for each image of the training data set, a probability value associated with a second set of classes.

11. 11. The patient-specific training system (100) of claim 1, wherein the selection unit (42) further comprises a training scenario subunit (424) configured to select an artificial neural network (41) of the training module (4).

12. 12. The patient-specific training system (100) of claim 11, wherein the training scenario subunit (424) is further configured to add patient-specific non-image data to a training dataset (D4) used to train the artificial neural network (41) of the training module (4).

13. 13. The patient-specific training system (100) of claim 12, wherein the patient-specific non-image data includes information regarding at least a previously diagnosed disease.

14. A computer-implemented patient-specific training method (M) for use with a patient-specific training system (100), preferably according to any one of claims 1 to 13, comprising: acquiring image data (D1) from a patient (M1); generating a first classification signal for each image of the acquired image data; Separating (M3) the image data (D1) into at least a first data set (D2) and a second data set (D3) based on the first classification signal and according to a first reliability criterion; A step (M4) of initializing an artificial neural network (41); training (M5) an artificial neural network (41) with the first data set (D2) in whole or in part; the computer-implemented patient-specific training method comprising:

15. A diagnostic decision support system comprising a patient-specific training system (100) according to any one of claims 1 to 13.