Ai-assisted drift detector to optimize a diagnosis process

The drift detector improves disease detection by using multi-modal data and expert knowledge to analyze feature patterns, addressing inaccuracies in existing methods and enabling timely intervention.

WO2026027038A1PCT designated stage Publication Date: 2026-02-05NEC LAB EURO GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/071589
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing drift detection methods in medical imaging are inaccurate due to measurement noise and variability in data quality, making it difficult to reliably detect new diseases.

Method used

A drift detector that utilizes multi-modal data at the explanation level, combining expert knowledge from reliable sources with real-time data to identify significant changes in feature patterns, using a comparator to analyze empirical distributions and generate reports for timely intervention.

Benefits of technology

Enhances the accuracy and reliability of detecting new diseases by reducing the impact of background noise, allowing for early outbreak detection and informed decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024071589_05022026_PF_FP_ABST
    Figure EP2024071589_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a drift detector to detect a drift in features of a diagnosis, the drift detector comprising a first collecting device configured to collect relevant explanations and raw data, a first database compiling the relevant explanations and raw data collected by the first collecting device, a second collecting device configured to collect relevant available expert knowledge from reliable sources, a second database compiling the relevant available expert knowledge from reliable sources collected by the second collecting device, a learned feature extractor configured to group initial data of the second database into a characteristic feature pattern for different diagnoses and to subsequently group the data from the first database into first feature patterns by diagnosis and from the second database into second feature patterns by diagnosis, a comparator configured to compare the empirical distributions of each of the grouped feature patterns in the databases with each other, and a processor configured to generate a report if results of the comparison do not comply with a predefined condition. Applications include medical applications such as recognizing new diseases.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AI-assisted drift detector to optimize a diagnosis process

[0002] TECHNICAL FIELD

[0003] [OOI] The present disclosure relates to artificial intelligence (Al) systems and methods for detecting, predicting, and preventing outbreaks of new pathogens causing diseases. Applications for the present disclosure include, but are not limited to, use cases in the medical sector and in healthcare as well as for decision making in such sectors. More specifically, the present disclosure relates to machine learning (ML) technologies that allow to detect new diseases, e.g. in a medical context. In response suitable counter measures can be identified, optimized, and automatically triggered to improve health of the population of persons e.g., by enhancing patient safety, triggering infection control, and managing resource allocation in a hospital or similar environments.

[0004] TECHNICAL BACKGROUND

[0005]

[0002] When CO VID first appeared, doctors were unsure whether x-ray images showed already known diseases, such as pneumonia, or if they are looking at a new disease. In fact, the detection of new diseases is a challenging problem because symptoms, measurements, or phenotypes (i.e. the features of the disease) might be similar to other diseases and just slightly deviate. Moreover, for example considering x-ray scans, the quality of x-ray images might vary over time, vaiy depending on the machine, and vary depending on the physician. Consequently, not every sign of a slight deviation from the known feature patterns for a certain disease is immediately an indicator for a new disease.

[0006]

[0003] In literature, the problem is tackled by analyzing databases over time with statistical frameworks for drift detection. In the following, several prior art technologies are discussed that form a part of the general background of the present disclosure.

[0004] Adaptive Concept Drift Detection as discussed in Dries A, Riickert U, ‘Adaptive Concept Drift Detection” published in Statistical Analysis and Data Mining, 2: 311-327. Available: https: / / doi.org / 1o.1oo2 / sam.1oo54 (in the following Ref. [1]) addresses the problem that the statistics underlying statistical hypothesis testing do not always perform very well. In particular, three drift detection tests are presented, whose test statistics are adapted.

[0007]

[0005] An overview of unsupervised drift detection methods as discussed in Gemaque R, Costa A, Giusti R, dos Santos E, “An overview of unsupervised drift detection methods” published in WIREs Data Mining Knowledge Discovery. 2020; 10: 01381 Available: htps: / / doi.org / 1o.1oo2 / widm.1381 (in the following Ref. [2]) presents a comprehensive overview of approaches that tackle concept drift in classification problems in an unsupervised manner.

[0008]

[0006] Automatic correction of performance drift under acquisition shift in medical image classification as discussed in Roschewitz M, Khara G, Yearsley J, Sharma N, James J, Ambrozay E, Heroux Adam, Kecskemethy P, Rijken T, Blocker B, “Automatic correction of performance drift under acquisition shift in medical image classification” published in Nature Communications, 14, 6608 (2023). Available: https: / / d0i.0rg / 10.1038 / s41467-023-42396-y (in the following Ref. [3]) proposes Unsupervised Prediction Alignment to adequately detect and correct performance drift due to changes in data acquisition.

[0009]

[0007] From concept drift to model degradation: An overview on performance-aware drift detectors as discussed in Bayram F, Bestoun A, Kassler A, “From concept drift to model degradation: An overview on performance-aware drift detectors” published in Knowledge-Based Systems, Volume 245 (2022). Available: https: / / d0i.0rg / 10.1016 / j.kn0sys.2022.108632 (in the following Ref. [4]) groups concept drift types by their mathematical definitions and survey the different terms used in the literature to build a consolidated taxonomy of the field. It further reviews and classifies performance-based concept drift detection methods proposed in the last decade. [oo8] In other words, in the prior art, the temporal data or streaming data is analyzed to identify statistically significant changes. The problem these frameworks face is that they must also deal with the changes mentioned above such as measurement noise of machines or similar, which act as confounding factors, and makes them less accurate in the reliable detection of drifts and therefore new diseases [3], [4].

[0010]

[0009] The invention aims at providing a more robust and accurate way to detect drifts in the feature patterns that are characteristic of occurrences of a new disease. As such, this invention can be used as an early detection of potential outbreaks and allow for doctors and healthcare institutions to respond appropriately. Broadening the field of useful applications, the invention is designed to be employed in a more general way, not only to detect human related diseases.

[0011] SUMMARY OF THE INVENTION

[0012]

[0010] The above-mentioned object is at least partially realized by a drift detector according to claim 1 or by a method for detecting a drift according to claim 8.

[0013]

[0011] In particular, the drift detector comprises a first collecting device configured to collect relevant explanations and raw data, a first database compiling the relevant explanations and raw data collected by the first collecting device, a second collecting device configured to collect relevant available expert knowledge from reliable sources, a second database compiling the relevant available expert knowledge from reliable sources collected by the second collecting device, a feature extractor configured to group initial data of the second database into a characteristic feature pattern for each diagnosis and to subsequently group the data from both databases into features by diagnosis, a comparator configured to compare the empirical distributions of each of the grouped features in the databases with each other, and a processor configured to generate a report if results of the comparison do not comply with a predefined condition. The first database can be seen as a temporal database whereas the second database corresponds to a reference database. In the databases, data containing the features of a diagnosis, such as symptoms, measurements, or phenotypes are compiled. When the comparison does not comply with the predefined condition, a significant drift between the first, temporal database and the second, reference database is detected.

[0014] Instead of basing the drift detection only on raw data as seen in the prior art, the present drift detector uses multi-modal data at explanation level to identify a drift. This way, background noise and unimportant features, which reduce the accuracy of drift detection systems, do not need to be considered.

[0015]

[0012] The first database is preferably configured to continuously receive data from users, and the second database is configured to retrieve its data automatically. In particular, the first, temporal database can be constantly updated over time based on provided explanations and raw data coming from doctors. The second, reference database is preferably automatically generated, for example by extracting information from reliable sources of the internet.

[0016] Because the proposed drift detector operates on multi-modal explanations automatically collected from the internet and is therefore expert informed, the curse of dimensionality typically observed with biomedical data is alleviated. Consequently, the drift detection shows a higher performance and reliability, leading to an earlier detection and management of outbreaks. In other words, a higher accuracy of the recognition of new diseases is achieved.

[0017]

[0013] To group initial data of the second, reference database into a characteristic feature pattern, the feature extractor may be configured to add a prototype to each feature pattern. Such prototypes operate as common attraction centers and can avoid that the different domains form sub-clusters and are disconnected.

[0018]

[0014] When comparing the empirical distributions of the grouped features, the comparator may use a two-sample test, such as a Brunner Munzel test, or another sophisticated approach, such as an approach based on kernel density estimation. The aim is to provide a robust and fast test to determine if a drift between two features is statistically relevant. The advantage of the Brunner Munzel (BM) test is that it is a nonparametric test and therefore does not require that some assumptions about an underlying distribution are satisfied. Moreover, such tests are mathematically well- motivated and understood. Similar holds for kernel density estimations (KDE). In general, the tests are kept as simple as possible to ensure that the application is fast and to avoid side-effects of complex methods such as the unknown generalization abilities of deep neural networks.

[0019]

[0015] To obtain the best results in drift detection, it is good practice to update the characteristic feature pattern upon instructions of experts in case of issuance of a report documenting a significant drift between reference and temporal databases. The characteristic feature pattern can be updated either by adding a new diagnosis or by adding new explanations reflecting the drift to the second, reference database.

[0020] When a drift is detected, the drift detector generates a report, which can also be in the form of an alarm, or where the report can be accompanied by an alarm. The report may include the collected explanations of the first, temporal database and the second, reference database, which are relevant to the drift. The report may also include an explanation of the drift detector (drift detector explains why a significant drift was detected).

[0021] Because the drift can have different origins, as it can indicate a new disease, a new diagnosis practice but also for example a failure or a calibration issue with a measuring device, the expert will look at the details of the report and perform a cross-examination to decide whether the second, reference database is to be updated to reflect the drift in diagnosis, or whether the second, reference database needs to be expanded by adding a new diagnosis reflecting the drift. In both cases, based on the output data an automatic report is generated, which can be used to initiate training of doctors, either for the new diagnosis of an existing disease or for a new disease. Moreover, if it is a new disease an automatic quarantine of patients with the new disease can be initiated.

[0022]

[0016] In some embodiments the drift detector can also comprise a communication module which automatically communicates updates of the characteristic feature pattern to relevant stakeholder such that they can update their own databases and adjust their diagnosis practice. This ensures that only the latest information is used in their diagnosis practice. The communicated, automatically generated reports with expert-based examples supports the diagnosis of a new disease by presenting the differences between the expert-knowledge database and the temporal database as identified by the drift detector.

[0023]

[0017] It may happen that the drift detector detects a significant drift between the (first) temporal database and the (second) reference database. To be able to react in a timely fashion, the drift detector can comprise an alarm module configured to raise an alarm. In the case of a new disease, this can prompt immediate quarantine measures, contact tracking or other protective measures.

[0024]

[0018] To reliably detect a drift, in a first step relevant explanations and raw data are collected and compiled in a first, temporal database. Analogously relevant available expert knowledge from reliable sources is collected and compiled in a second, reference database. For a given set of diagnoses a feature extractor is then trained to group the data of the reference database into a characteristic feature pattern. Subsequently, the feature extractor groups data from both the temporal database and the reference database into features by diagnosis. Once the data of both databases is grouped into features by diagnosis, a comparator can compare the empirical distributions of each of the grouped features in the first, temporal database and the second, reference database with each other. If the results of the comparison do not comply with a predefined condition, a processor will generate a report.

[0025] A predefined condition can for example be a defined threshold above which the differences between temporal and reference database are considered to be significant.

[0026] The feature extractor is preferably trained by means of an artificial intelligence (Al), such as machine learning (ML). Once the feature extractor has been trained, it will embed different domains of a given diagnosis, i.e group the features of a given diagnosis close to each other and it will embed different domains, ie. group the features of different diagnoses far away from each other.

[0019] The described drift detector can be used in different situations. Next to the use for which the drift detector was initially developed, namely to recognise a new disease, the drift detector can also be used to improve diagnostic methods by updating medical software, to improve training of doctors or to detect pest or disease outbreaks in crops and forests.

[0027]

[0020] When the drift detector is used to recognise a new disease, explanations and raw data coming from doctors, hospitals and similar, such as for example X-ray images with patches of the images marked by the doctor to be important or an additionally provided text, are collected in the temporal database, and expert knowledge from reliable sources, such as for example online books, international guidelines, textbooks, agreed standards, etc., is collected in the reference database;

[0028]

[0021] When the drift detector is used to update medical software, for example to more accurately recognise polyps, explanations and raw data gathered during a medical visit, such as an endoscopy, are collected in the temporal database. Thereby, the explanations relating to the visit can originate from a medical doctor or a humancentric AL Furthermore, the expert knowledge collected in the reference database comes from the medical software itself, such as for example user agreed reference images for malignant and benign polyps, which are built by endoscopy software for polyp recognition based on agreed standards and which include corresponding explanation.

[0029]

[0022] The drift detector can also be used to adaptively update diagnosis strategies in a database for doctors. Explanations and raw data coming from doctors, hospitals and similar, such as for example medical notes relating to e.g., x-ray scans, for each diagnosis and collected in the temporal database is compared to the expert knowledge, which is collected in the reference database and corresponds to the training databases for the doctors.

[0023] The drift detector can also be employed in a more remote field, which is not related to medical applications, such as to monitor a disease outbreak or an insect pest outbreak in crops and forests. Explanations and raw data coming from current longitudinal observations of an area such as for example satellite images, sensor data, are collected in the temporal database. In such case, the expert knowledge from reliable sources, corresponds to the longitudinal observations of an area at a starting point of time, such as for example satellite images or sensor data, and are collected in the reference database. Preferably the reference database also contains retrospective annotations on whether an outbreak is occurring, and on which features or patterns are indicative of such event.

[0030]

[0024] The various aspects disclosed herein may for example be helpful for hospitals to allow them to divert patients that are suspected of carrying infectious diseases or other urgent symptoms that they are not equipped to handle, and also ensures that they can adjust and prepare in an advance for these situations by having an eagle-eye view of their system. For instance, not all hospitals have the capacity to deal with novel diseases or new variants of locally previously eradicated pathogens arriving from abroad. Instead of posing risk to the hospital staff and other patients, redirecting the infected patient to another facility that would be forewarned of the incoming patient can (a) prevent outbreaks, be (b) lifesaving by assuring prompt and specialized care, (c) ensure a better resource management of staff and medical items.

[0031]

[0025] This invention introduces a drift detector that analyzes whether the data that is examined (e.g., x-ray images for patients) shows signs for a new disease by leveraging purposefully explanations provided by experts such as doctors and comparing them with explanations stored in a reference database created for example from knowledge from public databases such as the internet. The comparison is realized by using techniques known from drift detection methods. By focusing the comparison on the provided explanations, the invention overcomes the problem that the comparison is affected by background noise and clutter in the raw data such that, for example, the drift detection is steered by expert knowledge. SHORT DESCRIPTION OF THE DRAWINGS

[0032]

[0026] In the following, different aspects and preferred embodiments of the disclosure are disclosed by reference to the accompanying figures, which show:

[0033] Fig. 1 Overall architecture of the drift detector;

[0034] Fig. 2 More detailed overall architecture of the drift detector;

[0035] Fig. 3 Feature extractor initializing embedding the different domains;

[0036] Fig. 4 Constructed embedding space used for the drift detector;

[0037] Fig. 5 Overall workflow of the invention.

[0038] DETAILED DESCRIPTION OF EXEMPLARY IMPLEMENTATIONS

[0039]

[0027] In the following, some exemplary embodiments / implementations of the various aspects disclosed herein are described in more detail, with reference to the drawings. Naturally, the computing systems and apparatuses of the present disclosure may employ standard hardware components (e.g., a set of on-premises edge computing hardware and / or cloud-based computing resources connect to each other via conventional wired or wireless networking technology). In some implementations, application-specific hardware (e.g., circuitry for training a ML model and / or circuitry for executing a trained ML for pathogen outbreak prediction and prevention) may also be employed. Further, such computing hardware may be configured to execute software instructions (e.g., retrieved from collocated or remote memoiy circuitry) to execute the computer-implemented methods discussed herein.

[0040]

[0028] While specific feature combinations are described in the following paragraphs with respect to the exemplary embodiments of the present disclosure, it is to be understood that not all features of the discussed embodiments have to be present for realizing the disclosure, which is defined by the subject matter of the claims. The disclosed embodiments may be modified by combining certain features of one embodiment with one or more technically and functionally compatible features of other embodiments. Specifically, the skilled person will understand that features, components, processing steps and / or functional elements of one embodiment can be combined with technically compatible features, processing steps, components and / or functional elements of any other embodiment of the present disclosure as long as covered by the invention as specified by the appended claims.

[0041]

[0029] Moreover, the various embodiments discussed herein can be implemented in hardware, software or a combination thereof. For instance, the various modules of the systems and apparatuses disclosed herein maybe implemented via application specific hardware components such as application specific integrated circuits, ASICs, and / or field programmable gate arrays, FPGAs, and / or similar components and / or application specific software modules being executed on multi-purpose data and signal processing equipment such as CPUs, DSPs and / or systems on a chip, SOCs, or similar components or any combination thereof.

[0042]

[0030] For instance, the various computing (sub)-systems discussed herein may be implemented, at least in part, on multi-purpose data processing equipment such as edge computing servers. Similarly, ML model training subsystems or processes discussed herein may be implemented, at least in part, on multi-purpose cloud-based data processing equipment such as a set of cloud-severs and similar technology.

[0043]

[0031] This invention proposes a drift detector 1 incorporating medical explanations such that the detection of new disease becomes more accurate. For this, the drift detector 1 consists of two databases for each diagnosed disease as shown in Fig. 2. The first database 3, the temporal database, is continuously fed with explanations 3a and raw data 3b for a given diagnosis 3c coming from doctors 110, hospitals 120 and so on. The second database 5, the reference database, is automatically created based on available expert knowledge 150, 160. This could exemplarily be done by collecting raw data 5b and their explanations 5a for the given diagnosis 5c, such as a diagnosed disease, automatically from the internet.

[0032] As shown in Fig.i in a summarized overview of the architecture of the drift detector showing the working principle, after explanations 3a, 5a and raw data 5a, 5b are fed to the drift detector 1, the drift detector 1 processes the data and issues a report 81 when a drift is detected.

[0044]

[0033] In Fig. 2, a more detailed view if the architecture of the drift detector is shown. It can be seen that the temporal database 3 is fed with data collected by a first collecting device 2 and the reference database 5 is fed with data collected by a second collecting device 4. Once the data is collected, a learned feature extractor 6 separates the data into feature patterns 631, 63n, 651, 6511 before storing it in the respective database 3, 5.

[0045]

[0034] Then, a comparator 7 analyzes in a temporal, recurrent analysis, for each diagnosed disease 1 to n, the feature patterns 631, 63n, 651, 6511 stored in the databases 3, 5 on the explanation level for significant temporal changes, meaning drifts, in order to recognize early outbreaks of new diseases. By focusing the drift detection on the available and provided explanations, the general problem that drift detection must differentiate between information that is relevant to the accurate diagnosis of the disease, and background noise or confounders in the raw data, is overcome. In turn, this specific focus on the meaningful signal in the data leads to an increased reliability in the recognition of new diseases.

[0046]

[0035] To provide an accurate new disease recognition, the drift detector considers all possible data modalities coming from different domains di to dm, for explanations 3a, 5a. For instance, if the raw data 3b are images (data modality) coming from an x-ray machine (domain), the explanations 3a could be patches of the images marked by the doctor to be important (corresponding for example to a first data modality from a given domain di) or an additionally provided text (corresponding for example to a second data modality from a given domain d2). All these different data modalities coming from different domains di to dm, are considered by the comparator 7 for the drift detection.

[0036] Moreover, because the comparator 7 does the drift detection on the explanation level, it is possible to include expert knowledge from the literature or state-of-the-art (e.g., medical textbooks 150), in the form of reference explanations 5a (e.g., images) prototypical of certain diseases 1 to n. This knowledge is included in the reference database 5 and further improves the recognition accuracy of new diseases n+i and ensures that the reference databases 5 are initialized with appropriate and reliable information when embedding features of the different diagnosis in the initial embedding space.

[0047]

[0037] The result of the comparison between the data of the reference database with the data of the temporal database can be that there is no drift or that there is a drift. When a drift is detected, the processor 8 will generate a report 81. This report 81 is made available to experts to receive their feedback 9. The report contains statistically significant explanation differences detected by the drift detector in the explanations 3a and highlighted drifted regions of the raw data 3b. This way, the data can easily be cross-examined by the experts such as the doctors, to decide whether the updates 91 of the reference database should be and update of a certain diagnosis or an addition of a new disease.

[0048]

[0038] As mentioned above, the second collecting device 4 automatically collects expert-knowledge 150, 160 from the internet by collecting information 5a, 5b, 5c from reliable sources (e.g., online books 150, international guidelines 160 or similar) and compiles the collected expert-knowledge 150, 160 into the reference database 5. When the collecting is initially done, it provides the data for the initial reference database 5-i, with which the feature extractor 6 of the drift detector is trained, as can be schematically seen in Fig. 3. The data collected for the initial reference database is separated into features per domain (such as patches in X-ray images or explanatory test) and grouped into characteristic feature patterns 651-i, 652-i, ..., 65n-i for each monitored disease 1 to n. To obtain the characteristic feature patterns 651-i, 652-i, ..., 6sn-i, the features separated by domains are grouped with the help of a modality specific feature extractor 61, 62, usually a neural network. For instance, a first specific feature extractor 61 can be a convolutional neuronal network that processes images, and second specific feature extractor 62 can be a transformer architecture to process text. Both specific feature extractors 61 and 62 map into the same embedding space. In this embedding space, prototypes 60 are added as attraction centers to have an optimization target: the learned feature representations of disease 1 should be close to a first prototype 60, the learned feature representation of disease 2 close to a second prototype 60, ... Together, all specific feature extractors 61, 62, ... define the feature extractor 6.

[0049]

[0039] Even though in Fig. 2 the feature extractor 6 is represented twice through two distinct arrows, the feature extractor 6 is preferably shared among both the databases and the diagnoses.

[0050] To share the feature extractor 6 amongst diagnoses, it is preferably trained by a triplet loss where negative and positive samples are presented, such that the embedding of different data modalities coming from different domains di, ..., dm of a given diagnosis (such as a disease) becomes close in the shared embedding space. At the same time, when training the feature extractor 6, the embedding of different diagnoses (such as different diseases) should be far away. Presenting negative and positive samples during training forces the specific feature extractors to learn meaningful, good representations, thereby avoiding a potential collapse of the embedding, for example in which everything is projected to a zero vector.

[0051] At the same time the specific feature extractors 61, 62, ... are shared between the databases to allow a comparison of the distributions during the drift detection.

[0052]

[0040] Fig. 4 visualizes working principle of the feature extractor 6 when characteristic feature patterns are initially formed for each monitored diagnosis. It shows how the different domains di to dm are embedded in an embedding space, that is to say, how the features are grouped. In figure 4, different colors represent different diseases 1 to n. Thus, the different domains di to dm are grouped by diseases 1 to n. The training of such an embedding space can be realized by concepts known from multi-modal domain alignment studied in the field of vision-language representation learning [5], where prototypes 60 of each diagnosis 1 to n are purposefully added to the embedding space to avoid that the embedding of different domains form sub-clusters such that they are disconnected. The prototypes 6o operate as common attraction centers for the domain embedding. It becomes apparent that all features belonging to a certain diagnosis are grouped around the prototypes 6o and that the domains of the grouped data (such as x- rays, text, etc.) can differ in the characteristic feature pattern of each diagnosis.

[0053]

[0041] For the drift detection, the learned feature extractor 6 is used to embed the data coming from both the first, temporal database 2 and the second, reference databases 4. Then, the comparator 7 compares the empirical distributions 651, ..., 6511 of the embedded explanations 5a in the reference database 5, or the grouped features 651, 652 within the reference database 5 for a given diagnosis, with the empirical distributions

[0054] 631, 632, ..., 63n of the embedded explanations 3a in the temporal databases 3, or the grouped features 631, 632, ..., 63n within the temporal database 3 for said given diagnosis, using a two-sample test (e.g., Brunner Munzel test) or another sophisticated approach (e.g., based on kernel density estimation).

[0055]

[0042] If the test statistics do not comply with a predefined condition, for example if the test statistics exceed a defined threshold (and assuming that the measure is low in case of no drift), the two distributions 631, 651; 632, 652; ...; 63n, 6511 drifted significantly. If the two distributions 631, 651; 632, 652; ...; 63n, 6511 drifted significantly, the proposed drift detector 1 is used to generate a report 81 with the processor 8. In this report 81, the difference between the two distributions 631, 651;

[0056] 632, 652 can be highlighted by presenting the contrasting explanations and the corresponding available raw data. These reports 81 are generated upon the collected expert knowledge and are therefore reliable sources to investigate the differences. Moreover, the detected drift can be used to raise an alarm that there might be an outbreak of a new disease such that actions can be taken immediately.

[0057]

[0043] The generated reports 81 are used by experts to determine if the significant drift corresponds to a new disease or if a drift in diagnosis of a known disease has taken place. Based on the feedback of experts 9, the system is updated 91 accordingly. If the occurrence of a new disease is confirmed, a new diagnosis will be added to the existing ones in the reference database 5. Otherwise, the new explanations 5a will be added to the reference database 5 of the corresponding diagnosis to reflect the drift. Those updates 91 can also be automatically communicated to the healthcare professionals so that they can update their own domain knowledge and adjust their diagnosis practice.

[0058]

[0044] In Fig. 5 the overall workflow shows a preferred sequence of the different components of the drift detector. The workflow starts with identification of diagnoses 201, 202, ..., 2on to be monitored. For each diagnosis to be monitored, the first collecting device 2 and the second collecting device 4 collect relevant explanations, raw data and expert knowledge to feed the feature extractor 6. The feature extractor groups the data into characteristic feature patterns by diagnosis to be able to compare them with a comparator 7. Once the comparator has compared the grouped feature patterns with each other, it will determine if one grouped feature pattern shows a drift with respect to the other grouped feature pattern. If a drift is detected, a processor 8 will generate a report 81 containing all relevant information. Such a report is preferably submitted to an expert, who will provide feedback 9 on whether the reference database 5 describing the diagnoses 201, 202, ..., 2on to be monitored should be updated, either by updating the content of a monitored diagnosis of by adding a new diagnosis 2on+i.

[0059]

[0045] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the aspects to the precise form disclosed. Modifications and variations may be made in light of the above disclosure or may be acquired from practice of the aspects. As used herein, the term component is intended to be broadly construed as hardware, firmware, or a combination of hardware and software. As used herein, a processor is implemented in hardware, firmware, or a combination of hardware and software.

[0060]

[0046] It will be apparent that systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the aspects. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code— it being understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0061]

[0047] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of various aspects. In fact, many of these features maybe combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of various aspects includes each dependent claim in combination with every other claim in the claim set. A phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, ac- c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c).

[0062]

[0048] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items and maybe used interchangeably with “one or more.” Furthermore, as used herein, the terms “set” and “group” are intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, and / or the like), and may be used interchangeably with “one or more.” Where only one item is intended, the phrase “only one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” and / or the like are intended to be open-ended terms.

[0063]

[0049] As used herein, the phrase “based on” shall not be construed as a reference to a closed set of information, one or more conditions, one or more factors, or the like. In other words, the phrase “based on A” (where “A” may be information, a condition, a factor, or the like) shall be construed as “based at least on A” unless specifically recited differently.

[0050] As used herein, the term “or” is an inclusive “or” unless limiting language is used relative to the alternatives listed. For example, reference to “X being based on A or B” shall be construed as including within its scope X being based on A, X being based on B, and X being based on A and B. In this regard, reference to “X being based on A or B” refers to “at least one of A or B” or “one or more of A or B” due to “or” being inclusive. Similarly, reference to “X being based on A, B, or C” shall be construed as including within its scope X being based on A, X being based on B, X being based on C, X being based on A and B, X being based on A and C, X being based on B and C, and X being based on A, B, and C. In this regard, reference to “X being based on A, B, or C” refers to “at least one of A, B, or C” or “one or more of A, B, or C” due to “or” being inclusive. As an example of limiting language, reference to “X being based on only one of A or B” shall be construed as including within its scope X being based on A as well as X being based on B, but not X being based on A and B.

[0064] USE CASES AND ADDITIONAL APPLICATION SCENARIOS

[0065] New disease detection

[0066]

[0051] The described drift detector can be used to automatically scan for new diseases. To this end, the drift detector should preferably receive data input 3a, 3b, 3c and 5a, 5b, 5c on a global scale. The input collected by the first collecting device 2 to populate the first, temporal database should be at least medical notes, such as for example x-ray scans, diagnosis, and explanations. The input for the second, reference database is collected from textbook, agreed standards, etc. for each diagnosis. The comparator 7 of the drift detector 1 analyses the provided data for drift on the explanation lever for each disease and raises an alert if a drift is detected, as this means the detection of a potential new disease. The processor 8 can then generate a report summarizing the explanations and the raw data of the detected drift such that a cross examination can be performed, or automatic safety routines can be initiated. Such a report can be analysed automatically so that an automatic quarantine of patients with the new disease can be initiated. Preferably, after an automatic analysis, the report is further analysed by experts, and automatic measures are adapted if need be.

[0067] Updating of medical software

[0052] Another use case for the described drift detector 1 is to detect when medical software must be updated, thus improving the functioning of the software. Such a software can be for example the endoscopy software for polyp recognition, which is building user agreed reference images for malignant or benign polyps. These reference images are presented to the doctor during endoscopy to support the decision on whether the polyp is malignant or benign. Such reference images of malignant and benign polyps, which are based on agreed standards and which include the corresponding explanations, are stored in the reference database 5. The first database 3 stores explanations of polyp images coming from an expert, which are gathered during endoscopy. For example, the doctor or a human-centric computer vision-based method can mark regions that may explain the decision on whether a polyp is malignant or benign.

[0068]

[0053] The drift detection system compares the reference explanations with the temporal explanations and raises an alert if a drift is detected, as this could mean there is a new type of polyps or there is a system error, stemming for example from aging light sources. If a drift is detected, the drift detector furthermore generates a report 81 to explain, the detected drift. This report can be used to check the functionality of the endoscopy system or can be used to retrain the doctors. Furthermore, the report can be analysed automatically so that a software update can also be initiated automatically.

[0069] Adaptive training database for doctors

[0070]

[0054] Currently, training databases for doctors are rarely updated. A resulting problem is that, while doctors refine and adapt their personal diagnosis strategies over time with increasing experience, sometimes even unconsciously, this acquired knowledge is not necessarily reflected in the training databases. A drift detector 1 can be deployed on a global scale to analyse the training databases and the different diagnoses for drifts. If a drift is established and the drift is not a new disease, it is likely that the diagnosis strategy of a majority of doctors has changed, thereby revealing a swarm intelligence. The drift detector can help detect those drifts on explanation level and raise an alert if a drift is detected. The drift detector 1 can further be configured to automatically update the diagnosis best practice. Explanations 3a, 5a and the raw data 3b for the detected drift are summarized in an automatically generated report 81. Such a report 81 can furthermore be automatically sent to doctors to inform them about the new diagnosis strategy. It can also be the starting point to organise a more formal training of doctors about the changed diagnosis strategy.

[0071]

[0055] Here, the input to the first database 3 of the drift detector mainly consists of raw data 3b, such as medical notes (e.g., x-ray scans), diagnosis 3c and explanations 3a. For each diagnosis, the second database 5 is populated with reference explanations 5a collected from textbooks 150, agreed standards 160, etc...

[0072] Detection of disease and insect pest outbreaks in crops and forests

[0073]

[0056] The fourth use case exemplifies an application of the drift detector beyond the medical domain. The drift detector can be employed to monitor agricultural surfaces, forests and the like. Notably, the changing climate is increasing the risks of such events, with the apparition of diseases and pests that were previously foreign to the region of interest. When not detected and treated in time, pests and diseases, and in particular previously unknown pests and diseases can have devastating and long-lasting effects on crops and forests. Coincidentally to the risks of new pests and diseases due to the climate change, the growing availability of high frequency satellite images and sensor data opens the door to an accurate and continuous automatic monitoring of crops and forest.

[0074]

[0057] The reference data 5 base contains the longitudinal observations of an area, such as for example satellite images or sensor data, at a starting point of time. It can be useful to further provide the reference data base with retrospective annotations on whether an outbreak is occurring, and on which features or patterns are indicative of such event. The drift detector analyses the provided data for drifts by comparing the data at said starting point with the current data to detect whether an outbreak is currently occurring. If the data is provided with sufficient granularity, not only the outbreak itself can be detected but also if the outbreak is known or new. In the former case, the outbreak might already have been treated successfully with a specific treatment. In the latter case, when the outbreak of a disease or insect pest that was not previously observed in the area, the detection of a drift will prompt the need for finding new solutions.

[0058] As output to the drift detection, a report can be automatically sent to the farmers or rangers, informing them about the risks of an ongoing outbreak and, when possible, giving them recommendations on the course of action to take, based on past occurrences of a similar outbreak.

[0075] LIST OF REFERENCE SIGNS drift detector

[0076] 2 first collecting device

[0077] 3 first database

[0078] 3a explanations of the first database

[0079] 3b raw data of the first database

[0080] 3c monitored diagnosis corresponding to the explanations and the raw data of the first database

[0081] 4 second collecting device

[0082] 5 second database

[0083] 5a explanations of the second database

[0084] 5a-i initial explanations of the second database

[0085] 5b raw data of the second database

[0086] 5C monitored diagnosis corresponding to the explanations and the raw data of the second database

[0087] 6 feature extractor

[0088] 61 first specific feature extractor

[0089] 62 second specific feature extractor

[0090] 60 prototype for each diagnosis

[0091] 631. ..., 63n first feature pattern for each diagnosis

[0092] 651. ..., 6511 second feature pattern for each diagnosis

[0093] 651-i, 65n-i initial characteristic feature pattern for each diagnosis

[0094] 7 comparator

[0095] 8 processor

[0096] 81 report

[0097] 9 feedback of experts

[0098] 91 updates 110, 120 instance producing temporal data

[0099] 110 doctors

[0100] 120 hospital

[0101] 150, 160 expert knowledge 150 books

[0102] 160 guidelines

[0103] 201, ..., 20n diagnosis to be monitored di, ..., dm domains, or data modalities, for the explanations

[0104] 1, ..., n number of the diagnosis to be monitored

Claims

Claims1. Drift detector (i) to detect a drift in features of a diagnosis, the drift detector comprising a first collecting device (2) configured to collect relevant explanations (3a) and raw data (3b), a first database (3) compiling the relevant explanations (3a) and raw data (3b) collected by the first collecting device, a second collecting device (4) configured to collect relevant available expert knowledge (5a) from reliable sources (5c), a second database (5) compiling the relevant available expert knowledge from reliable sources collected by the second collecting device, a learned feature extractor (6) configured to group initial data of the second database into a characteristic feature pattern (651-i, 652-i) for different diagnoses and to subsequently group the data from the first database into first feature patterns (631, ..., 63n) by diagnosis and from the second database into second feature patterns (651, ..., 6sn) by diagnosis. a comparator (7) configured to compare the empirical distributions of each of the grouped feature patterns in the databases with each other, a processor (8) configured to generate a report (81) if results of the comparison do not comply with a predefined condition.

2. Drift detector according to claim 1, wherein the first database (3) is configured to continuously receive data from users, and the second database (5) is configured to retrieve its data automatically.

3. Drift detector according to claim 1 or 2, wherein the feature extractor (6) is configured to add a prototype (6o) to each characteristic feature pattern, the prototypes operating as common attraction centers.

4. Drift detector according to any preceding claim, wherein the comparator (7) is configured to use a two-sample test, such as a Brunner Munzel test, or another sophisticated approach, such as an approach based on kernel density estimation, to compare the empirical distributions of the grouped features in both databases.

5. Drift detector according to any preceding claim, configured to update (91) the characteristic feature pattern upon instructions of experts in case of issuance of a report documenting a significant drift between reference and temporal databases, wherein updating the characteristic feature pattern can be either adding a new diagnosis or new explanations reflecting the drift to the reference database.

6. Drift detector according to claim 5, further comprising a communication module configured to automatically communicate the updates to relevant stakeholder such that they can update their own databases and adjust their diagnosis practice.

7. Drift detector according to any preceding claim, further comprising an alarm module configured to raise an alarm in case of significant drift between reference and temporal databases8. Method for detecting a drift, the method comprising the steps of- collecting relevant explanations and raw data and compiling them in a first, temporal database (3);- collecting relevant available expert knowledge from reliable sources and compiling them in a second, reference database (5);- training a feature extractor (6) with the data of the reference database to group the data into an initial characteristic feature pattern (651-i, 652-i, 65n-i) for a given set of diagnoses,- grouping the data from both the reference (5) and the temporal (3) databases into features by diagnosis by means of the feature extractor (6);- comparing the empirical distributions of each of the grouped features in the reference and the temporal databases with each other by means of a comparator (7),- generating a report (81) if results of the comparison do not comply with a predefined condition.

9. Method according to claim 8, wherein the temporal database is continuously fed, and the reference database is automatically created.

10. Method according to claim 8 or 9, created wherein training the feature extractor includes adding a prototype (60) to the characteristic feature pattern that operates as common attraction centers.

11. Method according to any preceding claim, wherein the comparison between the empirical distributions of the grouped features in the first database and the second database is made using a two-sample test, such as a Brunner Munzel test, or another sophisticated approach, such as an approach based on kernel density estimation.

12. Method according to any preceding claim, further comprising the step of updating the characteristic feature pattern (651, 652, ..., 6sn) upon feedback of experts to the report in case of significant drift between first (3) and seconddatabases (5), wherein either a new diagnosis or new explanations reflecting the drift are added to the reference database.

13. Method according to claim 12, wherein after updating the characteristic feature pattern, the updates are automatically communicated to relevant stakeholder such that they can update their own domain knowledge and adjust their diagnosis practice.

14. Method according to any preceding claim, further comprising the step of raising an alarm in case of significant drift between reference and temporal databases15. Method according to any preceding claim, wherein, a) when the method is used to recognise a new disease,- explanations and raw data coming from doctors, hospitals and similar, such as for example X-ray images with patches of the images marked by the doctor to be important or an additionally provided text, are collected in the temporal database, and- expert knowledge from reliable sources, such as for example online books, international guidelines, textbooks, agreed standards, etc., is collected in the reference database; or b) when the method is used to update medical software,- explanations and raw data gathered during a medical visit, such as an endoscopy, whereby explanations relating to the visit can originate from a medical doctor or a human-centric Al, are collected in the temporal database, and- expert knowledge from medical software, such as user agreed reference images for malignant and benign polyps, which are built by endoscopy software for polyp recognition based on agreed standards and which include corresponding explanations, is collected in the reference database; orc) when the method is used to adaptively update diagnosis strategies in a database to train doctors,- explanations and raw data coming from doctors, hospitals and similar, such as for example medical notes relating to e.g., x-ray scans, for each diagnosis are collected in the temporal database, and- expert knowledge from training databases for doctors is collected in the reference database; or d) when the method is used to monitor a disease outbreak or an insect pest outbreak in crops and forests, - explanations and raw data coming from current longitudinal observations of an area, such as for example satellite images or sensor data, are collected in the temporal database, and- expert knowledge, i.e. the longitudinal observations of an area at a starting point of time, such as for example satellite images or sensor data, is collected in the reference database, preferably together with retrospective annotations on whether an outbreak is occurring, and on which features or patterns are indicative of such event.