A computer-implemented method for root cause analysis

A computer-implemented method using machine learning for medical devices automates root cause analysis by analyzing device data with TF-IDF and trained models, addressing inefficiencies and errors in manual methods, and enabling active feedback.

WO2026104321A1PCT designated stage Publication Date: 2026-05-21ROCHE DIABETES CARE GMBH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ROCHE DIABETES CARE GMBH
Filing Date
2025-11-10
Publication Date
2026-05-21

Smart Images

  • Figure EP2025082388_21052026_PF_FP_ABST
    Figure EP2025082388_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method and an investigation system for root cause analysis for investigation of a medical device is disclosed. The method comprises the following steps: i) retrieving feature data of the medical device by using at least one communication interface; ii) applying term frequency-inverse document frequency (TF-IDF) on the feature data and extracting a feature vector therefrom, iii) applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Roche Diabetes Care GmbH

[0002] A computer-implemented method for root cause analysis

[0003] Technical Field

[0004] The invention refers to a computer-implemented method for root cause analysis for investigation of a medical device, an investigation system configured for root cause analysis for investigation of a medical device, a computer program, a computer-readable storage medium, a non-transient computer-readable medium and several uses. The invention generally refers to the topic of root cause analysis methods.

[0005] Background art

[0006] In the field of root cause analysis, methods for medical devices usually require actions such as sending the defective device to the manufacturer or to another service provider, operations by an investigator comprising read out of text files from a memory of the device, and manual screening of the text files in view of maintenance, warning or error messages. Firstly, such root cause analysis methods are time consuming. Moreover, such root cause analysis methods are error-prone due to several factors. For example, the manual screening of large amount of data by the human investigator can lead to situations of finding a few relevant events in a mass of irrelevant events which is like asking to find the proverbial needle in a haystack. Moreover, classifying events can be nontrivial and may strongly relate to the skills of the investigator. In case of several root causes, the situation may get even more complicated. Moreover, such methods are only retrospectively and cannot be used for active feedback in the market.

[0007] In other technical fields, use of machine learning for root cause analysis is known, e.g. from US 11,176,464 Bl. However, these methods are not transferrable to root cause analysis for a medical device due to specific data of the medical device and specific requirements in view of product risk, health safety, regulatory requirements and the like for products on the medical field.

[0008] Problem to be solved

[0009] It is therefore desirable to provide methods and devices which address the above-mentioned shortcomings of known methods and devices. Specifically, methods and systems shall be proposed which allow for reducing susceptibility to errors due to manual action, providing the possibility for active feedback and increasing user friendliness for performing root cause analysis.

[0010] Summary

[0011] This problem is addressed by a computer-implemented method for root cause analysis for investigation of a medical device, an investigation system configured for root cause analysis for investigation of a medical device, a computer program, a computer-readable storage medium, a non-transient computer-readable medium and several uses with the features of the independent claims. Advantageous embodiments which might be realized in an isolated fashion or in any arbitrary combinations are listed in the dependent claims as well as throughout the specification.

[0012] As used in the following, the terms “have”, “comprise” or “include” or any arbitrary grammatical variations thereof are used in a non-exclusive way. Thus, these terms may both refer to a situation in which, besides the feature introduced by these terms, no further features are present in the entity described in this context and to a situation in which one or more further features are present. As an example, the expressions “A has B”, “A comprises B” and “A includes B” may both refer to a situation in which, besides B, no other element is present in A (i.e. a situation in which A solely and exclusively consists of B) and to a situation in which, besides B, one or more further elements are present in entity A, such as element C, elements C and D or even further elements.

[0013] Further, it shall be noted that the terms “at least one”, “one or more” or similar expressions indicating that a feature or element may be present once or more than once typically will be used only once when introducing the respective feature or element. In the following, in most cases, when referring to the respective feature or element, the expressions “at least one” or “one or more” will not be repeated, non-withstanding the fact that the respective feature or element may be present once or more than once.

[0014] Further, as used in the following, the terms "preferably", "more preferably", "particularly", "more particularly", "specifically", "more specifically" or similar terms are used in conjunction with optional features, without restricting alternative possibilities. Thus, features introduced by these terms are optional features and are not intended to restrict the scope of the claims in any way. The invention may, as the skilled person will recognize, be performed by using alternative features. Similarly, features introduced by "in an embodiment of the invention" or similar expressions are intended to be optional features, without any restriction regarding alternative embodiments of the invention, without any restrictions regarding the scope of the invention and without any restriction regarding the possibility of combining the features introduced in such way with other optional or non-optional features of the invention.

[0015] In a first aspect, a computer-implemented method for root cause analysis for investigation of a medical device is provided.

[0016] The method comprises the following method steps, which, as an example, may be performed in the given order. However, a different order is also feasible. Further, it is possible to perform two or more of the method steps simultaneously or in a fashion overlapping in time. Further, it is also possible to perform one, more than one or even all of the method steps repeatedly.

[0017] The method comprises the following steps

[0018] i) retrieving feature data of the medical device by using at least one communication interface;

[0019] ii) applying term frequency -inverse document frequency (TF-IDF) on the feature data and extracting a feature vector therefrom,

[0020] iii) applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause.

[0021] The method may comprise determining at least one root cause of an unknown complaint of a medical device, e.g. by means of machine learning based on completed complaints in a service investigation process. The present invention can allow for reliable root cause analysis for medical devices using machine learning features. The proposed extraction of a feature vector can allow considering a whole device history, e.g. as present in log files relating to the medical device. The method can cover root causes that are related to significant irregularities in this data. The machine learning can comprise considering domain knowledge e.g. for feature extraction, allowing for reduced effort but high accuracy of prediction results. Specifically, the determined root cause is independent from the patient’s allegation.

[0022] The term “computer-implemented” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a method involving at least one computer and / or at least one computer network. The computer and / or computer network may comprise at least one processor which is configured for performing at least one of the method steps of the method according to the present invention. Preferably each of the method steps is performed by the computer and / or computer network. The method may be performed completely automatically, specifically without user interaction. The term “automatically” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a process which is performed completely by means of at least one computer and / or computer network and / or machine, in particular without manual action and / or interaction with a user.

[0023] The term “medical device” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a device configured for performing at least one medical function. The medical device may be configured for performing at least one diagnostic purpose and / or at least one therapeutic purpose. The medical device specifically may comprise an assembly of two or more components capable of interacting with each other, such as in order to perform one or more diagnostic and / or therapeutic purposes, such as in order to perform the medical analysis and / or the medical procedure. For example, the medical device may be configured for qualitatively and / or quantitatively detecting at least one analyte in a bodily fluid, such as in a body fluid contained in a body tissue of a user. The medical device may be configured for continuous or non-con-tinuous monitoring of at least one analyte in a bodily fluid. For example, the two or more components of the medical device may be capable of performing at least one detection of the analyte in the bodily fluid and / or in order to contribute to the detection of the analyte in the bodily fluid. The medical device generally may also be or may comprise at least one of a sensor assembly, a sensor system, a sensor kit or a sensor device. For example, the medical device may comprise at least one analyte sensor, e.g. at least one at least partially insertable analyte sensor. The analyte sensor may be or may comprise at least one electrochemical sensor and / or at least one optical sensor. The term “analyte” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an arbitrary element, component or compound which may be present in a body fluid and the concentration of which may be of interest for a user. Specifically, the analyte may be or may comprise an arbitrary chemical substance or chemical compound which may take part in the metabolism of the user, such as at least one metabolite. As an example, the at least one analyte may be selected from the group consisting of glucose, cholesterol, triglycerides, lactate, ketone. Additionally or alternatively, however, other types of analytes may be determined and / or any combination of analytes may be determined. The medical device may comprise further elements such as at least one processing unit and / or at least one control unit and / or at least one communication interface.

[0024] For example, the medical device may comprise at least one medication device such as a medication pump, e.g. in addition or alternatively to the analyte sensor described above. The medical device may comprise at least one infusion kit, e.g. comprising at least one infusion cannula and the medication pump. The term “infusion cannula” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an arbitrary cannula being configured to introduce an infusion, i.e. a liquid substance, specifically a liquid substance comprising a medicine, into the body tissue, exemplarily directly into a vein of a patient. Therefore, the infusion cannula may be attached to a reservoir comprising the liquid substance, specifically via the ex vivo proximal end of the infusion cannula. The infusion cannula may be part of the infusion kit. The infusion kit may be configured for a conduction of an arbitrary infusion.

[0025] For example, the medical device is a device selected from the group consisting of: at least one insulin pump, a glucose meter, a home use device for analyte measurement. The glucose meter may be or may comprise at least one continuous glucose sensor. The home use device for analyte measurement may be a wearable device.

[0026] The term “root cause” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to at least one origin of at least one event which occurred during usage of the medical device. The term “usage of the medical device” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to any action after production of the medical device such as operation, storage, handling, transport and the like. The root cause may be at least one technical and / or physical cause such as malfunction of at least one physical component of the medical device, such as hardware failure and equipment malfunction. The root cause may be at least one human cause, e.g. occurred due to maloperation. The root cause may be a software problem or failure. For example, potential root causes may be one or more of a malfunction of at least one element of the medical device such as of at least one sensor of the medical device; a maloperation; at least one hardware problem or failure; at least one software problem or failure. The root cause may be a precursor appearing before the malfunction and / or failure.

[0027] The term “event” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a situation characterized by dynamics and / or at least one change such as in status and / or condition. The event may relate to a transition from one state to another. The event may be a history event. The history event may be an event occurred during usage of the medical device. The event may be an event relating to the medical function of the medical device and / or to any other functions of the medical device such as communication.

[0028] For example, the event may be a physical and / or technical event. For example, the event may be an event related to at least one user action and / or to usage of the medical device. For example, the event may be or may relate to an error warning and / or a reminder. For example, the event may be one or more of switching on the medical device; switching off the medical device; switching on or off of an element of the medical device, e.g. the communication interface; change in settings of the medical device; change in configuration of the medical device; change of a battery; a change in battery charge status such as a change in voltage of a battery , discharge detected or limited battery voltage detected (e.g. < IV and / or runtime < 30 minutes); change in coupling to a further device; an electronic error such as related to an interrupted buzzer or a corrupted response; detection of communication problems; a maintenance action; detection of an occlusion; detection of missing of an element such as of a removed reservoir; detection of an inserted reservoir; obtaining measurement data such as bolus values; software error detected; detecting problems with data storage; at least one user input e.g. a complaint.

[0029] The event may comprise information stored in at least one database. Said information may comprise one or more of: a time stamp indicating the time when the event occurred; information which event occurred such as a name and / or type; information about additional conditions of the medical device at which the event occurred. The medical device may comprise software configured for generating information about the event. The generation of the information of the event may be triggered by occurrence of the event. The information about the event may be generated in machine readable format.

[0030] The term “root cause analysis” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to at least one process of identifying at least one root cause. The root cause analysis may comprise determining how at least one event, in particular a problematic event, occurred. The problematic event may be an event whose occurrence should be prevented, e.g. for ensuring quality of the medical device, such as an error, malfunction, failure and the like.

[0031] The term “investigation of a medical device” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to studying functionality of the medical device such as for quality assurance. The method may be part of a maintenance process and / or at a complaint handling and / or may be performed at the end of life of the medical device and / or may be performed in reaction to user complaints and / or in regular intervals or continuously.

[0032] Step i) comprises retrieving feature data of the medical device by using at least one communication interface.

[0033] The term "communication interface" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an item or element forming a boundary configured for transferring information. The communication interface may be configured for transferring information from a computational device, e.g. a computer, such as to send or output information, e.g. onto another device. Additionally or alternatively, the communication interface may be configured for transferring information onto a computational device, e.g. onto a computer, such as to receive information. The communication interface may provide means for transferring or exchanging information. The communication interface may provide a data transfer connection, e.g. Bluetooth, NFC, inductive coupling or the like. As an example, the communication interface may be or may comprise at least one port comprising one or more of a network or internet port, a USB-port and a disk drive. The communication interface may be at least one web interface.

[0034] The term “retrieve” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to the process of obtaining data from a data source. The data source may, however, vary, in accordance with the specific application. Thus, in the context of the present invention, the retrieving of the feature data may take place by at least one of the following: downloading the feature data from at least one data source, such as from at least one data storage device of the medical device and / or from a web- or cloud-based data storage device; obtaining the feature data via at least one computer network, such as the Internet; obtaining the feature data via at least one wire-based and / or wireless interface. The retrieving may fully or partially take place automatically, such as by automatic download, and / or may fully or partially take place manually. Semi-automatic retrieving processes are also possible.

[0035] The term “feature data” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to data comprising information about at least one measureable property of the medical device that is being analyzed and / or data relevant for root cause analysis, e.g. indicative of a root cause under investigation. For example, the feature data may comprise one or more of: at least one event; at least one event sequence with temporal relationship; at least one loop of events and / or at least one event parameter. For example, the feature data may be or may comprise at least one log file relating to the medical device, e.g. generated by the medical device. The log file may comprise a list of events that occurred during operation of the medical device such as one or more of at least one problem, at least one error, operational information, at least one warning. Each type of event may have a corresponding predefined term which is used for generating the log file. The terms of the corresponding events occurred during operation may be recorded and listed in the log file. The possible terms which can occur in the log file may be pre-known. The knowledge about possible terms of the log file can be used for evaluation of the log file, in particular in step ii).

[0036] The term “sequence of events” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a plurality of events occurred successively along a predefined time chain. The plurality of events may comprise at least two events, preferably three events, more preferably four events, most preferably more than four events. The sequence of events may be a combination of predefined events having a temporal relationship. The sequence of events may be complete or incomplete, e.g. one or more of the events of the sequence of events may be missing. The events of the sequence of events may be of different type and / or name. The events of the sequence of events may be predefined events. The sequence of events may relate to a plurality of events occurred in a predefined order and / or with an arbitrary order but occurred during a predefined time range.

[0037] The term “loop of events” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to reoccurring events of the same type and / or name. The first occurrence of the event is considered as start event and start of the loop. The length of a loop, e.g. time duration and / or number of events per loop, may be predefined or may be open, e.g. may be determined by counting the number of occurrences.

[0038] The term “event parameter” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to any information specifying a status of the medical device. For example, the event parameter may be selected from the group consisting of temperature (e.g. 4°C), power status or battery voltage (e.g. 2V).

[0039] For example, the feature data comprise data selected from the group consisting of at least one error message, at least one error warning, therapy data, measurement data, further sensor data of at least one further sensor of the medical device such as temperature sensor, information about user interaction, information about a device status such as a power status and / or battery voltage, status of one or more physical sensors, or accessibility to one or more elements of the device, device functional data such as currents, system status of the medical device or of at least one element of the medical device. The feature data further may comprise fitness and / or health data, e.g. provided by input of a user or obtained from a further device, e.g. a smartphone, a cloud or the like.

[0040] The method may comprise uploading the feature data to a database. The uploading may be performed when the medical device is under investigation at a complaint handling and / or investigation unit, and / or at the end of the lifetime of the medical device and / or in regular intervals or continuously. For example, the uploading may be performed wirelessly, e.g., in regular intervals or continuously when the medical device is in use by the patient.

[0041] The database may be at least partially cloud-based and / or a local database of the medical device and / or an external the database, for example of a least one portable communication device such as a smartphone. The term "database" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an organized collection of data, generally stored and accessed electronically from a computer or computer system. The database may comprise or may be comprised by a data storage device. The database may comprise at least one data base management system, comprising a software running on a computer or computer system, the software allowing for interaction with one or more of a user, an application or the database itself, such as in order to capture and analyze the data contained in the database. The database management system may further encompass facilities to administer the database. The database, containing the data, may, thus, be comprised by a data base system which, besides the data, comprises one or more associated applications. The database may be part of the processing device or may be external to the processing device. The processing device and / or the database may be at least partially cloud-based. The term “cloud-based” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an outsourcing of the processing device or of parts of the processing device to at least partially interconnected external devices, specifically computers or computer networks having larger computing power and / or data storage volume. The external devices may be arbitrarily spatially distributed. The external devices may vary over time, specifically on demand. The external devices may be interconnected by using the internet. The external devices may each comprise at least one communication interface. The feature data may comprise machine-readable data. For example, the machine-readable data is textual data or binary data.

[0042] The retrieving of the feature data may comprise retrieving at least one log file of the medical device and / or retrieving the feature data from at least one database comprising feature data read out from a log file and stored in the database. The term “log file” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a computer-generated data file comprising information about and / or of the medical device, in particular the feature data. The log file may be a document. The log file may be an .xml file or may have another format.

[0043] For example, the log file may be retrieved completely. Alternatively, the log file may be preprocessed. The preprocessing may comprise classifying the data into relevant data for root cause analysis and further data. Only information considered as relevant, e.g. such as the frequency or timestamps of certain commands, may be retrieved in step i). The selection which information is to be considered as relevant may be stored in a database, e.g. of the medical device. The preprocessing can allow reducing the amount of data, e.g. by excluding detailed information not relevant for root cause analysis. The excluded data, however, may be retrieved in addition, e.g. in case of further investigation would be required.

[0044] Step ii) comprises applying TF-IDF on the feature data and extracting a feature vector therefrom. The extracting may comprise determining a feature vector. The term “feature vector” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an ordered list of numerical properties of observed features. The feature vector may represent features which can be used for determining the potential root cause in step iii).

[0045] A feature may be a term in the log file. The method may comprise calculating a TF-IDF score for each feature. The extracting of the feature vector may comprise calculating a feature value for every feature and putting the calculated feature value(s) into the feature vector, wherein the feature value is the TF-IDF score. For example, the determination of the feature vector for input into the first trained model may be performed as follows. Firstly, the term frequency tf may be determined and normalized for each term of a predefined set of terms, e.g.

[0046] tf(t, p) = Frequency of term (t) within the prediction document (p)

[0047] normalized tf(t, p) = tf(t, p) / Total number of terms of the prediction document. Subsequently, the document frequency (df) may be determined. E.g. as dfs the dfs determined from training may be used. Subsequently, the inverse document frequencies idfs may be determined. The idfs may be determined by dividing the total number of documents +1 for the prediction document by the number of documents containing the term (df(t)), and then taking the logarithm of that quotient:

[0048] idf(t) = log (Total number of documents + 1 / (1 + df(t))),

[0049] wherein +1 is added to the df(t) to prevent a division by zero. The determination of TF-IDFs may comprise determining TF-IDFs of prediction document determined by multiplying the normalized tf(t, p) with the idf(t) :

[0050] tfidf(t, pd) = normalized tf(t, p) * idf(t).

[0051] The TF-IDFs may be transferred into a feature vector. The TF-IDFs of the prediction document are the input for the first trained model, e.g. the kNN algorithm or Neural Network, that is trained with c-TF-IDFs from a training data set.

[0052] Step iii) comprises applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause. The determining of the potential root cause may comprise predicting a root cause. The applying the first trained model on the feature vector comprises using the feature vector as input for the first trained model.

[0053] The term “processing unit”, also denoted as processor, as generally used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an arbitrary logic circuitry configured for performing basic operations of a computer or system, and / or, generally, to a device which is configured for performing calculations or logic operations. In particular, the processing unit may be configured for processing basic instructions that drive the computer or system. As an example, the processing unit may comprise at least one arithmetic logic unit (ALU), at least one floatingpoint unit (FPU), such as a math co-processor or a numeric coprocessor, a plurality of registers, specifically registers configured for supplying operands to the ALU and storing results of operations, and a memory, such as an LI and L2 cache memory. In particular, the processing unit may be a multi-core processor. Specifically, the processing unit may be or may comprise a central processing unit (CPU). Additionally or alternatively, the processing unit may be or may comprise a microprocessor, thus specifically the processing unit’s elements may be contained in one single integrated circuitry (IC) chip. Additionally or alternatively, the processing unit may be or may comprise one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs) or the like. The processing unit specifically may be configured, such as by software programming, for performing one or more evaluation operations.

[0054] The trained model may be trained by training on at least one training data set. The term “training” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a process of determining parameters of at least one machine learning and / or deep learning model, specifically of the algorithm of the machine learning and / or deep learning model, specifically on at least one training data set or set of training data. The training specifically may comprise at least one optimization or tuning process, wherein a best parameter combination, e.g. according to at least one optimization procedure, is determined.

[0055] The training may comprise providing a trainable model. The term “trainable model”, also denoted as model, as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The trainable model may be or may comprise at least one mathematical model configured for transforming one or more input values into one or more output values by using one or more parameters which may be adjusted in order to enable the model to be trained. The trainable model specifically may be a trainable mathematical model which is trainable on at least one training data set using one or more of machine learning, deep learning, neural networks, or other form of artificial intelligence. The term specifically may refer, without limitation, to the fact that the trainable model can be further trained, optimized or updated based on additional training data. Specifically, the trainable model is trained on a training dataset. The trainable model may be trained by using machine learning. The trainable model may be at least partially data-driven by being trained on data from historical training data. The term “providing a trainable model” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to generating a model and / or retrieving a model from a database. The trainable model may be a model taken from a software library, providing a plurality of trainable models. As an example, the Scikit-learn open-source software library or Keras may be mentioned. Additionally or alternatively, also, as an example, reference may be made to the extreme Gradient Boosting (XGBoost) open-source software library, specifically providing gradient boosting trainable models, being programmed in one or more of the programming languages C++, Java, Python, R, Julia, Perl and Scala. Other libraries, however, are also feasible, as well as customized trainable models not retrieved from software libraries.

[0056] The trainable model of the first trained model may comprise at least one model selected from the group consisting of: a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees.

[0057] The term “trained model” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a trainable model which has gone through at least one training process as defined above, by applying, at least once, a set of training data to the trainable model. Specifically, one or more parameters of the trainable model might have been adapted, on the basis of the training data, in order to transform the trainable model into a trained model. Specifically, the trained model may be a trainable model which was trained on at least one training dataset, also denoted training data. The trained model specifically may be or may comprise a classifier, configured for classifying an object, on the basis of one or more input variables, describing the object. The term “classifying” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a process of categorizing into classes. The term “class” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a unique root cause.

[0058] The computer-implemented method for root cause analysis for investigation of a medical device may be for root cause analysis for investigation of the medical device of a patient who used the device. The training data may be historical log files of medical devices of the same type. E.g., the medical device of the patient who used the device may be an Accu-Chek Solo pump (in short: Accu-Chek Solo) or Accu-Chek Insight pump (in short Accu-Chek Insight). As an example, a data set of about 640 historical log files of Accu-Chek Solo or about 1000 of Accu-Chek Insight devices may be used that had been classified to relate to specific root causes. It is also possible to include historical log files of medical devices which were not classified to relate to a specific root cause, but where the medical device functioned flawless. A training and retraining may be performed each time a new log file is received. The data set may be split into a training data set and a test data set. The splitting may comprise a x%-y% splitting, with x% being in the range 60 to 90, specifically in the range 65 to 75, and more specifically x%= 70, and with y%=100-x. Therein, x% denotes the training data set and y% denotes the test data set. For example, 70 % of the data set may be used as training data set and 30 % for validation. The number of root causes under investigation may be from 1 to 100 root causes, e.g. from 5 to 25 root causes.

[0059] The first trained model was trained supervised and / or unsupervised, in particular by using machine learning. The term “machine learning” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a method of using artificial intelligence (Al) for automatically model building. The term “supervised” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to machine learning based on labeled training data. The labeled training data comprises of a set of training examples labeled with known classes of root causes. For example, the first trained model was trained on annotated historic feature data having known root causes. Supervised training can be advantageous in case there is already the knowledge of the characteristics of the root causes. The term “unsupervised” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to machine learning based on unlabeled data, e.g. the classes of the training data are not known. The target may be to generalize from the training data to unseen instances.

[0060] The training specifically may comprise splitting the training data set into a training data, test data and / or validation data. For example, 70 % of the training data set may be used as training data and 30 % may be used for testing and / or validation. Other splitting ratios, however, are possible.

[0061] The training may further comprise at least one feature extraction procedure. The feature extraction procedure may comprise extracting significant and / or important features from the plurality of features of the feature data. The significance and / or importance of a feature may be quantified, such as by quantifying the influence a variation of the value of the respective feature has on the accuracy of the model. The term “accuracy” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a fraction of correct predictions, e.g. number of correct predictions divided by a total number of predictions. The feature extraction procedure may then, as an example, also comprise considering a feature importance hierarchy. Thus, as an example, the feature extraction procedure may comprise selecting all features having a predetermined feature importance, such as a predetermined feature importance of at least one threshold value. Additionally or alternatively, the features may be ordered according to their feature importance, followed by selecting the most important features such that these selected features, in combination, result in a model accuracy metrics of at least one accuracy threshold value. Additionally or alternatively, the feature extraction procedure may comprise selecting a predetermined number of features, e.g., a predetermined number of features highest in ranking with respect to their feature importance.

[0062] For example, for the trained model the following determinations are made for each document of the training data set. Firstly, the term frequency tf may be determined and normalized for each term of a predefined set of terms, e.g.

[0063] tf(t, d) = Frequency of term (t) within document (d)

[0064] normalized tf(t, d) = tf(t, d) / Total number of terms of document.

[0065] There are several variants of the TF known to the skilled person, e.g. as described in en.wik-ipedia.org / wiki / Tf%E2%80%93idf.

[0066] Subsequently, the document frequency (df) and / or the document frequency of the same class may be determined, e.g. as follows

[0067] df(t) = Number of documents containing term (t) at least one time df(t, c) = Number of documents belonging to same class (c) and containing term at least once. Subsequently, the inverse document frequency idf and / or the class-based inverse document frequency may be determined. The_inverse document frequency may be determined by dividing the total number of documents by the number of documents containing the term (df(t)) and then taking the logarithm of that quotient:

[0068] idf(t) = log (Total number of documents / (1 + df(t)))

[0069] The “+1” may be added to the df(t) to prevent a division by zero. Several variants of the idf may be thinkable such as described in e.g. en.wikipedia.org / wiki / Tf%E2%80%93idf. c-IDFs may be determined by multiplying the idf(t) with the quotient of the number of documents within the current class on which the terms appears (df(t, c)) and the number of documents of the current class (df(c)):

[0070] idf(t, c) = idf(t) * (df(t, c) / df(c))

[0071] This formula is based on “An improvement of TFIDF weighting in text categorization”, Mingy ong Liu, Jiangang Yang (www.semanticscholar.org / paper / An-improvement-of-TFIDF-weighting-in-text-Liu-Yang / 6876855902ca8fc7bfd74fc3eb8ecl39c3bfel45, and www.quora.com / Is-there-something-like-tf-idf-for-classes). For example, subsequently, the c-TF-IDFs may be determined. The c-TF-IDFs may be determined by multiplying the normalized tf(t) with the idf(t, c):

[0072] ctfidf(t, c) = normalized tf(t) * idf(t, c).

[0073] There are several variants of the c-TF-IDF thinkable, e.g. as described in arxiv.org / abs / 2203.05794, maartengr.github.io / BERTopic / api / ctfidf.html and personal. eur.nl / firasincar / papers / WIMS201 l / wims2011.pdf.

[0074] The c-TF-IDFs may be transferred into a feature vector for training the first trained model, e.g. the kNN algorithm or Neural Network.

[0075] For example for the training, the extracting the feature vector may thus comprise using a class-based term frequency-inverse document frequency (c-TF-IDF). The classes are defined with respect to expected root cause. A class may relate to a root cause code. The c-TF-IFD may be designed as described in www.capitalone.com / tech / machine-learning / understand-ing-tf-idf or towardsdatascience.com / creating-a-class-based-tf-idf-with-scikit-learn-caea7bl5b858. For example, the feature vector may be extracted as described in US 10, 740, 374 B2, WO 2022 / 015918 Al. However, other embodiments are possible.

[0076] For example, for investigation importance and / or significance, specifically for supervised training, at least one heat map can be used for visualization of feature values (c-TF-IDFs) versus predefined root causes. For example, an interactive heat map view can be used which shows all determined feature values grouped by root cause code of the training data as a heat map. For example, on the y-axis are the complaints, (e.g. grouped by classified root cause) the x-axis are the features and the z-values are the c-TF-IDFs shown as color and / or gray scale representation. The interactive heat map can provide the possibility that on each entry a tooltip can show the information from the x- and y-axis as well as the z-value. The heat map can visualize the features that are relevant for a specific root cause, e.g. by considering the common feature values of the complaints of a specific root cause, and also shows the similarities of the root causes, e.g. by considering similar feature values among the root causes.

[0077] For example, in case of supervised training, the extraction of features may be performed as follows. Firstly, the model may be trained using supervised machine learning considering history events and history event and event parameter combinations. Then, by using test and / or validation data on the trained model, the accuracy of the prediction is determined, e.g. manually or automatically. Next, the feature extraction may comprise checking if features could be removed or restricted to isolate the significant features to reduce complexity. This could be done manually with the general domain knowledge about the medical device, e.g. by knowing which features would have no impact at all, and / or using the heat map and considering if they have a very low feature value compared to other features. The feature extraction may comprise repeating the supervised training and accuracy determination including loops of events in the feature creation process. For example, loops of specific events may be included which in accordance of general domain knowledge about the medical device were identified as key events for specific root causes. The feature extraction may comprise identifying confusions between root causes by checking the confusion matrix. The term “confusion matrix” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a table layout that allows the visualization of the performance of a machine learning algorithm by showing the confusions between classes. The confusion matrix may show a heat map with the counts of confusions between each pair of root cause codes as z-values that follow a color representation as shown in a diagram legend. The confusion matrix may also show an overlaying table with accuracy, precision and recall for each root cause code for information. The term “recall” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to the true positive rate or sensitivity, e.g. what proportion of actual positives was identified correctly (e.g. determined by True Positives divided by the sum of True Positives and False Negatives). The term “recall” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to the positive prediction value, e.g. what proportion of positive identifications was actually correct (determined by True Positives divided by the sum of True Positives and False Positives). The feature extraction may comprise identifying the complaints and corresponding features that led to a confusion. The feature extraction may comprise repeating the supervised training and accuracy determination including event sequences, e.g. which are tailored for specific root causes that represent the domain knowledge of completed complaints. This can allow reducing confusions between the root causes. For example, the event sequences may be typical event sequences that has been identified during the complaint handling process. The feature extraction may comprise training and working on fine-tuning of tailored event sequence features until the accuracy of the prediction is optimized. The described training can learn the best parameters for its algorithms on its own during the training phase by conducting a wide range of experiments.

[0078] For example, the determining of the potential root cause in step iii) may be performed by using a trained k nearest neighbors (kNN) algorithm. One or more of the following parameters may be determined during the training process of the kNN algorithm: the Manhattan distance, the centroid weight, the Euclidean distance, the inverse weight, a first-neighbordistance, the reverse weight, the uniform weight. The term “Manhattan distance” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a distance between two neighbors measured along axes at right angles. The Manhattan distance may be a parameter used for determining the distance from the root case to be predicted to complaints of the training data. The term “centroid weight” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a weight parameter of the kNN algorithm, wherein each of the k nearest neighbors is assigned their centroid weight for prediction and the centroid weight is determined by the rank order instead of the distances themselves (e.g., for rank 1: (1 + 1 / 2 + ... + 1 / k) / k) and so on until rank k: (1 / k) / k). The term “Euclidean distance” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a straight line distance between two points, wherein this parameter is used to determine the distance from the root cause to be predicted to complaints of the training data. The term “inverse weight” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a weight parameter of the kNN algorithm, wherein each of the k nearest neighbors is assigned their inverse weight for prediction, wherein the inverse weight is determined by computing the inverse of each distance, finding the sum of the inverses and then dividing each inverse by the sum. The term “reverse weight” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a weight parameter of the kNN algorithm, wherein each of the k nearest neighbors is assigned their reverse weight for prediction, wherein the reverse weight is determined by computing the sum of all distances, dividing all distances by the sum, then reversing the order. The term “Uniform Weigh” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to a kNN weight parameter of the kNN algorithm, wherein each of the k nearest neighbors is assigned the same weight for prediction (e.g., majority rule and no weighting). The first-neighbor-distance may be determined during the training process. For example, the first-neighbor-distance may correspond to the 85% decile or a higher decile or the 90% decile or a higher decile of the overall known first-neighbor-distances of correct predictions as determined from the training data set. This can help the user to qualify the accuracy of the current prediction. The value of k of the kNN may be different depending on the training data set, e.g. in case of a limited amount of training data a smaller k may be used. The weighting of neighbors may be such that each neighbor has the same weight or different weight. The Euclidean distance and the Manhattan distance may be determined by experiment. For example, the kNN may be run on the Accu-Chek Solo log files with k=7, Euclidean distance and centroid weights, a ratio of 70% training data and 30 % test data. For example, the kNN may be run on the Accu-Chek Insight log files with k=6, Manhatten distance and centroid weights, a ratio of 70% training data and 30 % test data.

[0079] The kNN algorithm may be designed for performing a multiclass classification based on the extracted feature vector. The method may comprises visualization of information about the determined root cause and / or information about non-classification into potential root causes via at least one userinterface. The visualization may comprise showing machine learning features and the prediction results. This can allow enabling the detection of unknown relationships between anomalies and root causes by a human being for further investigation.

[0080] For example, in case of using a kNN, the potential root cause may be visualized by one or more of: a heat map showing the determined potential root cause, the feature values of the determined potential root case and the feature values of the nearest neighbors (having at least one feature value > 0); a diagram showing weights of root causes of the determined potential root cause and from the set of the nearest neighbors; a diagram showing distances and weights of the nearest neighbors and of the determined potential root cause. In case the distance of the predicted case to the first neighbor is greater than the predefined first-neighbordistance, a warning message may be issued indicating that the prediction might be inaccurate.

[0081] For example, additionally or alternatively, the determining of the potential root cause in step iii) may be performed by using a neural network. The result may be the probabilities for certain root causes. The feature values may be used to train a Neural Network that is used for prediction of the potential root cause. One or more of the following parameters may be determined during the training process of the neural network: weight of a class, number of hidden layers; nodes per layer; drop out. In case of imbalanced classes the ones with lower counts get more weight. For example, the Neural Network may comprise 5 layers, wherein 3 layers may be hidden layers. The first layer may be the input layer and the fifth layer may be the output layer providing the root cause(s). For example, the number of nodes of the input layer may be 141 features, wherein the number of features corresponds to the number of nodes. The number of output root causes may be from 1 to 100, e.g. 20 for Accu-Chek Insight or e.g. 9 for Accu-Chek Solo. For example, the number of nodes for the three hidden layers may be 128, 64, 32 for Accu-Chek Insight. For example, the number of nodes for the three hidden layers may be 256, 64, 64 for Accu-Chek Solo. Randomly neurons (nodes) may be omitted. This drop out, e.g. in the first hidden layer, can allow to get the random principle in. A learning rate may be 0.01. The Keras Tenor Flow library may be used. For example, the NN may be run on the Accu-Chek Insight log files with three hidden layers with 128, 64 and 32 nodes in the three hidden layers respectively, a drop out in the first hidden layer, a learning rate of 0.01 and a ratio of 70% training data and 30 % test data. For example, the NN may be run on the Accu-Chek Solo log files with three hidden layers with 256, 64 and 64 nodes in the three hidden layers respectively, a drop out in the first hidden layer, a learning rate of 0.01 and a ratio of 70% training data and 30 % test data.

[0082] The determining of the at least one root cause may comprise applying a second trained model on the feature data. The method may comprises determining a combined information about the root cause. The second trained model may comprise at least one model selected from the group consisting of a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees. In step iii), the determining of the root cause may be performed with two different algorithms. This can allow to follow the approach of diverse redundancy and to increase the confidence in the prediction.

[0083] In one example of the medical device being an insulin pump, the potential root cause may at least be one of a craze in mainboard; a defective buzzer; a battery issue; a communication issue; an occlusion not handled.

[0084] The method may comprise performing at least one action depending on the determined root cause. The action may comprise one or more of issuing at least one feedback to a technician and / or customer; issuing at least one warning such as to at least one customer being affected and / or a responsible person; issuing at least one early warning to at least one customer and / or a responsible person in case a root cause is determined which will lead to an error or warning in the foreseeable future with a certain probability; issuing at least one warning to at least one customer and / or a responsible person not yet being affected; issuing at least one warning message; issuing a predefined warning code; at least one debugging action; adapting at least one software and / or at least one hardware component of the medical device; adapting of a manual in case of a user error; supporting complaint handling such as batch identification; initiating a recall of the medical device and / or a batch of medical devices. The action may comprise using at least two warning levels depending on severity of the root cause.

[0085] The method may further comprise at least one retraining step. The first model and / or second trained model may be retrained considering the extracted feature data and the determined root cause. The method may comprise at least one feedback loop. The first trained model may be retrained considering feedback from one or more of a customer, a responsible person, a technician, a complaint handling and / or investigation unit received in response to the performed action. The methods as proposed herein provides a large number of advantages over methods of similar kind. Specifically, the above-mentioned technical challenges may be addressed. The method may comprise providing a most likely prediction. This prediction can be confirmed quickly by an investigator before a detailed and time consuming analysis of a complaint is required. It supports the service and investigation process to save time and resources in consideration of an increasing amount of cases during the ramp up of a product. The extracted feature vector from complex event-based histories and determined feature values from the machine learning and the presented prediction results are comprehensible in every sense due to the used algorithms and conceived visualizations. Unknown dependencies between irregularities and root causes may be detected in the visualizations of the solution to enhance the knowledge about symptoms of a root cause to support further investigation. In case of an available machine readable history specification of a medical device, the features for the machine learning process can be extracted automatically to speed up the feature extraction process during the training phase, e.g. the features do not need to be described manually. Root cause tailored features, e.g. event sequences, can be configured to include the domain knowledge about root causes that has been acquired during the service and investigation process to complement the automatically extracted features. These features reduce the likelihood of confusion between the predicted root causes and optimize the overall performance. The automatic extraction of features and the possibility to extend these by configured root cause tailored features enables a quick extension to support additional devices. The conceived first-neighbor-distance warning threshold can support the user to qualify the accuracy of the prediction. The confidence in the prediction may be increased by using two different algorithms following the approach of diverse redundancy, e.g. by using kNN and a neural network.

[0086] In a further aspect of the present invention, an investigation system configured for root cause analysis for investigation of a medical device is proposed. The investigation system comprises:

[0087] at least one communication interface configured for retrieving feature data of the medical device;

[0088] at least one processing unit configured for applying term frequency-inverse document frequency (TF-IDF) on the feature data, extracting a feature vector therefrom and for applying at least one first trained model on the feature vector thereby determining at least one potential root cause. The term "system" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically may refer, without limitation, to an arbitrary set of interacting or interdependent components parts forming a whole. Specifically, the components may interact with each other in order to fulfill at least one common function. The at least two components may be handled independently or may be coupled or connectable.

[0089] The investigation system is configured for performing a method for root cause analysis according to the present invention, such as according to any one of the embodiments of the method described above and / or according to any one of the embodiments of the method described in further detail below. With respect to definitions and embodiments reference is made to the description of the method described above and / or as described in more detail below.

[0090] The investigation system may further comprise at least one user interface configured for visualizing of at least one information about the determined root cause. The term "user interface" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term may refer, without limitation, to a feature of the investigation system which is configured for interacting with its environment, such as for the purpose of unidirectionally or bidirectionally exchanging information, such as for exchange of one or more of data or commands. For example, the user interface may be configured to share information with a user and to receive information by the user. The user interface may be a feature to interact visually with a user, such as a display, or a feature to interact acoustically with the user. The user interface, as an example, may comprise one or more of: a graphical user interface; a data interface, such as a wireless and / or a wire-bound data interface.

[0091] The investigation system may be at least partially an element of one or more of the medical device, a diabetes management system, a smartphone, a cloud.

[0092] Further disclosed and proposed herein is a computer program including computer-executable instructions for performing the method according to the present invention in one or more of the embodiments enclosed herein when the instructions are executed on a computer or computer network. Specifically, the computer program may be stored on a computer-readable data carrier and / or on a computer-readable storage medium. As used herein, the terms “computer-readable data carrier” and “computer-readable storage medium” specifically may refer to non-transitory data storage means, such as a hardware storage medium having stored thereon computer-executable instructions. The computer-readable data carrier or storage medium specifically may be or may comprise a storage medium such as a random-access memory (RAM) and / or a read-only memory (ROM).

[0093] Thus, specifically, one, more than one or even all of method steps i) to iii) as indicated above may be performed by using a computer or a computer network, preferably by using a computer program.

[0094] Further disclosed and proposed herein is a computer program product having program code means, in order to perform the method according to the present invention in one or more of the embodiments enclosed herein when the program is executed on a computer or computer network. Specifically, the program code means may be stored on a computer-readable data carrier and / or on a computer-readable storage medium.

[0095] Further disclosed and proposed herein is a data carrier having a data structure stored thereon, which, after loading into a computer or computer network, such as into a working memory or main memory of the computer or computer network, may execute the method according to one or more of the embodiments disclosed herein.

[0096] Further disclosed and proposed herein is a non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to one or more of the embodiments disclosed herein.

[0097] Further disclosed and proposed herein is a computer program product with program code means stored on a machine-readable carrier, in order to perform the method according to one or more of the embodiments disclosed herein, when the program is executed on a computer or computer network. As used herein, a computer program product refers to the program as a tradable product. The product may generally exist in an arbitrary format, such as in a paper format, or on a computer-readable data carrier and / or on a computer-readable storage medium. Specifically, the computer program product may be distributed over a data network. Finally , disclosed and proposed herein is a modulated data signal which contains instructions readable by a computer system or computer network, for performing the method according to one or more of the embodiments disclosed herein.

[0098] Referring to the computer-implemented aspects of the invention, one or more of the method steps or even all of the method steps of the method according to one or more of the embodiments disclosed herein may be performed by using a computer or computer network. Thus, generally, any of the method steps including provision and / or manipulation of data may be performed by using a computer or computer network. Generally, these method steps may include any of the method steps, typically except for method steps requiring manual work, such as providing the samples and / or certain aspects of performing the actual measurements.

[0099] Specifically, further disclosed herein are:

[0100] a computer or computer network comprising at least one processor, wherein the processor is adapted to perform the method according to one of the embodiments described in this description,

[0101] a computer loadable data structure that is adapted to perform the method according to one of the embodiments described in this description while the data structure is being executed on a computer,

[0102] a computer program, wherein the computer program is adapted to perform the method according to one of the embodiments described in this description while the program is being executed on a computer,

[0103] a computer program comprising program means for performing the method according to one of the embodiments described in this description while the computer program is being executed on a computer or on a computer network,

[0104] a computer program comprising program means according to the preceding embodiment, wherein the program means are stored on a storage medium readable to a computer,

[0105] a storage medium, wherein a data structure is stored on the storage medium and wherein the data structure is adapted to perform the method according to one of the embodiments described in this description after having been loaded into a main and / or working storage of a computer or of a computer network, and

[0106] a computer program product having program code means, wherein the program code means can be stored or are stored on a storage medium, for performing the method according to one of the embodiments described in this description, if the program code means are executed on a computer or on a computer network. In a further aspect of the present invention, a use of an investigation system according to the present invention, such as according to any one of the embodiments of the investigation system described above and / or according to any one of the embodiments of the investigation system described in further detail below, in an application selected from the group consisting of: a medical application, in particular for an investigation of an insulin pump, or other medical devices.

[0107] Summarizing and without excluding further possible embodiments, the following embodiments may be envisaged:

[0108] Embodiment 1. A computer-implemented method for root cause analysis for investigation of a medical device, wherein the method comprises the following steps i) retrieving feature data of the medical device by using at least one communication interface;

[0109] ii) applying term frequency-inverse document frequency (TF-IDF) on the feature data and extracting a feature vector therefrom,

[0110] iii) applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause.

[0111] Embodiment 2. The method according to the preceding embodiment, wherein the feature data comprises information about one or more of: at least one event; at least one event sequence with temporal relationship; at least one loop of events and / or at least one event parameter.

[0112] Embodiment 3. The method according to any one of the preceding embodiments, wherein the feature data comprise data selected from the group consisting of: at least one error message, at least one error warning, therapy data, measurement data, further sensor data of at least one further sensor of the medical device such as temperature sensor, information about user interaction, information about a device status such as a power status and / or battery voltage, status of one or more physical sensors, or accessibility to one or more elements of the device, device functional data such as currents, system status of the medical device or of at least one element of the medical device.

[0113] Embodiment 4. The method according to any one of the preceding embodiments, wherein the feature data further comprises fitness and / or health data of a further device. Embodiment 5. The method according to any one of the preceding embodiments, wherein the medical device is a device selected from the group consisting of: at least one insulin pump, a glucose meter, a home use device for analyte measurement.

[0114] Embodiment 6. The method according to any one of the preceding embodiments, wherein the potential root causes are one or more of: a malfunction of at least one element of the medical device such as of at least one sensor of the medical device; a maloperation; at least one hardware problem or failure; at least one software problem or failure; a precursor appearing before the malfunction and / or failure.

[0115] Embodiment 7. The method according to any one of the preceding embodiments, wherein the feature data comprise one or more features and wherein extracting the feature vector comprises calculating a feature value for the one or more features, respectively and putting the calculated feature values into the feature vector, wherein the feature values are a TF-IDF scores.

[0116] Embodiment 8. The method according to any one of the preceding embodiments, wherein the retrieving of the feature data comprises retrieving at least one log file of the medical device and / or retrieving the feature data from at least one database comprising feature data read out from a log file and stored in the database.

[0117] Embodiment 9. The method according to the preceding embodiment, wherein the log file is retrieved completely, and / or wherein the log file is preprocessed.

[0118] Embodiment 10. The method according to any one of the two preceding embodiments, wherein the database is at least partially cloud based and / or a local database of the medical device and / or an external the database, for example of a least one portable communication device such as a smartphone.

[0119] Embodiment 11. The method according to any one of the preceding embodiments, wherein the method comprises uploading the feature data to a database, wherein the uploading is performed when the medical device is under investigation at a complaint handling and / or investigation unit, and / or at the end of the life time of the medical device and / or in regular intervals or continuously. Embodiment 12. The method according to any one of the preceding embodiments, wherein the feature data comprises machine-readable data, wherein the machine-readable data is textual data or binary data.

[0120] Embodiment 13. The method according to any one of the preceding embodiments, wherein the first trained model was trained supervised and / or unsupervised.

[0121] Embodiment 14. The method according to the preceding embodiment, wherein the method comprises at least one retraining step, wherein the first trained model is retrained considering the extracted feature data and the determined root cause.

[0122] Embodiment 15. The method according to any one of the two preceding embodiments, wherein the training comprises a feature extraction procedure, wherein at least one heat map is used for visualization of feature values versus predefined root causes.

[0123] Embodiment 16. The method according to any one of the preceding embodiments, wherein the first trained model comprises at least one model selected from the group consisting of: a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees.

[0124] Embodiment 17. The method according to any one of the preceding embodiments, wherein the first trained model was trained on annotated historic feature data having known root causes and wherein the training comprised applying class-based term frequency inverse document frequency (c-TF-IDF) to the historic feature data and extracting historic feature vectors therefrom and training a k nearest neighbors algorithm or a Neural Network therewith.

[0125] Embodiment 18. The method according to any one of the preceding embodiments, wherein the determining of the at least one root cause comprises applying a second trained model on the feature data, wherein the method comprises determining a combined information about the root cause, wherein the second trained model comprises at least one model selected from the group consisting of: a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees.

[0126] Embodiment 19. The method according to any one of the preceding embodiments, wherein the method comprises visualization of information about the determined root cause and / or information about non-classification into potential root causes via at least one user-interface.

[0127] Embodiment 20. The method according to any one of the preceding embodiments, wherein the method comprises performing at least one action depending on the determined root cause, wherein the action comprises one or more of: issuing at least one feedback to a technician and / or customer; issuing at least one warning such as to at least one customer being affected and / or a responsible person; issuing at least one early warning to at least one customer and / or a responsible person in case a root cause is determined which will lead to an error or warning in the foreseeable future with a certain probability; issuing at least one warning to at least one customer and / or a responsible person not yet being affected; issuing at least one warning message; issuing a predefined warning code; at least one debugging action; adapting at least one software and / or at least one hardware component of the medical device; adapting of a manual in case of a user error; supporting complaint handling such as batch identification; initiating a recall of the medical device and / or a batch of medical devices.

[0128] Embodiment 21. The method according to the preceding embodiment, wherein the action comprises using at least two warning levels depending on severity of the root cause.

[0129] Embodiment 22. The method according to any one of the two preceding embodiments, wherein the method comprises at least one feedback loop, wherein the first trained model is retrained considering feedback from one or more of a customer, a responsible person, a technician, a complaint handling and / or investigation unit received in response to the performed action.

[0130] Embodiment 23. An investigation system configured for root cause analysis for investigation of a medical device, wherein the investigation system comprises:

[0131] at least one communication interface configured for retrieving feature data of the medical device;

[0132] at least one processing unit configured for applying term frequency-inverse document frequency (TF-IDF) on the feature data, extracting a feature vector therefrom and for applying at least one first trained model on the feature vector thereby determining at least one potential root cause, and

[0133] wherein the investigation system is configured for performing the method according to any one of the preceding embodiments . Embodiment 24. The investigation system according to the preceding embodiment, further comprising at least one user interface configured for visualizing of at least one information about the determined root cause.

[0134] Embodiment 25. The investigation system according to any one of the preceding embodiments referring to an investigation system, wherein the investigation system is at least partially an element of one or more of the medical device, a diabetes management system, a smartphone, a cloud.

[0135] Embodiment 26. A computer program comprising instructions which, when the program is executed by the investigation system according to any one of the preceding embodiments referring to an investigation system, cause the investigation system to perform the method according to any one of the preceding embodiments referring to a method.

[0136] Embodiment 27. A computer-readable storage medium comprising instructions which, when the instructions are executed by the investigation system according to any one of the preceding embodiments referring to an investigation system, cause the investigation system to perform the method according to any one of the preceding embodiments referring to a method.

[0137] Embodiment 28. A non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of the preceding embodiments referring to a method.

[0138] Embodiment 29. A use of an investigation system according to any one of the preceding embodiments referring to an investigation system, in an application selected from the group consisting of: a medical application, in particular for an investigation of an insulin pump, or other medical devices.

[0139] Short description of the Figures

[0140] Further optional features and embodiments will be disclosed in more detail in the subsequent description of embodiments, preferably in conjunction with the dependent claims. Therein, the respective optional features may be realized in an isolated fashion as well as in any arbitrary feasible combination, as the skilled person will realize. The scope of the invention is not restricted by the preferred embodiments. The embodiments are schematically depicted in the Figures. Therein, identical reference numbers in these Figures refer to identical or functionally comparable elements.

[0141] In the Figures:

[0142] Figure 1 shows a flow chart of an embodiment of a computer-implemented method for root cause analysis for investigation of a medical device;

[0143] Figure 2 shows an exemplary heat map for visualization of feature values versus predefined root causes;

[0144] Figures 3 and 4 show results of an exemplary prediction of an unknown complaint using a kNN algorithm;

[0145] Figures 5A and 5B show an embodiment of an investigation system configured for root cause analysis for investigation of a medical device (Figure 5A) and a flow chart of a further embodiment of a computer-implemented method for root cause analysis for investigation of a medical device (Figure 5B); and

[0146] Figures 6A and 6B show an embodiment of an investigation system configured for root cause analysis for investigation of a medical device (Figure 6A) and a flow chart of a further embodiment of a computer-implemented method for root cause analysis for investigation of a medical device (Figure 6B).

[0147] Detailed description of the embodiments

[0148] Figure 1 shows a flow chart of an embodiment of a computer-implemented method for root cause analysis 110 for investigation of a medical device (not shown in Figure 1). The method comprises the following method steps, which, as an example, may be performed in the given order. However, a different order is also feasible. Further, it is possible to perform two or more of the method steps simultaneously or in a fashion overlapping in time. Further, it is also possible to perform one, more than one or even all of the method steps repeatedly.

[0149] The method comprises the following steps

[0150] i) (denoted by reference number 112) retrieving feature data of the medical device by using at least one communication interface;

[0151] ii) (denoted by reference number 114) applying term frequency-inverse document frequency (TF-IDF) on the feature data and extracting a feature vector therefrom, iii) (denoted by reference number 116) applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause.

[0152] The method may comprise determining at least one root cause of an unknown complaint of a medical device, e.g. by means of machine learning based on completed complaints in a service investigation process. The present invention can allow for reliable root cause analysis for medical devices using machine learning features. The proposed extraction of a feature vector can allow considering a whole device history, e.g. as present in log files relating to the medical device. The method can cover root causes that are related to significant irregularities in this data. The machine learning can comprise considering domain knowledge e.g. for feature extraction, allowing for reduced effort but high accuracy of prediction results. Specifically, the determined root cause is independent from the patient’s allegation.

[0153] Examples and possible embodiments of features and aspects of the method will be discussed in the following with respect to Figures 2 to 6B. Thus, for a detailed description of examples and possible embodiments related to the computer-implemented method for root cause analysis 110 for investigation of a medical device, reference is made to the description of Figures 2 to 6B. In the examples of Figures 2 to 6B, the medical device comprises at least one medication device such as a medication pump, e.g. an insulin pump. However, other examples of medical devices are also feasible.

[0154] Figure 2 shows an exemplary heat map for visualization of feature values versus predefined root causes. Specifically, the exemplary heat map of Figure 2 shows different predefined root causes 118 on the y-axis, optionally numbered as complaints 120 and features on the x-axis. In this example, the root causes 118 may comprise: craze in mainboard 122; defective buzzer 124; battery issue 126; communication issue 128; occlusion not handled 130. The features may comprise: electronic error (buzzer interruption) 132; electronic error (response corrupted) 134; communication lost 136; maintenance (occlusion detected) 138; change battery 140; loop of communication lost 142; loop of electronic error (response corrupted) 144; typical battery issue 146, e.g. sequence of change battery, battery voltage (value < IV and delta < 30 min); typical occlusion not handled 148, e.g. loop of maintenance (occlusion detected), missing remove reservoir, missing insert reservoir, maintenance (occlusion detected).

[0155] Further, in the heat map of Figure 2, values of a class-based term frequency -inverse document frequency are shown for each pair of root cause 118 and feature. Step ii) may comprise extracting a feature vector from the feature data by using term frequency -inverse document frequency (TF-IDF). The extracting may comprise determining a feature vector. For example, the extracting the feature vector may comprise using a class-based term frequency-inverse document frequency (c-TF-IDF). The classes are defined with respect to expected root cause 118. The c-TF-IFD may be designed as described in www.capitalone.com / tech / ma-chine-leaming / understanding-tf-idf or towardsdatascience.com / creating-a-class-based-tf-idf-with-scikit-leam-caea7b 15b 858.

[0156] The trained model, e.g. the first trained model applied in step iii) of the method shown in Figure 1, may be trained by training on at least one training data set. The training may comprise providing a trainable model. The trainable model of the first trained model may comprise at least one of a k nearest neighbors algorithm and a Neural Network. The training data may be historical log files of Accu-Chek Solo and / or Accu-Chek Insight.

[0157] The heat map as shown in Figure 2 may be used for feature extraction. The first trained model may be trained supervised, in particular by using machine learning. The training may further comprise at least one feature extraction procedure. For example, for investigation importance and / or significance, specifically for supervised training, at least one heat map can be used for visualization of feature values versus predefined root causes 118. For example, an interactive heat map view can be used which shows all determined feature values grouped by root cause code of the training data as a heat map. As shown in the example of Figure 2, on the y-axis are the complaints 120, the x-axis are the features and the z-values are the c-TF-IDFs shown as color and / or gray scale representation. The interactive heat map can provide the possibility that on each entry a tooltip can show the information from the x-and y-axis as well as the z-value. The heat map can visualize the features that are relevant for a specific root cause 118, e.g. by considering the common feature values of the complaints of a specific root cause 118, and also shows the similarities of the root causes 118, e.g. by considering similar feature values among the root causes 118.

[0158] Further, as an example, the extraction of features may be performed as follows. Firstly, the model may be trained using supervised machine learning considering history events and history event and event parameter combinations. Then, by using test and / or validation data on the trained model, the accuracy of the prediction is determined, e.g. manually or automatically. Next, the feature extraction may comprise checking if features could be removed or restricted to isolate the significant features to reduce complexity. This could be done manually with the general domain knowledge about the medical device, e.g. by knowing which features would have no impact at all, and / or using the heat map and considering if they have a very low feature value compared to other features. The feature extraction may comprise repeating the supervised training and accuracy determination including loops of events in the feature creation process. For example, loops of specific events may be included which in accordance of general domain knowledge about the medical device were identified as key events for specific root causes 118. The feature extraction may comprise identifying confusions between root causes 118 by checking the confusion matrix. The feature extraction may comprise identifying the complaints and corresponding features that led to a confusion. The feature extraction may comprise repeating the supervised training and accuracy determination including event sequences, e.g. which are tailored for specific root causes 118 that represent the domain knowledge of completed complaints. This can allow reducing confusions between the root causes 118. For example, the event sequences may be typical event sequences that has been identified during the complaint handling process. The feature extraction may comprise training and working on fine-tuning of tailored event sequence features until the accuracy of the prediction is optimized. The described training can learn the best parameters for its algorithms on its own during the training phase by conducting a wide range of experiments.

[0159] Further, as an example, the determining of the potential root cause in step iii) may be performed by using a trained k nearest neighbors (kNN) algorithm. The result of the determining of the potential root cause using a trained kNN algorithm is shown in Figure 3. Figure 3 shows the result of an exemplary prediction of an unknown complaint 120 (labeled by number #99 in Figure 3) with k = 6. Specifically, Figure 3 shows a filtered view of the heat map shown in Figure 2 of common c-TF-IDFs and the TF-IDFs of the predicted case and of nearest neighbors whereby any z-value is >0. The features 136, 140 and 146 where the TF- IDF of the predicted case and any of the cTF-IDF of the nearest neighbors is >0 are highlighted (framed) in Figure 3. In Figure 3, the nearest neighbor number 150 is ordered by increasing distance from bottom-up.

[0160] The kNN algorithm may be designed for performing a multiclass classification based on the extracted feature vector. The potential root cause may be visualized by one or more of: a heat map showing the determined potential root cause, the feature values of the determined potential root case and the feature values of the nearest neighbors (having at least one feature value > 0); a diagram showing weights of root causes of the determined potential root cause and from the set of the nearest neighbors; a diagram showing distances and weights of the nearest neighbors and of the determined potential root cause.

[0161] The result of the prediction using the kNN algorithm can be seen in Table 1 :

[0162]

[0163] Table 1: Exemplary prediction of unknown complaint #99 with kNN (k = 6)

[0164] Further, Figure 4 shows the result of an exemplary prediction of an unknown complaint 120 (labeled by number #99 in Figure 3) with k = 6. Specifically, as outlined above, Figure 4 shows a diagram showing distances (denoted by reference number 152) and weights (denoted by reference number 154) of the nearest neighbors and of the determined potential root cause. In case the distance of the predicted case to the first neighbor is greater than a predefined first-neighbor-distance, a warning message may be issued indicating that the prediction might be inaccurate. In the diagram of the distances, the predefined first-neighbor-distance 155 is exemplarily shown.

[0165] Alternatively or additionally, the determining of the at least one root cause may comprise applying a second trained model on the feature data. The method may comprises determining a combined information about the root cause. The second trained model may comprise at least one model selected from the group consisting of: a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees. In step iii), the determining of the root cause may be performed with two different algorithms. This can allow to follow the approach of diverse redundancy and to increase the confidence in the prediction.

[0166] For example, a neural network may be used as the second trained model. The training of the second trainable model may be performed similarly to the training of the first trainable model, as outlined above. Thus, in this example, the determining of the potential root cause in step iii) may be performed by using the neural network, specifically in addition to the kNN algorithm. The result may be the probabilities for certain root causes. The feature values may be used to train a neural network that is used for prediction of the potential root cause. One or more of the following parameters may be determined during the training process of the neural network: weight of a class, number of hidden layers; nodes per layer; drop out. In case of imbalanced classes, the ones with lower counts get more weight.

[0167] The result of the prediction of the unknown complaint #99 is shown in Table 2:

[0168]

[0169] Table 2: Exemplary prediction of unknown complaint #99 with neural network

[0170] Figures 5 A and 5B show an embodiment of an investigation system 156 configured for root cause analysis for investigation of a medical device 158 (Figure 5 A) and a flow chart of a further embodiment of a computer-implemented method for root cause analysis for investigation of a medical device 158 (Figure 5B).

[0171] As shown in Figure 5 A, the investigation system 156 comprises at least one communication interface 160 configured for retrieving feature data of the medical device 158. For example, the medical device 158 may be configured for transferring the feature data via the communication interface 160 to a remote controller 162, e.g. e remote controller of a portable communication device.

[0172] The investigation system 156 further comprises at least one processing unit 164 configured for extracting the feature vector from the feature data by using term frequency-inverse document frequency (TF-IDF) and for applying at least one first trained model on the feature vector thereby determining at least one potential root cause. As shown in Figure 5A, the remote controller 162 and the processing unit 164 may be connected via a further communication interface 166, wherein the remote controller 162 may be configured for forwarding the feature data to the processing unit 164 via the further communication interface 166. The processing unit 164 may specifically be a cloud-based processing unit. Thus, in this example, the cloud-based processing unit may be configured for performing steps i) to iii) of the method for root cause analysis 110. The processing unit 164 may be configured, depending on the determined root cause, for issuing at least one early warning to at least one customer, e.g. via the further communication interface 166 to the remote controller 162, and / or to a responsible person (denoted by reference number 168) in case a root cause is determined which will lead to an error or warning in the foreseeable future with a certain probability. The responsible person 168 may be one or more of a parent, a diabetologist or the like.

[0173] Further, as shown in Figure 5 A, the investigation system 156 may further comprise at least one user interface 170 configured for visualizing of at least one information about the determined root cause. For example, the remote controller 162 may be configured for forwarding the early warning to at least one customer (denoted by reference number 172), e.g. at least one patient, via the user interface 170. Further, the customer 172 may be able to provide feedback to the remote controller 162 via the user interface 170, e.g. if the early warning was helpful or the like. The feedback may be used for retraining the first and / or second trained model.

[0174] The investigation system 156 may comprise a plurality of medical devices 158, remote controllers 162, customers 172 and / or responsible persons 168. Thus, as indicated by reference number 174, the processing unit 164 may be configured for communicating with a plurality of units 174 comprising the medical device 158, the remote controller 162, the customer 172 and / or the responsible person 168. The investigation system 156 may be configured for performing a method for root cause analysis 110 according to the present invention, such as according to the exemplary embodiment shown in Figure 5B and / or according to any other embodiment disclosed herein. In the exemplary embodiment of Figure 5B, the medical device 158 may transfer the feature data via the communication interface 160 to the remote controller 162 (denoted by reference number 176). The remote controller 162 may forward the feature data to the processing unit 164 via the further communication interface 166 (denoted by reference number 178). As outlined above, in the exemplary embodiment of Figure 5 A, the processing unit 164 may be a cloud-based processing unit. Thus, in this exemplar, steps i) to iii) may be performed on the cloud-based processing unit.

[0175] The method may proceed at decision node 180. The method may comprise performing at least one action depending on the determined root cause. For example, in case no upcoming issue may be detected, the method may proceed with updating the database and / or at least one of the first and second trained model (denoted by reference number 182). Alternatively, in case a root cause is determined which will lead to an error or warning in the foreseeable future with a certain probability, the method may proceed with issuing at least one early warning to the customer 172, e.g. via the further communication interface 166 to the remote controller 162, (denoted by reference number 184) and / or to the responsible person 168 (denoted by reference number 186). Further, at least one of the customer 172 and the responsible person 168 may be prompted to provide feedback to the remote controller 162, e.g. via the user interface 170, e.g. if the early warning was helpful or the like (denoted by reference number 188). The feedback may be send back to update the database and / or at least one of the first and second trained model (denoted by reference number 190).

[0176] Figures 6A and 6B show another embodiment of an investigation system 156 configured for root cause analysis for investigation of a medical device 158 (Figure 6A) and a flow chart of a further embodiment of a computer-implemented method for root cause analysis 110 for investigation of a medical device 158 (Figure 6B). The embodiments of Figures 6A and 6B widely correspond to the embodiments shown in Figures 5 A and 5B. Thus, for a detailed description of the embodiments of Figures 6A and 6B, reference is made to the description of Figures 5 A and 5B.

[0177] As an alternative to the embodiments of Figures 6 A and 6B, the remote controller 162 may comprise the processing unit 164. Thus, steps i) to iii) may be performed on the remote controller 162. Further, in this example, the investigation system 156 may comprise a database 192. The database 192 may specifically be a cloud-based database. The remote controller 162 may be connected with the database 192 via the further communication interface 166. For example, the remote controller 162 may be configured for retrieving at least one of the first and second trainable model from the database, e.g. via the further communication interface 166. Thus, as shown in Figure 6B, the method may additionally comprise retrieving the trainable model from the database 192 (denoted by reference number 194). List of reference numbers

[0178] method for root cause analysis

[0179] retrieving feature data

[0180] extracting a feature vector

[0181] applying a first trained model on the feature vector

[0182] root cause

[0183] complaint

[0184] craze in mainboard

[0185] defective buzzer

[0186] battery issue

[0187] communication issue

[0188] occlusion not handled

[0189] electronic error (buzzer interruption)

[0190] electronic error (response corrupted)

[0191] communication lost

[0192] maintenance (occlusion detected)

[0193] change battery

[0194] loop of communication lost

[0195] loop of electronic error (response corrupted)

[0196] typical battery issue

[0197] typical occlusion not handled

[0198] nearest neighbor number

[0199] distance of nearest neighbors and determined potential root cause weight of nearest neighbors and determined potential root cause first neighbor distance

[0200] investigation system

[0201] medical device

[0202] communication interface

[0203] remote controller

[0204] processing unit

[0205] further communication interface

[0206] a responsible person

[0207] user interface

[0208] customer

[0209] unit transfer feature data to remote controller

[0210] forward feature data to the processing unit decision node

[0211] updating database, first and / or second trained model issuing an early warning to customer

[0212] issuing an early warning to responsible person prompting to provide feedback

[0213] send back feedback

[0214] database

[0215] retrieving trainable model from database

Claims

1. - 43 -2.Roche Diabetes Care GmbH3.Claims1. A computer-implemented method for root cause analysis for investigation of a medical device, wherein the method comprises the following steps5.i) retrieving feature data of the medical device by using at least one communication interface;6.ii) applying term frequency-inverse document frequency (TF-IDF) on the feature data and extracting a feature vector therefrom,7.iii) applying, by using at least one processing unit, at least one first trained model on the feature vector thereby determining at least one potential root cause.

2. The method according to the preceding claim, wherein the feature data comprises information about one or more of: at least one event; at least one event sequence with temporal relationship; at least one loop of events and / or at least one event parameter.

3. The method according to any one of the preceding claims, wherein the feature data comprise data selected from the group consisting of: at least one error message, at least one error warning, therapy data, measurement data, further sensor data of at least one further sensor of the medical device such as temperature sensor, information about user interaction, information about device status, device functional data such as currents, system status of the medical device or of at least one element of the medical device.

4. The method according to any one of the preceding claims, wherein the medical device is a device selected from the group consisting of: at least one insulin pump, a glucose meter, a home use device for analyte measurement.

5. The method according to any one of the preceding claims, wherein the potential root causes are one or more of: a malfunction of at least one element of the medical device such as of at least one sensor of the medical device; a maloperation; at least one hardware problem or failure; at least one software problem or failure; a precursor appearing before the malfunction and / or failure.- 44 -6. The method according to any one of the preceding claims, wherein the feature data comprise one or more features and wherein extracting the feature vector comprises calculating a feature value for the one or more features, respectively and putting the calculated feature values into the feature vector, wherein the feature values are TF- IDF scores.

7. The method according to any one of the preceding claims, wherein the retrieving of the feature data comprises retrieving at least one log file of the medical device and / or retrieving the feature data from at least one database comprising feature data read out from a log file and stored in the database.

8. The method according to any one of the preceding claims, wherein the method comprises uploading the feature data to a database, wherein the uploading is performed when the medical device is under investigation at a complaint handling and / or investigation unit, and / or at the end of the life time of the medical device and / or in regular intervals or continuously.

9. The method according to any of the preceding claims, wherein the method comprises at least one retraining step, wherein the first trained model is retrained considering the extracted feature data and the determined root cause.

10. The method according to any one of the preceding claims, wherein the first trained model comprises at least one model selected from the group consisting of: a k nearest neighbors algorithm, a Neural Network, a linear regression classifier, xgboost, random forest, decision trees.

11. The method according to any one of the preceding claims, wherein the first trained model was trained on annotated historic feature data having known root causes and wherein the training comprised applying class-based term frequency inverse document frequency (c-TF-IDF) to the historic feature data and extracting historic feature vectors therefrom and training a k nearest neighbors algorithm or a Neural Network therewith.

12. The method according to any one of the preceding claims, wherein the method comprises performing at least one action depending on the determined root cause, wherein the action comprises one or more of: issuing at least one feedback to a technician and / or customer; issuing at least one warning such as to at least one customer being affected and / or a responsible person; issuing at least one early warning to at least one- 45 -19.customer and / or a responsible person in case a root cause is determined which will lead to an error or warning in the foreseeable future with a certain probability; issuing at least one warning to at least one customer and / or a responsible person not yet being affected; issuing at least one warning message; issuing a predefined warning code; at least one debugging action; adapting at least one software and / or at least one hardware component of the medical device; adapting of a manual in case of a user error; supporting complaint handling such as batch identification; initiating a recall of the medical device and / or a batch of medical devices.

13. The method according to the preceding claim, wherein the action comprises using at least two warning levels depending on severity of the root cause.

14. An investigation system configured for root cause analysis for investigation of a medical device, wherein the investigation system comprises:22.at least one communication interface configured for retrieving feature data of the medical device;23.at least one processing unit configured for applying term frequency-inverse document frequency (TF-IDF) on the feature data, extracting a feature vector therefrom and for applying at least one first trained model on the feature vector thereby determining at least one potential root cause, and24.wherein the investigation system is configured for performing the method according to any one the preceding claims.

15. A computer program comprising instructions which, when the program is executed by an investigation system comprising a processing unit and at least one communication interface configured for retrieving feature data of a medical device cause the investigation system to perform the method according to any one of the preceding claims referring to a method.