Cross-site re-identification of patients while preserving privacy
Patent Information
- Application Number
- DE102025106920
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2026-08-27
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a system for providing training data for training a machine learning model and a method for using the system. To train machine learning algorithms for analyzing medical images, it is common practice to divide the available training data into several cohorts. For example, the available training data can be divided into at least training data, validation data, and test data. It is also common practice to remove all private health information (PHI) using an anonymization tool to protect data privacy.If the same patient visits the same medical facility multiple times, or if the same patient visits multiple medical facilities whose medical data are aggregated in a central location to train the machine learning algorithm, the medical data of the same patient, who may exhibit the same symptoms, can be assigned to different database cohorts—such as training data, validation data, and test data—which weakens the accuracy and / or efficiency of the machine learning algorithm's training and / or validation. Regardless of the grammatical usage, the term includes individuals of male, female, or other gender identities. According to data available to the inventor, this so-called patient overlap can be estimated at approximately 3%.The higher the density of medical facilities in a given area, such as a metropolitan region, the greater the resulting patient overlap. Therefore, there remains a need to create a system for identifying patient overlap while considering data privacy aspects in order to improve training data. One object of the present invention is to improve the training data for training the machine learning algorithm. Accordingly, a system for improving the training data for training a machine learning model is created. The system comprises: a medical imaging unit for capturing a medical image of a current patient; an overlap detection unit for identifying a patient overlap between the captured medical image and a previous dataset stored in a shared database and correlated with a previous patient; and a patient allocation unit which, in the event of identifying a patient overlap, assigns the captured medical image to one of several database cohorts for training the machine learning model, according to the previous dataset. The machine learning model is trained, and / or will be trained, to analyze medical images. For example, the machine learning model can be trained, and / or will be trained, to identify anatomical structures or anomalies in medical images. Training the machine learning model can be done through supervised and / or unsupervised training. This document can be applied to any machine learning model used to analyze medical images. The medical imaging unit can be, for example, a computed tomography (CT) scanner, a magnetic resonance imaging (MRI) scanner, an X-ray machine, or any other device capable of acquiring medical images. In this context, a medical image can be understood as a representation of a patient's anatomical and / or medical structures. For example, the medical image could be cross-sections of a patient's body acquired by a CT scanner. The shared database may reside in the current medical facility, in another medical facility, and / or on a cloud server. A record can be understood as a collection of data correlated with a medical examination, specifically a single examination and / or multiple examinations over time on a particular individual patient. The record may include one or more medical images, findings correlated with the medical image(s), and means of identification to correlate the data in the record with the specific patient. The shared database may contain a single previous record or a set of previous records.In the latter case, the overlap detection unit can be designed to identify a patient overlap between the captured medical image and the set of previous records. However, it is preferred that the overlap detection unit be designed to identify the patient overlap of the captured medical image with each previous record in the set of previous records individually. If no patient overlap is identified for a particular record in the set of previous records, another previous record in the set of previous records can be used to identify patient overlaps. In the following, findings can be understood as medical and / or anatomical anomalies of a patient and / or diagnoses that correlate with the medical and / or anatomical anomalies. Landmarks can be understood as distinctive locations within a medical image in relation to the patient's body, such as the top of the skull, the dome of the liver, the apex of the lungs, or others. Landmarks can help in determining the location and / or size of findings within the patient's body. The previous patient can be the same person as the current patient or a different patient. The term "previous" patient means that the previous patient is correlated with a previous data record. Specifically, a previous medical image and / or previous landmarks and / or findings from a previous data record show the body of the previous patient and / or are correlated with it. Furthermore, the term "previous patient" can represent multiple patients. In the following, patient overlap can be understood as the situation where the same patient visits the same medical facility multiple times and / or the same patient visits multiple medical facilities whose medical data are merged, while the patient's medical data is assigned to different training cohorts for training and / or validating the machine learning algorithm. According to the data available to the inventor, patient overlap is estimated at approximately 3%. The database cohorts represent different data sets used to train the machine learning algorithm. At a minimum, the database cohorts include training data, validation data, and test data. The training data can be used, for example, to train the machine learning algorithm for a specific purpose, such as identifying findings in medical images. The training data can be used during the learning process to adjust parameters (e.g., weights), such as those of a classifier. The validation data can be used to fine-tune hyperparameters (i.e., the architecture) of the machine learning algorithm. The test data can be used for testing purposes related to the trained machine learning algorithm. For example, the test data can be used to evaluate performance (i.e.,The generalization of the trained machine learning algorithm may be affected. In cases where different medical images of the same person are used for different database cohorts, particularly if they show identical or similar findings and / or symptoms, the efficiency and / or accuracy of the machine learning algorithm training may be compromised. Furthermore, using the same or similar data for training and validation and / or testing purposes may violate the principle of validating and / or testing the machine learning algorithm. In cases where a patient overlap is identified, the captured medical image is assigned to one of the N2 database cohorts according to the previous dataset. If the previous dataset is assigned to the validation data, the captured medical image is consequently also assigned to the validation data. In this way, identifying patient overlaps can help ensure that all data, particularly all medical images correlated with the same patient, are assigned to the same database cohort. Therefore, identifying patient overlaps can increase the efficiency and / or accuracy of training the machine learning algorithm. According to one embodiment, the overlap detection unit is designed to identify patient overlap by determining whether the current patient is identical to the previous patient. Specifically, equality between the current patient and the previous patient occurs if at least one of a number N1 of similarity checks indicates similarities between the captured medical image and the previous data record, where N1 ≥ 1. Specifically, the N1 similarity checks include a first feature check, a second feature check, an image check, and / or a text check. The first feature check and / or the second feature check includes an examination of whether higher-order features correlated with the captured medical image and the previous dataset are identical and / or similar. The image check includes an examination of whether the captured medical image shows optical and / or visual similarity and / or similarity to the previous dataset. The text check includes an examination of whether a medical report based on the captured medical image shows similarity and / or similarity to another medical report based on the previous dataset. Alternatively, similarity between the current patient and the previous patient may be established if at least one of the N1 similarity checks indicates similarity and / or similarity between a standardized medical image and the previous dataset. According to a further embodiment, the overlap detection unit comprises a patient number assignment unit, wherein the patient number assignment unit is configured to generate a current patient number that identifies the current patient and / or the current medical facility of a number N2 of medical facilities, where N2 ≥ 1. In particular, the current patient number comprises an identification code that identifies the current medical facility and / or an identification code that identifies the current patient. In particular, the medical imaging unit is located in the current medical facility, with each of the N2 medical facilities being coupled to the common database.In particular, the previous data record is assigned to a previous medical image that was captured in the current medical facility or in another of the N2 medical facilities. The patient number assignment unit can be configured to store the current patient number in an internal system database and / or in the shared database. Specifically, the patient number assignment unit can be configured to store personal health information (PHI) relating to the current patient in the internal database, such as name, address, medical registration number, insurance company, and / or similar information. The data in the internal database should not be shared with other N2 healthcare facilities, even if the N2 healthcare facilities are linked to the data in the shared database. That is, the N2 healthcare facilities are in data communication and / or data exchange with the shared database.Medical facilities, as defined here, include hospitals, universities, institutions, and / or organizations that conduct medical examinations involving the capture of medical images. This also includes departments and / or subunits of the aforementioned entities. The previous medical image may have been captured in one of the N2 medical facilities. Consequently, the previous data record may have been generated in one of the N2 medical facilities. The previous medical image may be contained within the previous data record. According to a further embodiment, the overlap detection unit comprises a standardization unit, wherein the standardization unit is designed to standardize the captured medical image. In particular, the standardization unit is designed to normalize the intensity and / or amplitude of the captured medical image to a given range, to crop the captured medical image, to align the captured medical image, to rescan the captured medical image, and / or to subscan the captured medical image. Standardization of the captured medical image can also be understood as applying optical corrections to the captured medical image. Standardization compensates for the fact that medical images may be captured at different times, by different technicians, in different medical facilities, and / or using different medical imaging units and / or protocols. Standardization can help ensure that data originating from multiple medical facilities fall within the same acceptable range. The features relevant for identifying patient overlaps are identical between the captured medical image and the standardized medical image.Therefore, identifying a patient overlap of the standardized medical image has the same meaning as identifying a patient overlap of the captured medical image. According to a further embodiment, the system also includes a user interface for receiving a current set of landmarks and / or findings correlated with the captured medical image from a user of the system. In particular, the system further includes an alignment unit for aligning the current set of landmarks and / or findings. Specifically, the alignment unit is designed to calculate a distance and / or an affine transformation of the current set of landmarks and / or findings. In Euclidean geometry, an affine transformation is a geometric transformation that preserves straight lines and parallelism, but not necessarily Euclidean distances and angles. Examples of affine transformations include translation, scaling, homotheticity, similarity, reflection, rotation, hyperbolic rotation, shear transformation, and combinations thereof in any order. The alignment unit can be designed to adjust the current set of landmarks and / or findings so that, for example, the landmarks of different medical images are located at the same image position. As a result, similar findings exhibit congruence. This facilitates the comparison of findings for potential similarity and / or temporal evolution. The user interface can be located on the medical imaging unit or elsewhere.The user interface can also be designed to receive a command from the user to capture the medical image and / or a data removal command. The data removal command can be issued in response to a specific patient's request that all of their data be deleted and / or not used for training the machine learning algorithm. The user interface can be a computer, tablet, mobile phone, and / or other electronic device. The user interface can also be implemented as a speech recognition device. The user can be a technician and / or a physician. The user can interact with the user interface directly and / or via a network connection and / or the internet. According to a further embodiment, the overlap detection unit comprises a transformation unit, wherein the transformation unit is designed to transform the standardized medical image into a private image section and / or a shared image section. In particular, the transformation unit is designed to transform the standardized medical image by performing a Fourier analysis to derive phases and / or amplitudes, and / or by applying a trained machine learning model. Specifically, the derived phases constitute the shared image section or the private image section, while the derived amplitudes constitute the other image section, respectively. In other words, the transformation unit is designed to derive features from the standardized medical image. Some of these features form the shared image section, while the remaining features form the private image section. The transformation unit can be configured to store the private image section in the internal database and / or the shared image section in the shared database. Storing only the shared image section in the shared database ensures data privacy. Both the private and shared image sections are accessible only within the current medical facility, thus enabling the display of the standardized medical image in its entirety. According to a further embodiment, the overlap detection unit comprises a volume generation unit, wherein the volume generation unit is designed to generate a first reconstruction image by combining previous higher-order features from the previous data set with the current aligned set of landmarks and / or findings, and / or to generate a second reconstruction image by combining current higher-order features derived from the standardized medical image with a previously aligned set of landmarks and / or findings from the previous data set. The first and / or second reconstruction image represents an artificially generated medical image based on a combination of given higher-order features and a given set of landmarks and / or findings. Using this numerical data, the volume generation unit generates or reconstructs a visual medical image. The volume generation unit may be designed to execute a generative algorithm, such as conditional GANs. However, it may also be designed to execute other algorithms. The previous higher-order features and / or the previously aligned set of landmarks and / or findings were extracted from the previous medical image, for example, during a prior examination.The current higher-order features are extracted from the standardized medical image, while the current aligned set of landmarks and / or findings is correlated with the standardized medical image and specified by the user of the system, for example, a technician and / or a physician. According to another embodiment, the overlap detection unit comprises a feature extraction unit for extracting higher-order features from an input image. In particular, the feature extraction unit is designed to extract the higher-order features by applying a transformation. In particular, the feature extraction unit is designed to extract the current higher-order features from the standardized medical image, to extract first higher-order features from the first reconstruction image, and / or to extract second higher-order features from the second reconstruction image. The feature extraction unit can be designed to apply a machine learning model, such as a deep learning model, trained on large datasets applicable to a variety of use cases. The deep learning model can be a base model. The base model can be an autoencoder. The autoencoder is a type of artificial neural network used to learn efficient coding of unlabeled data (unsupervised learning). The autoencoder learns two functions: an encoding function that transforms input data and a decoding function that reconstructs the input data from the coded representation. The feature extraction unit, specifically the autoencoder, learns efficient representation of medical images for dimensional reduction.In this process, higher-order features with lower dimensions are generated based on the standardized medical image, such as edges, patterns, colors and / or the like. According to a further embodiment, the overlap detection unit comprises a testing unit for performing the first feature check, the second feature check, the image check, and / or the text check, wherein the first feature check comprises a comparison of the first higher-order reconstruction features with the current higher-order features, and wherein the second feature check comprises a comparison of the second higher-order reconstruction features with the previous higher-order features. In particular, the image check comprises a visual comparison of the standardized medical image with the first reconstruction image.In particular, the overlap detection unit includes a medical report generator for automatically generating a current medical report based on the standardized medical image and / or a medical reconstruction report based on the first reconstruction image, with the text check including a comparison of the current medical report with the medical reconstruction report. Alternatively or additionally, image review can include a visual comparison of the standardized medical image with the second reconstruction image. Similarly, the medical report generator can be designed to automatically create the medical reconstruction report based on the second reconstruction image. The term "medical report" can be understood as a text-based collection and / or summary of findings correlated with a particular medical image, similar to those written by physicians and / or technicians. Therefore, the reconstruction features can be designed to automatically identify findings in the given medical image and generate the text-based medical report, including a classification and / or specification of the findings regarding their location, size, significance, or the like. According to another embodiment, the overlap detection unit is designed to iteratively select one of the available datasets stored in the shared database as the previous dataset in order to identify the patient overlap of the standardized medical image with the available datasets.In particular, the overlap detection unit is designed to select only those of the available datasets that meet at least one of a number of selection criteria, the selection criteria comprising a location threshold based on a spatial distance between the current medical facility and one of the N2 medical facilities in which the respective available dataset was generated, a landmark threshold based on distances between landmarks of the current aligned set of landmarks and / or findings and landmarks of the respective available dataset, and / or a finding threshold based on the proximity between the findings of the current aligned set of landmarks and / or findings and the findings of the respective available dataset. Using these selection criteria, only those of the available datasets where patient overlap is highly likely are selected. This allows the process of identifying patient overlaps to be accelerated without significantly reducing performance or accuracy. According to another embodiment, the patient allocation unit is designed to update the shared database in the case where the current patient is identical to the previous patient by correlating the current patient number with a previous patient number stored in the previous record, and to assign the standardized medical image to one of the database cohorts according to the previous record, and / or in the case where the current patient is not identical to the previous patient, to assign the standardized medical image to one of the database cohorts based on a design selection and to update the shared database by creating a new record in the shared database containing the current patient number, the shared section of the standardized medical image,to update the current higher-order features and / or the current aligned set of landmarks and / or findings. In other words, if an equality is found, the standardized medical image of the same cohort in the database is assigned to the same cohort to which the previous medical image stored in the previous record was assigned. According to another embodiment, the user interface is designed to receive a data unlearning command correlated with the current patient, wherein the overlap detection unit is designed to unlearn all available records in the shared database that show a patient overlap with the current patient. Data unlearning can be achieved by deleting the data or by assigning a marker to the data indicating that it should not be used to train the machine learning algorithm. The data unlearning command can be received by the system user via the internet and / or other networks. According to another embodiment, the system is designed to train the machine learning algorithm based on the available datasets in the shared database. The machine learning algorithm can be implemented as a deep learning algorithm or using other techniques. Training can be performed using data from the shared database, particularly the available datasets. Training can be conducted using a supervised and / or unsupervised approach. The system continuously and / or periodically identifies patient overlaps and assigns the captured medical image and / or available datasets to them, thus improving the quality of the training. According to another embodiment, the system is designed to use the trained machine learning algorithm, in particular for detecting findings, anatomical structures and / or anomalies within the captured medical image. Anomalies can include, for example, nodules, blocked arteries, and / or other diseases. Because the machine learning algorithm is trained on training data that accounts for patient overlap, the trained algorithm can be more accurate and / or precise in identifying anomalies and / or anatomical structures within the captured medical image of the current patient. As a result, the current patient's treatment can be improved and / or expedited. Furthermore, technical and / or medical devices can be targeted and / or used based on the detected findings, anatomical structures, and / or anomalies. A second aspect is a method for using a system to provide training data for training a machine learning model. The system comprises a medical imaging unit, an intersection detection unit, and a patient allocation unit. Specifically, the system is implemented according to the above specifications.The process includes: the medical imaging unit capturing a medical image of a current patient; the overlap detection unit identifying any patient overlap between the captured medical image and a previous record stored in a shared database and correlated with a previous patient; and, if patient overlap is identified, the patient allocation unit assigning the captured medical image to one of a number of database cohorts according to the previous record for training the machine learning model. The entity in question, e.g., the patient allocation unit, can be implemented in hardware and / or in software. If the entity is implemented in hardware, it can be embodied as a device, e.g., a computer or a processor, or as part of a system, e.g., a computer system. If the entity is implemented in software, it can be embodied as a computer program product, a function, a routine, program code, or an executable object. Each embodiment of the first aspect can be combined with any embodiment of the second aspect to obtain a further embodiment of the second aspect. According to a third aspect, the invention relates to a computer program product comprising program code for executing the method described above when executed on at least one computer. The embodiments and features described with reference to the proposed system apply accordingly to the proposed method and / or the proposed computer program product. Other possible implementations or alternative solutions of the invention also include combinations—not expressly mentioned here—of the features described above or below in relation to the embodiments. Those skilled in the art may also add individual or isolated aspects and features to the most basic form of the invention. Further embodiments, features and advantages of the present invention will become apparent from the following description and the dependent claims in conjunction with the accompanying drawings; which show: Fig. 1 a schematic representation of an embodiment of a system for providing training data for training a machine learning model; Fig. 2 a schematic representation of data flows that occur when using the system according to Fig. 1; and Fig. 3 a schematic block diagram of an embodiment of a method for using the system according to Fig. 1. In the figures, identical reference symbols denote identical or functionally equivalent elements, unless otherwise specified. Fig. 1 shows a schematic representation of an embodiment of a system 1 for improving training data not shown for training a machine learning algorithm not shown for analyzing medical images. System 1 comprises a medical imaging unit 2 for acquiring a medical image MI of a current patient CP. The medical imaging unit 2 can be, for example, a computed tomography (CT) scanner, a magnetic resonance imaging (MRI) scanner, an X-ray scanner, or the like. The acquired medical image MI can, for example, show one or more sections or slices of the current patient CP. System 1 also comprises a user interface 3 for receiving user commands from a user of System 1 (not shown). The user commands can include a command to acquire the medical image MI, to detect a possible patient overlap, to unlearn data correlated with the current patient CP, or other commands.User Interface 3 is further designed to receive a current set of landmarks and / or findings CL that correlate with the acquired medical image MI. Landmarks can be, for example, the skullcap, the dome of the liver, the apex of the lungs, or others. The landmarks are used to specify the location of findings within a patient's body. Findings can be, for example, medical structures and / or anomalies, such as nodules. The current set of landmarks and / or findings CL can be specified by the user of System 1, for example, a physician or a technician. The medical imaging unit 2 and User Interface 3 are located in a current medical facility 4. The current medical facility 4 can be a hospital, a university, a corresponding department of a hospital or university, or the like.System 1 may be located in the current medical facility 4, in another medical facility not shown, and / or on a cloud server not shown. System 1 further includes an internal database 5 located within the current medical facility 4. The internal database 5 is used to store all of the current patient's (CP) private health information (PHI), such as the current patient's name, medical registration number (MRM), and similar information. For privacy reasons, this data should not be shared outside of the current medical facility 4. The internal database 5 may be a hard drive and / or a server located within the current medical facility 4. System 1 further includes a shared database 6, which may be stored in the current medical facility 4, in another medical facility not shown, on a cloud server, and / or in other locations. The shared database 6 contains a number of available data records 7 that can be used to train the machine learning algorithm. These available data records 7 can be accessed by any medical facility, including the current medical facility 4. In other words, the shared database 6 is used to exchange data between all of the medical facilities. System 1 further comprises an overlap detection unit 8 for detecting a possible patient overlap between the acquired medical image MI and the available data records 7 stored in the shared database 6. The components of the overlap detection unit 8 are explained below with reference to Fig. 1. A data flow generated by the overlap detection unit 8 for identifying the patient overlap is explained in detail with reference to Fig. 2. In Fig. 2, multiple representations of features do not indicate multiple existence of the features, but rather multiple use of the same feature. Fig. 1 and Fig. 2 are referred to simultaneously below. The individual steps performed by the overlap detection unit 8 are explained in detail below with reference to the block diagram in Fig. 3. The overlap detection unit 8 includes a patient number assignment unit 9. The patient number assignment unit 9 is designed to identify whether the current patient CP has already been examined at the current medical facility 4. To this end, the patient number assignment unit 9 is designed to access the internal database 5 to detect previous examinations of the current patient CP. If the current patient CP has already been examined at the current medical facility 4, the captured medical image MI is assigned to the same cohort in the database as in the context of the current patient CP's previous examinations. If the current patient CP is new to the current medical facility 4, there may still be previous examinations of the current patient CP at other medical facilities.To detect such patient overlaps, the patient number assignment unit 9 is designed to generate a current patient number 10 (see Fig. 2) that includes a signature to identify the current medical facility 4 and a signature to identify the current patient CP. For example, the current patient number 10 might be “1-276”, where the digit “1” identifies the current medical facility 4 and the digits “276” identify the current patient CP. The current patient number 10 is not correlated with the current patient CP’s medical registration number or any other official numbers protected by data privacy laws.The patient number assignment unit 9 is designed to store the current patient number 10 in the internal database 5 along with the current patient CP's private health information (PHI), such as the patient CP's name, medical registration number, or similar information. Furthermore, the patient number assignment unit 9 is designed to store the current patient number 10 in the shared database 6 (see Fig. 2). The overlap detection unit 8 further includes an alignment unit 11 for aligning the current set of landmarks and / or findings CL. The alignment includes calculating the distances of the respective landmarks visible in the captured medical image MI and calculating their respective affine transformations. Alternatively or additionally, the alignment includes calculating the distances of the respective findings correlated with the captured medical image MI and calculating the affine transformations of the respective findings. The purpose of the alignment is to numerically transform the positions and / or sizes of the landmarks and / or findings for further numerical processing of the data, including comparison with other data stored in the shared database 6.In the following, the aligned current set of landmarks and / or findings CL will be referred to as the current aligned set of landmarks and / or findings CAL. The overlap detection unit 8 further includes a standardization unit 12, which is designed to standardize the acquired medical image (MI). Standardization includes, for example, normalizing the acquired medical image (MI) to a specific range. It may include data processing steps such as intensity and / or amplitude normalization, cropping, rescanning, and subsampling of the acquired medical image (MI), or others. The purpose of standardization is to account for the effect that similar medical images may vary in intensity or other parameters. Furthermore, patients may be positioned differently during each examination with the medical imaging unit 2. Hereinafter, the output of the standardization unit 12 is referred to as the standardized medical image (SMI). The overlap detection unit 8 further comprises a transformation unit 13. The transformation unit 13 is designed to transform the standardized medical image SMI into a private image section PP, which is stored in the internal database 5, and a shared image section SP, which is stored in the shared database 6. The transformation unit 13 can transform the standardized medical image SMI by performing a Fourier analysis to derive phases and amplitudes and / or by applying a machine learning model for the transformation. The derived phases can form the shared image section SP, while the derived amplitudes can form the private image section PP. However, this assignment can also be reversed.Data protection is ensured by storing only a section of the standardized medical image SMI in the shared database 6. The overlap detection unit 8 further comprises a feature extraction unit 14. The feature extraction unit 14 is designed to extract higher-order features from a given input image. Specifically, the feature extraction unit 14 is designed to extract these higher-order features by applying a basic model, in particular an autoencoder. These higher-order features can include, for example, edges, patterns, colors, textures, and / or other elements of the given input image. The feature extraction unit 14 is designed to extract a number of current higher-order features (CHF) from the standardized medical image (SMI). The extracted current higher-order features (CHF) are stored in the shared database 6. The overlap detection unit 8 further comprises a volume generation unit 15. The volume generation unit 15 can execute a generative algorithm, such as conditional GANs, to generate an artificial visual medical image based on given numerical data, in particular based on a combination of given higher-order features and landmarks and / or findings. In simplified terms, the volume generation unit 15 is designed to derive information from the higher-order features about which elements are visible in the generated medical image, while it derives information from the landmarks and / or findings about where these elements are located in the generated medical image. The overlap detection unit 8 further includes a medical report generator 16 for automatically generating a text-based medical report based on a given medical image. In particular, the medical report generator 16 is designed to automatically identify findings in the given medical image and to generate the text-based medical report, which includes a classification and / or a specification of the findings with regard to their location, size, significance, or other aspects. The overlap detection unit 8 further includes a verification unit 17 for performing a comparison between two or more input data sets and for identifying similarities between the input data sets. The input data sets can be higher-order features, medical images, and / or text-based medical reports. To identify a patient overlap, the overlap detection unit 8 is designed to select one of the available records 7 from the shared database 6. The overlap detection unit 8 can be configured to select only those of the available records 7 that meet at least one of a number of selection criteria. These selection criteria can include a location threshold based on the spatial distance between the current medical facility 4 and the specific medical facility where the respective available record 7 was created. In other words, only those of the available records 7 that were created near the current medical facility 4 might be selected.The selection criteria may further include a landmark threshold based on the distances between landmarks in the current aligned set of landmarks and / or findings (CAL) and landmarks in the respective available dataset 7. The selection criteria may also include a finding threshold based on the proximity between findings in the current aligned set of landmarks and / or findings (CAL) and findings in the respective available dataset 7. The purpose of the selection criteria is to expedite the identification of patient overlaps by comparing the standardized medical image (SMI) only with those available datasets 7 where a potential patient overlap is likely. Subsequently, the standardized medical image (SMI) is iteratively compared with all selected available datasets 7, i.e., sequentially.In the following, the selected of the available data sets 7 will be referred to as the previous data set 18. After selecting the previous dataset 18, the volume generation unit 15 is configured to load the previous higher-order features PHF stored in the previous dataset 18. Furthermore, the volume generation unit 15 is configured to load the current aligned set of landmarks and / or findings CAL. The volume generation unit 15 is also configured to generate an artificial medical image based on the loaded previous higher-order features PHF and the current aligned set of landmarks and / or findings CAL. The generated medical image is referred to as the first reconstruction image FRI. The feature extraction unit 14 is configured to extract first higher-order reconstruction features FRHF from the generated first reconstruction image FRI.Subsequently, test unit 17 is designed to compare the first higher-order reconstruction features FRHF with the current higher-order features CHF. Simultaneously or subsequently, the volume generation unit 15 is designed to load the previously aligned set of landmarks and / or findings PAL stored in the previous dataset 18. Furthermore, the volume generation unit 15 is designed to load the current higher-order features CHF. Then, the volume generation unit 15 is designed to generate another artificial medical image based on the previously aligned set of landmarks and / or findings PAL and the current higher-order features CHF. The additional medical image generated is referred to as the second reconstruction image SRI. Consequently, the feature extraction unit 14 is designed to extract second higher-order reconstruction features SRHF from the second reconstruction image SRI.Subsequently, the test unit 17 is designed to compare the second higher-order reconstruction features SRHF with the previous higher-order features PHF. Simultaneously, subsequently, or in the event that the previous two comparisons of test unit 17 show a similarity, test unit 17 is designed to compare the standardized medical image SMI with the first reconstruction image FRI. Simultaneously, subsequently, or if the previous three comparisons of the test unit 17 show a similarity, the medical report generator 16 is designed to automatically create a current medical report CMR based on the standardized medical image SMI. Furthermore, the medical report generator 16 is designed to automatically create a medical reconstruction report RMR based on the first reconstruction image FRI. Then, the test unit 17 is designed to compare the current medical report CMR with the medical reconstruction report RMR. In the event that at least one, two, three, or all of the comparisons performed by the testing unit 17 show a similarity, the standardized medical image SMI is assumed to represent a patient overlap with the previous data set 18. Consequently, the current patient CP is considered identical to an unseen previous patient who correlates with the previous data set 18. System 1 further includes a patient allocation unit 19. Patient allocation unit 19 is designed, in the case where the current patient CP is considered identical to the previous patient, to update the shared database 6 by correlating the current patient number 10 with an unseen previous patient number that correlates with the previous patient and / or is stored in the previous record 18, adding the current patient number 10 to the previous record 18, and assigning the standardized medical image SMI to one of the database cohorts according to the previous record 18. In other words, if the previous record 18 is assigned to the validation data, the standardized medical image SMI is also assigned to the validation data.Alternatively or additionally, the patient allocation unit 19 is designed, in the event that the current patient CP is not identical to the previous patient, to update the shared database 6 by creating a new record 20 in the shared database 6, which includes the current patient number 10, the shared section of the standardized medical image SMI, the current aligned higher-order features CHF, and / or the current aligned set of landmarks and / or findings CAL. Furthermore, the patient allocation unit 19 is designed to assign the standardized medical image SMI to one of the database cohorts for training the machine learning algorithm, based on a predefined design choice. The user interface 3 can be configured to receive an unseen data learning command. Upon receiving this data learning command, the overlap detection unit 8 is configured to identify all patient overlaps of the respective patient requesting the data learning command with all available records 7 in the shared database 6. Then, all identified available records 7 that have a patient overlap with the respective patient can either be flagged to prevent their use in training the machine learning algorithm, and / or the identified records can be deleted from the shared database 6.This ensures that the respective patient data is not used for training the machine learning algorithm, not only in the current medical facility 4, but also in all medical facilities accessing the shared database 6. Established unlearning techniques used in the medical facilities, such as joint, isolated, decomposed, and unified training, can be employed to unlearn the respective patient data. Fig. 3 shows a schematic block diagram for the use of System 1. In a first step (S1), the current patient number 10 is generated by the patient number assignment unit 9. In a further step (S2), the medical image MI is acquired by the medical imaging unit 2. In a further step (S3), the current set of landmarks and / or findings CL is received by the user interface 3. In a further step (S4), the current set of landmarks and / or findings CL is aligned by the alignment unit 11 to form the current aligned set of landmarks and / or findings CAL. In a further step (S5), the acquired medical image MI is standardized by the standardization unit 12 to form the standardized medical image SMI. In a further step (S6), the standardized medical image SMI is transformed by the transformation unit 13 to form the private image section PP and the shared image section SP.In a further step S7, the current higher-order features CAF are extracted by the feature extraction unit 14. In a further step S8, one of the available datasets 7, stored in the shared database 6, is selected to form the previous dataset 18. Additionally, the previously aligned set of landmarks and / or findings PAL and previous higher-order features PHF are loaded from the selected previous dataset 18. In a further step S9A, the first reconstruction image FRI is generated by the volume generation unit 15 based on the currently aligned set of landmarks and / or findings CAL and the previous higher-order features PHF. In a further step S10A, the first higher-order reconstruction features FRHF are extracted from the first reconstruction image FRI by the feature extraction unit 14. In a further step S11A, the first feature verification is performed by comparing the first higher-order reconstruction features FRHF with the current higher-order features CHF by the verification unit 17. In a further step S9B, the second reconstruction image SRI is generated by the volume generation unit 15 based on the previously aligned set of landmarks and / or findings PAL and the current higher-order features CHF. In a further step S10B, the second higher-order reconstruction features SRHF are extracted from the second reconstruction image SRI by the feature extraction unit 14. In a further step S11B, the second feature verification is performed by comparing the previous higher-order features PHF with the second higher-order reconstruction features SRHF by the verification unit 17. In a further step S11C, the image check is carried out by comparing the standardized medical image SMI and the first reconstruction image FRI by the test unit 17. In a further step, S9D, the current medical report CMR is automatically generated by the medical report generator 16 based on the standardized medical image SMI. In a further step, S10D, the medical reconstruction report RMR is automatically generated by the medical report generator 16 based on the first reconstruction image FRI. In a further step, S11D, the text is checked by comparing the current medical report CMR with the medical reconstruction report RMR using the verification unit 17. The four branches mentioned above, S9A-S11A, S9B-S11B, S11C and / or S9D-S11D, can be executed simultaneously or sequentially. In a further step S12A, the results of the first feature check S11A, the second feature check S11B, the image check S11C, and / or the text check S11D are analyzed by the check unit 17 to determine any possible patient overlap of the standardized medical image SMI with the previous data record 18. If a patient overlap is identified, the shared database 6 is updated in a further step S13A by adding the current patient number 10 to the previous data record 18. The procedure then continues in a step S14, which is explained below.If no patient overlap of the standardized medical image SMI with the previous dataset 18 is identified, step S12B determines whether any of the available datasets 7 can be analyzed for a possible patient overlap with the standardized medical image SMI. If so, the previous routine is repeated starting with step S8, and a possible patient overlap between the standardized medical image SMI and the other available dataset 7 is analyzed. If all available datasets 7 have been analyzed for a possible patient overlap, possibly taking the selection criteria into account, a new dataset 20 is created in the shared database 6 in a further step S13B. In step S14, system 1 checks whether the data learning command was entered via user interface 3. If the data learning command was given, the patient data is unlearned from the shared database 6 in step S15. Otherwise, in a further step S16, the machine learning algorithm is trained using all available data records 7 in the shared database 6. Although the present invention has been described in accordance with preferred embodiments, it is obvious to those skilled in the art that variations are possible in all embodiments. REFERENCE MARK 1 System 2 Medical Imaging Unit 3 User Interface 4 Current Medical Facility 5 Internal Database 6 Shared Database 7 Available Records 8 Overlap Detection Unit 9 Patient Number Assignment Unit 10 Current Patient Number 11 Alignment Unit 12 Standardization Unit 13 Transformation Unit 14 Feature Extraction Unit 15 Volume Generation Unit 16 Medical Report Generator 17 Verification Unit 18 Previous Record 19 Patient Assignment Unit 20 New Record CAL Current Aligned Set of Landmarks and / or Findings CHF Current Higher-Order Features CL Current Set of Landmarks and / or Findings CP Current Patient FRHF First Higher-Order Reconstruction Features FRI First Reconstruction Image MI Medical Image PAL Previous Aligned Set of Landmarks and / or Findings PHF Previous Higher-Order Features PP Private Image Section RMR Medical Reconstruction Report SMIStandardized medical image S1-S16 step SP shared image section SRHF second higher-order reconstruction features SRI second reconstruction image
Claims
System (1) for improving training data for training a machine learning model, wherein the system (1) comprises: a medical imaging unit (2) for acquiring a medical image (MI) of a current patient (CP); an overlap detection unit (8) for identifying a patient overlap of the acquired medical image (MI) and a previous data set (18) stored in a shared database (6) and correlated with a previous patient; and a patient allocation unit (19) for assigning the acquired medical image (MI) to one of a number of database cohorts according to the previous data set (18) for training the machine learning model in the event that the patient overlap is identified. System according to claim 1, wherein the overlap detection unit (8) is designed to identify patient overlap by determining whether the current patient (CP) is identical to the previous patient, in particular that there is equality between the current patient (CP) and the previous patient if at least one of a number N1 of similarity checks shows similarities between the captured medical image (MI) and the previous data record (18), wherein N1 ≥ 1, wherein in particular the N1 similarity checks comprise a first feature check, a second feature check, an image check and / or a text check. System according to claim 1 or 2, wherein the overlap detection unit (8) comprises a patient number assignment unit (9), wherein the patient number assignment unit (9) is configured to generate a current patient number (10) that identifies the current patient (CP) and / or a current medical facility (4) from a number N2 of medical facilities, with N2 ≥ 1, in particular wherein the current patient number (10) comprises an identification code that identifies the current medical facility (4) and / or an identification code that identifies the current patient (CP), in particular wherein the medical imaging unit (2) is located in the current medical facility (4), wherein each of the N2 medical facilities is coupled to the shared database (6), in particular wherein the previous record (18) is assigned to a previous medical image.that was admitted to the current medical facility (4) or to another of the N2 medical facilities. System according to claim 3, wherein the overlap detection unit (8) comprises a standardization unit (12), wherein the standardization unit (12) is designed to standardize the captured medical image (MI), in particular wherein the standardization unit (12) is designed to normalize the intensity and / or the amplitude of the captured medical image (MI) to a given range, to crop the captured medical image (MI), to align the captured medical image (MI), to rescan the captured medical image (MI) and / or to subscan the captured medical image (MI). System according to claim 3 or 4, wherein the system (1) further comprises a user interface (3) for receiving a current set of landmarks and / or findings (CL) correlated with the captured medical image (MI) from a user of the system (1), in particular wherein the system (1) further comprises an alignment unit (11) for aligning the current set of landmarks and / or findings (CL), in particular wherein the alignment unit (11) is designed to calculate a distance and / or an affine transformation of the current set of landmarks and / or findings (CL). System according to claim 4 or 5, wherein the overlap detection unit (8) comprises a transformation unit (13), wherein the transformation unit (13) is designed to transform the standardized medical image (SMI) into a private image section (PP) and / or a shared image section (SP), in particular wherein the transformation unit (13) is designed to transform the standardized medical image (SMI) by performing a Fourier analysis to derive phases and / or amplitudes and / or by applying a trained machine learning model for transformation, in particular wherein the derived phases form the shared image section (SP) or the private image section (PP), while the derived amplitudes form the other image section. System according to claim 5 or 6, wherein the overlap detection unit (8) comprises a volume generation unit (15), the volume generation unit (15) being configured to generate a first reconstruction image (FRI) by combining higher-order features (PHF) from the previous data set (18) with the current aligned set of landmarks and / or findings (CAL) and / or to generate a second reconstruction image (SRI) by combining current higher-order features (CHF) derived from the standardized medical image (SMI) with a previous aligned set of landmarks and / or findings (PAL) from the previous data set (18). System according to claim 7, wherein the overlap detection unit (8) comprises a feature extraction unit (14) for extracting higher-order features from an input image, in particular wherein the feature extraction unit (14) is designed to extract the higher-order features by applying a transformation, in particular wherein the feature extraction unit (14) is designed to extract the actual higher-order features (CHF) from the standardized medical image (SMI), to extract first higher-order reconstruction features (FRHF) from the first reconstruction image (FRI) and / or to extract second higher-order reconstruction features (SRHF) from the second reconstruction image (SRI). System according to claim 8, wherein the overlap detection unit (8) comprises a test unit (17) for performing the first feature check, the second feature check, the image check and / or the text check, wherein the first feature check comprises a comparison of the first higher-order reconstruction features (FRHF) with the current higher-order features (CHF), and wherein the second feature check comprises a comparison of the second higher-order reconstruction features (SRHF) with the previous higher-order features (PHF), in particular wherein the image check comprises a comparison of the standardized medical image (SMI) with the first reconstruction image (FRI),in particular wherein the overlap detection unit (8) comprises a medical report generator (16) for automatically generating an up-to-date medical report based on the standardized medical image (SMI) and / or a medical reconstruction report (RMR) based on the first reconstruction image (FRI), wherein the text check comprises a comparison of the up-to-date medical report with the medical reconstruction report (RMR). System according to any one of claims 4 to 9, wherein the overlap detection unit (8) is designed to iteratively select one of the available datasets (7) stored in the shared database (6) as the previous dataset (18) in order to identify the patient overlap of the standardized medical image (SMI) with the available datasets (7), in particular wherein the overlap detection unit (8) is designed to select only those of the available datasets (7) that satisfy at least one of a number of selection criteria, wherein the selection criteria are a location threshold based on a spatial distance between the current medical facility (4) and one of the N2 medical facilities in which the respective available dataset was generated, a landmark threshold,which are based on distances between landmarks of the current aligned set of landmarks and / or findings (CAL) and landmarks of the respective available dataset, and / or a finding threshold based on proximity between the findings of the current aligned set of landmarks and / or findings (CAL) and the findings of the respective available dataset. System according to any one of claims 7 to 10, wherein the patient allocation unit (19) is configured, in the case where the current patient (CP) is identical to the previous patient, to update the shared database (6) by correlating the current patient number (10) with a previous patient number stored in the previous record (18) and to assign the standardized medical image (SMI) to one of the database cohorts according to the previous record (18), and / or, in the case where the current patient (CP) is not identical to the previous patient, to assign the standardized medical image (SMI) to one of the database cohorts based on a design selection and to update the shared database (6) by creating a new record (20) in the shared database (6) containing the current patient number (10), the shared section of the standardized medical image (SMI),includes the current higher-order features (CHF) and / or the current aligned set of landmarks and / or findings (CAL). System according to claim 10 or 11, wherein the user interface (3) is designed to receive a data unlearning command correlated with the current patient (CP), and wherein the overlap detection unit is designed to delete all available data records (7) in the shared database (6) that show a patient overlap with the current patient (CP). System according to any one of claims 1 to 12, wherein the system (1) is designed to train the machine learning algorithm on the basis of the available data sets (7) in the shared database (6). System according to claim 13, wherein the system (1) is designed to use the trained machine learning algorithm, in particular for detecting findings, anatomical structures and / or anomalies within the captured medical image (MI). Method for using a system (1) for providing training data for training a machine learning model with a medical imaging unit (2), an overlap detection unit (8) and a patient assignment unit (19), wherein the system (1) is implemented in particular according to any one of claims 1 to 14, wherein the method comprises: capturing a medical image (MI) of a current patient (CP) by the medical imaging unit (2); identifying a patient overlap of the captured medical image (MI) with a previous data set (18) stored in a shared database (6) and correlated with a previous patient by the overlap detection unit (8);and in the case where patient overlap is identified, the patient allocation unit (19) assigns the captured medical image (MI) to one of a number of database cohorts according to the previous dataset (18) for training the machine learning model.
Citation Information
Patent Citations
Systems and methods for associating medical images with a patient
US20190237198A1
Determining image similarity by analysing registrations
US20230087494A1