A system and method for anonymizing medical data in a health analysis platform.
A two-stage anonymization process for medical data addresses privacy and confidentiality issues by masking identifying and clinical study information, enhancing data privacy and reducing resource consumption.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MEDIDATA SOLUTIONS INC
- Filing Date
- 2025-10-17
- Publication Date
- 2026-05-19
AI Technical Summary
Existing medical data anonymization methods fail to adequately protect patient privacy and confidentiality, as they do not sufficiently mask identifying information and clinical study details, leading to potential breaches and unauthorized access.
A system and method that performs two stages of anonymization: a first anonymization operation on directly identifying information and a second on clinical research information, using specific rules to modify and remap data fields and values, generating a standardized dataset.
Enhances privacy and fidelity of clinical research data, reducing the risk of breaches and minimizing computational resources required to address violations, ensuring consistent and efficient data sharing.
Smart Images

Figure 2026082708000001_ABST
Abstract
Description
Technical Field
[0001] This specification generally relates to systems and methods for anonymizing medical data in a health analysis platform by evaluating and modifying physiological measurements of filtered medical data.
Background Art
[0002] Generally, maintaining the privacy of medical data, including clinical trial data, is important for protecting patient information and preventing unauthorized access to confidential data. Privacy violations can lead to harmful consequences, including the improper use and distribution of an individual's health information. To mitigate these risks, anonymization or data masking techniques may be required to protect and maintain privacy in medical data.
Summary of the Invention
[0003] The implementation provided in this disclosure includes a system for improving data security in a computerized health analytics platform. The system includes a hardware storage device, at least one processor, and a memory subsystem communicatively coupled to the at least one processor. The memory subsystem, when executed by the at least one processor, stores instructions causing the at least one processor to perform an operation comprising: (i) reading clinical research datasets relating to one or more clinical studies from the hardware storage device; (ii) performing a first computer-executable anonymization function on each of the clinical research datasets to generate a modified clinical research dataset; (iii) generating a standardized dataset based on the modified clinical research datasets; (iv) performing a second computer-executable anonymization function on the standardized dataset to generate a modified standardized dataset; (v) generating a first data structure representing the modified standardized dataset; and (vi) storing the first data structure using the hardware storage device. Each clinical research dataset includes (i) one or more first data fields that store one or more first values representing identifying information associated with one or more clinical research and entities participating in one or more clinical research, and (ii) one or more second data fields that store one or more second values representing clinical research information collected during one or more clinical research. Performing a first computer-executable anonymization function involves masking one or more first values in the clinical research dataset. Performing a first computer-executable anonymization function involves modifying one or more first values in the clinical research dataset according to a first set of rules in order to generate an anonymized representation of the identifying information.Generating a standardized dataset involves concatenating a modified clinical study dataset into a standardized dataset and at least one of the following: (i) remapping at least one of the first data fields of the standardized dataset to the respective standardized first data field; (ii) remapping at least one of the second fields of the standardized dataset to the respective standardized second data field; (iii) remapping at least one of the first values of the standardized dataset to the respective standardized first value; or (iv) remapping at least one of the second values of the standardized dataset to the respective standardized second value. Performing a second computer-executable anonymization function involves masking at least one of the study design, data collection, or treatment information of one or more clinical studies. Performing a second computer-executable anonymization function involves modifying at least a portion of the standardized dataset according to a second set of rules in order to generate an anonymized representation of the clinical study information.
[0004] The implementation provided in this disclosure includes a method for improving data security in a computerized health analytics platform. The method includes (i) reading clinical research datasets relating to one or more clinical studies from a hardware storage device; (ii) performing a first computer-executable anonymization function on direct identifiers contained in each of the clinical research datasets to generate a modified clinical research dataset; (iii) generating a standardized dataset based on the modified clinical research datasets; (iv) performing a second computer-executable anonymization function on the clinical research information portion contained in the standardized dataset to generate a modified standardized dataset; (v) generating a first data structure representing the modified standardized dataset; and (vi) storing the first data structure in a hardware storage device. Each clinical research dataset includes (i) one or more first data fields that store one or more first values representing identification information associated with one or more clinical studies and entities participating in one or more clinical studies; and (ii) one or more second data fields that store one or more second values representing clinical research information collected during one or more clinical studies. Performing a first computer-executable anonymization function involves masking one or more first values in a clinical research dataset. Performing a first computer-executable anonymization function also involves modifying one or more first values in a clinical research dataset according to a first set of rules in order to generate an anonymized representation of the identifying information.Generating a standardized dataset involves concatenating a modified clinical study dataset into a standardized dataset and at least one of the following: (i) remapping at least one of the first data fields of the standardized dataset to the respective standardized first data field; (ii) remapping at least one of the second fields of the standardized dataset to the respective standardized second data field; (iii) remapping at least one of the first values of the standardized dataset to the respective standardized first value; or (iv) remapping at least one of the second values of the standardized dataset to the respective standardized second value. Performing a second computer-executable anonymization function involves masking at least one of the study design, data collection, or treatment information of one or more clinical studies. Performing a second computer-executable anonymization function involves modifying at least a portion of the standardized dataset according to a second set of rules in order to generate an anonymized representation of the clinical study information.
[0005] Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions or operations described herein. One or more computer systems can be configured to perform a particular action by installing software, firmware, hardware, or a combination thereof on the system that causes the system to perform an action during operation. One or more computer programs can be configured to perform a particular action by including instructions that cause the device to perform an action when executed by a data processing device.
[0006] Details of one or more embodiments of the subject matter of this specification are described in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]
[0007] [Figure 1] This figure shows an example of a medical data anonymization system.
[0008] [Figure 2] This figure shows an example implementation of data anonymization software or algorithms used by electronic devices.
[0009] [Figure 3] This is a flowchart illustrating an exemplary process for anonymizing medical data.
[0010] [Figure 4] This is a flowchart illustrating an exemplary process for anonymizing medical data.
[0011] [Figure 5] This is a diagram illustrating an exemplary computer system. [Modes for carrying out the invention]
[0012] Similar reference numbers and names in various drawings refer to the same elements.
[0013] The privacy of clinical research data (e.g., clinical trial data) is important in various aspects of healthcare. For example, privacy is essential to protect patient identification. Furthermore, if not protected, a privacy breach could lead to the disclosure of other sensitive information, such as the efficacy of a product or drug, trial results and data, entity demographics, trial dates, or other information identifying the clinical research (e.g., clinical trial). Such information can be confidential not only for the patient's privacy but also for stakeholders, including sponsors of the clinical research involved in the trial.
[0014] Furthermore, even if data related to the identification of a clinical study is anonymized, some clinical study information, such as study design, data collection, and treatment information, may still be identifiable. For example, some clinical study information may be identifiable based on the study design (including references to specific time points regarding the patient being treated, treatment groups, visit schedules, administration stages, products, trials, or tests), data collection (including trial or test data, biomarkers, vital signs, etc.), and / or treatment information (including treatment names and treatment-related information).
[0015] Therefore, to enhance the privacy and fidelity of clinical research data, it may be necessary to anonymize both clinical research identification information and clinical research information at multiple stages. Furthermore, medical evaluations, treatment decisions, or medical research that rely on such clinical research data can benefit from mitigating privacy and confidentiality breaches and minimizing potential harm if the data is made public to third parties or the public.
[0016] The implementation provided in this disclosure addresses the above-mentioned data problem by at least [1] performing a first anonymization operation on medical data (e.g., clinical research data, real-time data, other medical data) by modifying at least the directly identifying information portion of the medical data (e.g., corresponding values in data fields) according to a first set of rules; [2] generating standardized medical data; and [3] performing a second anonymization operation on the standardized medical data by at least modifying the clinical research information portion of the standardized medical data (e.g., corresponding values in data fields) according to a second set of rules.
[0017] For example, the directly identifying information portion may include demographic information, entities (e.g., individuals, users, patients, subjects, etc.), a clinical study, or identifiers associated with parties related to a clinical study, location information related to a clinical study, treatment information, date information associated with a clinical study, or other directly identifying identifiers that can be used to identify entities, clinical studies, locations, or parties associated with a clinical study. Furthermore, for example, the clinical study information portion may include study design, data collection, treatment information, or information that does not directly identify the clinical study.
[0018] The direct identification information portion may be included as or represented as the first value in the first data field of the medical data. For example, the clinical research information portion may be included as or represented as the second value in the second data field of the medical data.
[0019] Performing a first anonymization operation may include performing a first computer-executable anonymization function that can mask a first value in a first data field of medical data related to a directly identifying portion. For example, the first computer-executable anonymization function may, based on a first set of rules, modify one or more of the first values associated with the directly identifying information in order to generate an anonymized representation of the directly identifying information.
[0020] Generating a standardized dataset can include concatenating a modified medical dataset (e.g., modified based on a first anonymization operation) to the standardized dataset and at least one of: (i) remapping at least one of the first data fields of the standardized dataset to respective standardized first data fields, (ii) remapping at least one of the second fields of the standardized dataset to respective standardized second data fields, (iii) remapping at least one of the first values of the standardized dataset to respective standardized first values, or (iv) remapping at least one of the second values of the standardized dataset to respective standardized second values.
[0021] Performing a second anonymization operation can include performing a second computer-executable anonymization function on the standardized dataset to generate a modified standardized dataset. For example, the second computer-executable anonymization function can modify at least a portion of the standardized dataset to generate an anonymized representation of clinical research information based on a second set of rules.
[0022] For example, the second set of rules can include rules for modifying one or more of the second values associated with the clinical research information portion.
[0023] For example, the second computer-executable anonymization function can include performing data aggregation and / or data cohorting on the standardized medical data before and / or after modifying a portion of the standardized medical data based on the second set of rules.
[0024] Once the anonymization operation is complete, a data structure representing the modified and standardized dataset can be generated. Such a data structure can be stored in a data store, transmitted to a computerized health analysis platform, or output to a user interface for display.
[0025] By doing so (e.g., performing multiple stages of anonymizing both direct identifying information and clinical research information of medical data), the privacy and fidelity of clinical research data or medical data can be enhanced.
[0026] Furthermore, based on the enhanced privacy and fidelity of clinical research data or medical data, privacy and confidentiality violations can be reduced, minimizing potential damages that may occur if such data is accessed by unauthorized third parties or made publicly available.
[0027] Additionally, based on the enhanced privacy and fidelity of medical data and / or the reduced privacy and confidentiality violations, the embodiments described herein can also reduce or eliminate the expenditure of computer resources consumed to address potential violations. For example, addressing violations often requires significant CPU usage, network bandwidth, memory, and storage resources, all of which can increase the computational load during the processing and analysis of clinical research data, medical data, or other data related to the violation. The embodiments described herein can prevent the occurrence of such violations, thereby minimizing the computer resources (e.g., CPU usage, memory, storage, etc.) required to handle such incidents.
[0028] Furthermore, once a data structure representing a modified, standardized dataset is transmitted (for example, to a computerized health analytics platform), such a data structure can be efficiently shared with various parties that have access to it via a computer network. For example, because the first data structure includes a standardized data structure, the data is compatible with various systems, ensuring consistency in data processing, thus promoting more efficient and reliable data processing.
[0029] Figure 1 shows an example of a medical data anonymization system 100. Specifically, the medical data anonymization system 100 performs the following steps: [1] a first anonymization operation on medical data, [2] standardized medical data, and [3] a second anonymization operation on the standardized medical data.
[0030] The medical data anonymization system 100 may include electronic devices 120 and sensor devices 110 that are connected to each other in a communicative manner (for example, via one or more wired or wireless communication links 150). Generally, the medical data anonymization system 100 can access data structures (for example, medical data such as clinical research data stored in a data store including a database module 122, or accessible to the electronic devices 120, for example, through a server) and perform anonymization of the information contained in such data structures through processing methods according to the implementations described herein. Furthermore, in some implementations, the medical data anonymization system 100 may use the sensor devices 110 to acquire sensor data about a user and use the electronic devices 120 to process the sensor data.
[0031] Generally, the electronic device 120 can include any number of devices configured to receive, process, and transmit data. Examples of the electronic device 120 include client computing devices (e.g., desktop or notebook computers), server computing devices (e.g., server computers or cloud computing systems), mobile computing devices (e.g., cellular phones, smartphones, tablets, personal digital assistants, and networked notebook computers), wearable computing devices (e.g., smartphones or headsets), and other computing devices capable of receiving, processing, and transmitting data. In some implementations, the electronic device 120 can include computing devices that operate using one or more operating systems (e.g., Microsoft Windows, Apple macOS, Linux, Unix, Google Android, and Apple iOS, among others) and one or more architectures (e.g., x86, PowerPC, and ARM, among others).
[0032] The sensor device 110 includes one or more sensors 112 configured to acquire measurements of the user's physiological functions, user behavior, and / or any other characteristics of the user. For example, the sensor device 110 may include or correspond to a wearable device (e.g., a smartwatch), a smartphone, a medical monitoring system, or an experimental apparatus. As an example, the sensor device may include one or more sensors 112 configured to acquire physiological parameters, including vital signs such as blood glucose levels, heart rate, blood pressure, respiratory rate, and body temperature. For example, one or more sensors may be an optical sensor (e.g., a PPG), a pulse pressure sensor (PP), a pressure sensor, an electrocardiogram (ECG), a bioimpedance sensor, a skin electrical response sensor, an intraocular pressure measuring / contact sensor, an accelerometer, a gyroscope, an acoustic sensor, an electromechanical motion sensor, and / or an electromagnetic sensor. Furthermore, for example, if the sensor device takes the form of an experimental apparatus, it may also measure physiological parameters or perform blood tests, such as analyzing blood glucose levels, cholesterol, and other biomarkers.
[0033] Furthermore, the sensor device 110 includes a communication module 116 configured to transmit data to and / or receive data from the electronic device 120. For example, the communication module 116 may include one or more receivers, transmitters, and / or transceivers. In some implementations, the communication module 116 may communicate with the electronic device 120 via one or more wireless links (e.g., serial link, Ethernet link, etc.) and / or wireless links (e.g., Wi-Fi link, Bluetooth link, etc.).
[0034] In some examples, the electronic device 120 can be configured to receive sensor data (e.g., physiological parameter data such as clinical parameters) acquired by the sensor device 110 and to process that sensor data. Furthermore, the electronic device 120 can be configured to present information about biomarkers and any other information to the user and / or another user (e.g., a healthcare provider).
[0035] In Figure 1, the electronic device 120 is shown as a single component. However, in practice, the electronic device 120 can be implemented in one or more computing devices (for example, each computing device may include at least one processor, such as a microprocessor or microcontroller). For example, the electronic device 120 may be a single computing device, such as a single smartphone. As another example, the electronic device 120 may include multiple computing devices connected via a network (for example, the Internet, a local area network, etc.), and the components of the electronic device 120 may be maintained and operated on some or all of these computing devices. For example, the electronic device 120 may include multiple computing devices, and the components of the electronic device 120 may be distributed across one or more of these computing devices.
[0036] Furthermore, the electronic device 120 is shown as a component separate from the sensor device 110. However, while the electronic device 120 can be a separate component from the sensor device 110, it can also include the sensor device 110, be coupled to the sensor device 110, or be adjacent to the sensor device 110 (for example, within a housing). For example, the electronic device 120 could be a wearable device that includes the sensor device 110, is coupled to the sensor device 110, or is adjacent to the sensor device 110.
[0037] As shown in Figure 1, the electronic device 120 includes a database module 122, a communication module 124, a processing module 126, and a user interface module 128. The operation module can be provided as one or more computer-executable software modules, hardware modules, or a combination thereof. For example, one or more operation modules can be implemented as a block of software code containing instructions that cause one or more processors to perform the operations described herein. Additionally or alternatively, one or more operation modules can be implemented in electronic circuits such as programmable logic circuits, field-programmable logic arrays (FPGAs), or application-specific integrated circuits (ASICs).
[0038] The communication module 124 is configured to transmit data to and / or receive data from the sensor device 110. For example, the communication module 124 may include one or more receivers, transmitters, and / or transceivers. In some implementations, the communication module 124 can communicate with the sensor device 110 via one or more wired links (e.g., serial link, Ethernet link, etc.) and / or wireless links (e.g., Wi-Fi link, Bluetooth link, etc.) (for example, via the communication module 116).
[0039] The database module 122 maintains information related to the operation of the medical data anonymization system 100.
[0040] For example, the database module 122 can store input data 122a, which may include or correspond to medical data subject to anonymization. In some implementations, the input data 122a may include at least some of the sensor data generated by the sensor device 110.
[0041] As another example, the database module 122 can store output data 122b generated by the electronic device 120. For example, the output data 122b may include standardized or anonymized data generated by the electronic device 120 based on the input data 122a.
[0042] Furthermore, the database module 122 can store processing rules 122c that specify how the data within the database module 122 is processed in order to perform the operations described herein.
[0043] For example, processing rule 122c may include one or more rules specifying how input data 122a is formatted, parsed, and processed in order to anonymize or standardize the medical data.
[0044] As another example, processing rule 122c may include one or more rules that specify the conditions under which data is presented to the user (for example, using user interface module 128) and the manner in which the data is presented.
[0045] As another example, processing rule 122c may include one or more rules that specify how to store data for future retrieval and / or processing (for example, using database module 122).
[0046] Examples of data processing techniques are described in more detail below.
[0047] The processing module 126 processes data stored in the electronic device 120, or data that is otherwise accessible to the electronic device 120. For example, the processing module 126 can be used to perform one or more of the operations described herein (for example, by executing processing rule 122c with respect to input data 112a to generate output data 122b).
[0048] The user interface module 128 is configured to present information to the user and / or receive input from the user. As an example, the user interface module 128 may include one or more display devices (e.g., display screens, touchscreens, etc.) configured to present a user interface (e.g., a graphical user interface, GUI) that allows the user to interact with the electronic device 120 and / or the sensor device 110. Exemplary interactions include browsing data, sending data from one component to another, and / or issuing commands to the electronic device 120 and / or the sensor device 110. Commands may include, for example, arbitrary user instructions to one or more of the electronic device 120 and / or the sensor device 110 to perform a particular action or task. In some implementations, the user interface module may also present information to the user audibly (e.g., using one or more speakers) and / or via haptic feedback (e.g., using one or more haptic generators such as vibration generators).
[0049] In some implementations, software applications can be used to facilitate the performance of the tasks described herein. For example, an application can be installed on the electronic device 120. Furthermore, the user can interact with the application to input data and / or commands to the electronic device 120 and to verify the data generated by the electronic device 120.
[0050] Figure 2 shows an exemplary implementation of data anonymization software or algorithms used by processor-based electronic devices (for example, electronic device 120 in the medical data anonymization system 100 in Figure 1, and computing device (which may be a server) in the system 500 in Figure 5). Specifically, the software or algorithms are used by the electronic device to [1] perform a first anonymization operation on medical data, [2] generate standardized medical data, and [3] perform a second anonymization operation on the standardized medical data.
[0051] An exemplary implementation configuration 200 shows a data store 210 and data anonymization software 260.
[0052] The data store 210 may include or correspond to a data store of an electronic device (which may be a server). For example, the data store 210 may be a database module 122 of an electronic device 120 and one or more storage devices 530 of the computing devices (which may be servers) of the system 500. The data store 210 can communicate data with the electronic device (which may be a server).
[0053] The data store 210 may contain one or more of the following: clinical research datasets 212, standardized datasets 240, and first data structures 250. The clinical research dataset 212 may contain directly identified data 220 and clinical research data 230.
[0054] Directly Identifiable Data 220 may include direct identifiers that can be used to identify entities (e.g., individuals, users, patients, subjects, etc.), clinical studies, locations, or parties associated with a clinical study. For example, for illustrative purposes, such direct identifiers may include the entity's demographics 222, identifiers 224 associated with one or more entities, one or more clinical studies, or one or more parties associated with a clinical study, location information 226 relating to a clinical study, dates 228 associated with a clinical study, and so on. Furthermore, such directly identifiable data 220 may include more or fewer categories and / or different types of data, and are not limited to the data categories shown above, also shown in Figure 2. Directly identifiable data 220 may be represented by one or more first values in (or stored in) one or more first data fields.
[0055] Such directly identifiable data 220 may be anonymized during a first anonymization operation, as described below with respect to data anonymization software 260.
[0056] Clinical research data 230 may include information about the study design 232, data collection 234, treatment 236, etc. For example, the study design 232 may include treatment groups, visit schedules, administration stages, references to specific points in time regarding patients being targeted for the product, trial, or test, and the category of the trial or test. Furthermore, for example, data collection 234 may include information about trial or test data, biomarkers, vital signs, etc., and treatment information may include the treatment name and other treatment-related information. Furthermore, for example, treatment 236 may include information about the treatment name and treatment group. Such study designs 232, data collection 234, and treatment 236 may include more or fewer categories and / or different types of data, and are not limited to the data categories described above. Clinical research data 230 may be represented by one or more second values in (or stored in) one or more second data fields.
[0057] Such clinical research data 230 may be anonymized during a second anonymization operation, as described below with respect to data anonymization software 260.
[0058] The clinical research dataset 212 can be used by an electronic device to generate an anonymized clinical research dataset based on a first anonymization operation via data anonymization software 260. Such an anonymized clinical research dataset can also be stored in the data store 210 and used by the electronic device to generate a standardized dataset 240.
[0059] Furthermore, the standardized dataset 240 can be used by an electronic device, via data anonymization software 260, to generate a modified standardized dataset based on a second anonymization operation. Such a modified standardized dataset can then be used to generate the first data structure 250.
[0060] In some implementations, the clinical research dataset 212, the standardized dataset 240, and the first data structure 250 can be stored in different data stores. For example, the clinical research dataset 212 can be stored in a different server (e.g., a computing device in system 500, which may also be a server), and the standardized dataset 240 and / or the first data structure 250 can be stored in the data store of an electronic device (for example, assuming in this example that the electronic device does not take the form of a server), and vice versa.
[0061] Furthermore, at least some of the data anonymization and / or standardization can be implemented as respective software programs executable by electronic devices. The software programs, stored in memory (such as the database module 122 in Figure 1, memory 520, and storage device 530 in Figure 5), and executed by a processor, may include machine-readable instructions that cause processor-based electronic devices to execute the instructions of the software program. As illustrated, the data anonymization software 260 may include a first anonymization tool 262, a data standardization tool 264, a second anonymization tool 266, and / or a first data structure generation tool 268. In some implementations, the data anonymization software 260 may include more or fewer tools. In some implementations, some of the tools may be combined, some of the tools may be divided into more tools, or a combination thereof may occur. In some implementations, the data anonymization software 260 may run on a server (for example, a computing device in System 500, which may also be a server), or it may run on both an electronic device and a server.
[0062] In some implementations, the data standardization tool 264 and / or the first data structure generation tool 268 may take the form of software separate from the data anonymization software 260 and run on a server, while other tools (e.g., the first anonymization tool 262, the second anonymization tool 266) may take the form of the data anonymization software 260 and run on an electronic device communicating with the server, and vice versa. Further variations are possible regarding the first anonymization tool 262, the data standardization tool 264, the second anonymization tool 266, and the first data structure generation tool 268 being separate software and running on an electronic device, a server, or a combination thereof.
[0063] The first anonymization tool 262 can be used to perform a first anonymization operation. For example, the first anonymization tool 262 may include a first computer-executable anonymization function for each of the clinical research datasets 212 in order to generate a modified clinical research dataset. The first computer-executable anonymization function can mask one or more first values (or directly identifiable data 220) within the clinical research dataset 212. For example, masking data such as one or more values (e.g., a first value, a second value) may include or correspond to modifying data by selectively obfuscating, deleting, hiding, or otherwise altering portions of the data.
[0064] For example, performing a first computer-executable anonymization function may include modifying one or more first values in a clinical research dataset 212 according to a first set of rules in order to generate an anonymized representation of the directly identifiable data 220.
[0065] In some implementations, the first set of rules may include: (i) determining that one or more first identifiers (e.g., identifiers for identifier 224) are associated with one or more entities; (ii) determining that one or more second identifiers (e.g., identifiers for identifier 224) are associated with one or more parties associated with one or more clinical studies, wherein one or more parties provide at least one of the drugs, products, or treatments used in one or more clinical studies; (iii) removing one or more first identifiers from the clinical study dataset; and (iv) replacing one or more second identifiers in the clinical study dataset with one or more alphanumeric characters or symbols.
[0066] In some implementations, the first set of rules may include (i) determining that date information (e.g., date 228) includes a calendar day, and (ii) replacing calendar days in a clinical study dataset (e.g., clinical study dataset 212) with relative dates, where each relative date represents a time offset to a predetermined event. For example, the predetermined event could be the start date of participation in one or more clinical studies by each of the entities in one or more clinical studies.
[0067] For example, the first set of rules could include the following anonymization techniques based on specific attributes, as outlined in Table 1 below.
[0068] [Table 1] JPEG2026082708000003.jpg158162
[0069] The data standardization tool 264 can be used to generate a standardized dataset 240 based on a modified clinical research dataset (for example, a dataset modified using the first anonymization tool 262).
[0070] For example, generating a standardized dataset 240 may include concatenating a modified clinical research dataset to the standardized dataset 240 and at least one of the following: (i) remapping at least one of the first data fields of the standardized dataset 240 to its respective standardized first data field; (ii) remapping at least one of the second fields of the standardized dataset 240 to its respective standardized second data field; (iii) remapping at least one of the first values of the standardized dataset 240 to its respective standardized first value; or (iv) remapping at least one of the second values of the standardized dataset 240 to its respective standardized second value.
[0071] A second anonymization tool 266 can be used to perform a second anonymization operation. For example, the second anonymization tool 266 may include a second computer-executable anonymization function with respect to the standardized dataset 240 in order to generate a modified standardized dataset. The second computer-executable anonymization function can mask one or more second values in the standardized dataset 240 (for example, values performed from clinical research data 230 to the standardized dataset 240, and standardized second values remapped from the second values).
[0072] Implementing a second computer-executable anonymization function may include modifying at least a portion of the standardized dataset 240 according to a second set of rules in order to generate an anonymized representation of clinical research information. In some implementations, the second set of rules may include modifying naming conventions in the standardized dataset 240 related to study design, data collection, treatment information, or other treatment-related data. For example, the second set of rules may include modifying naming conventions in the standardized dataset 240 for at least one of the following: (i) visit information, (ii) administration stage, (iii) trial or test, (iv) treatment name, (v) treatment group, (vi) reference to a specific point in time regarding the patient being treated for the product, trial, or test, or (vii) category of trial or test.
[0073] The naming conventions for test or test categories may be broader and more generalized than the naming conventions for individual tests or tests. For example, while test naming conventions are specific to individual tests or evaluations, the naming conventions for test categories may be broader and encompass a group of related tests or evaluations.
[0074] Table 2 below shows an example of modifying the naming convention for visit information. In this example, the second value related to the naming convention for visit schedules is modified.
[0075] [Table 2]
[0076] In the example above, the original visit naming convention for the visit schedule is more specific, allowing us to identify three different studies in Table 2. However, after modifying the naming convention for the visit information, it becomes more difficult, and practically impossible, to identify the three different studies based on the visit value.
[0077] Furthermore, Table 3 below shows an example of modifying the naming convention for administration stages (for example, the time-series information shown in Table 3). In this example, the second value related to the naming convention for administration stages is modified.
[0078] [Table 3]
[0079] Furthermore, Table 4 below shows an example of modifying the naming convention for a test or inspection. In this example, the second value related to the naming convention for the test or inspection is modified.
[0080] [Table 4]
[0081] In some cases, a standardized dataset may include groups associated with entities that share one or more common criteria or characteristics. In some implementations, performing a second computer-executable anonymization function may include performing data aggregation by (i) combining one or more second values from two or more of the groups (for example, second values performed from clinical research data 230 to standardized dataset 240, or standardized second values remapped from second values), and (ii) generating one or more larger, less specific groups based on one or more additional common criteria or characteristics shared by one or more larger, less specific groups. In some implementations, performing a second computer-executable anonymization function may include performing data cohorting by (i) generating additional groups beyond the number of groups, and (ii) regrouping entities and one or more second values in the standardized dataset into additional groups.
[0082] Once the first and second anonymization operations are complete, the first data structure generation tool 268 can generate the first data structure 250. For example, the first data structure 250 can represent a modified, standardized dataset. Such a data structure 250 can be stored in the data store 210, transmitted to a computerized health analysis platform, or output to a user interface for display.
[0083] Exemplary process Figure 3 is a flowchart of an exemplary process 300 for anonymizing medical data. In particular, the exemplary process 300 [1] performs a first anonymization on medical data, [2] generates standardized medical data, and [3] performs a second anonymization on the standardized medical data. Process 300 can be implemented by a processor-based system such as the medical data anonymization system 100 and system 500 as described herein, and in combination with the exemplary implementation form 200.
[0084] In 302, clinical research datasets relating to one or more clinical studies are accessed or read from a hardware storage device. For example, an electronic device (e.g., electronic device 120 of the medical data anonymization system 100 or a computing device of system 500) can be used to access or read the clinical research dataset from the hardware storage device. For example, the hardware storage device may correspond to the database module 122 of electronic device 120, or to a storage device of a computing device (e.g., one or more storage devices 530 of the computing devices in Figure 5, including a server). In some implementations, the electronic device may correspond to a server, and the hardware storage device may correspond to the server's data store.
[0085] Furthermore, each clinical research dataset (for example, clinical research dataset 212) may include, for example, one or more first data fields that store one or more first values representing identification information associated with one or more clinical research and entities participating in one or more clinical research (for example, information represented by direct identification data 220), and (ii) one or more second data fields that store one or more second values representing clinical research information collected during one or more clinical research (for example, information represented by clinical research data 230).
[0086] In some implementations, the identifying information includes at least one of the following: (i) demographic information relating to an entity; (ii) an identifier relating to an entity, one or more clinical studies, or one or more parties relating to one or more clinical studies; (iii) location information relating to one or more clinical studies; (iv) treatment information relating to at least one of the drugs, products, or treatments used in one or more clinical studies; or (v) date information relating to one or more clinical studies.
[0087] In some implementations, clinical research information includes information about the study design, data collection, and treatment associated with one or more clinical studies, as described above in relation to the description of the exemplary implementation 200 in Figure 2. In some implementations, at least some of the clinical research information may be represented by a naming convention that includes at least one of the following: (i) visit information, (ii) administration stage, (iii) study or test, (iv) treatment name, (v) treatment group associated with multiple entities, (vi) a reference to a specific point in time regarding patients targeted by the product, study, or test, or (vii) a category of study or test.
[0088] In 304, a first computer-executable anonymization function is performed on each of the clinical research datasets in order to generate a modified clinical research dataset. For example, an electronic device can be used to perform the first computer-executable anonymization function on the clinical datasets.
[0089] For example, performing a first computer-executable anonymization function can mask one or more first values in a clinical research dataset. For example, performing a first computer-executable anonymization function may include modifying one or more first values in a clinical research dataset according to a first set of rules (for example, the set of first rules described above in relation to the description of the exemplary implementation form 200 in Figure 2) in order to generate an anonymized representation of the identifying information.
[0090] In some implementations, the first set of rules may include replacing at least one of the following in a clinical research dataset—demographic information, identifiers, location information, or treatment information—with one or more alphanumeric characters or symbols.
[0091] In some implementations, the first set of rules may include: (i) determining that one or more first identifiers are associated with one or more entities; (ii) determining that one or more second identifiers are associated with one or more parties associated with one or more clinical studies, wherein one or more parties provide at least one of the drugs, products, or treatments used in one or more clinical studies; (iii) removing one or more first identifiers from the clinical study dataset; and (iv) replacing one or more second identifiers in the clinical study dataset with one or more alphanumeric characters or symbols.
[0092] In some implementations, the first set of rules may include (i) determining that date information includes calendar days, and (ii) replacing calendar days in the clinical study dataset with relative dates, each of which represents a time offset to a predetermined event. For example, the predetermined events could be the start date of participation in one or more clinical studies by each of the entities in one or more clinical studies.
[0093] In 306, a standardized dataset (e.g., standardized dataset 240) is generated based on the modified clinical research dataset. For example, an electronic device can generate a standardized dataset based on a modified clinical research dataset (e.g., one modified based on the execution of a computer-executable anonymization function).
[0094] For example, generating a standardized dataset may involve concatenating a modified clinical research dataset into a standardized dataset and include at least one of the following: (i) remapping at least one of the first data fields of the standardized dataset to each standardized first data field; (ii) remapping at least one of the second fields of the standardized dataset to each standardized second data field; (iii) remapping at least one of the first values of the standardized dataset to each standardized first value; or (iv) remapping at least one of the second values of the standardized dataset to each standardized second value.
[0095] In 308, a second anonymization function is performed on the standardized dataset to generate a modified, standardized dataset. For example, an electronic device can be used to perform a second computer-executable anonymization function on the standardized dataset.
[0096] For example, performing a second computer-executable anonymization function can mask at least one of the research design, data collection, or treatment information of one or more clinical studies. For example, performing a second computer-executable anonymization function may include modifying at least a portion of a standardized dataset according to a second set of rules (e.g., the second set of rules described above in relation to the description of the exemplary implementation form 200 in Figure 2) in order to generate an anonymized representation of the clinical study information.
[0097] In some implementations, the second set of rules may include modifying naming conventions related to study design, data collection, treatment information, or other treatment-related data in a standardized dataset. For example, the second set of rules may include modifying naming conventions in a standardized dataset for at least one of the following: (i) visit information, (ii) administration stage, (iii) trial or test, (iv) treatment name, (v) treatment group, (vi) reference to a specific point in time regarding the patient being treated for the product, trial, or test, or (vii) category of trial or test.
[0098] In some cases, a standardized dataset may contain groups associated with entities that share one or more common criteria or characteristics. In some implementations, performing a second computer-executable anonymization function may include performing data aggregation by (i) combining one or more second values from two or more of the groups (e.g., second values performed from clinical research datasets to a standardized dataset, or standardized second values remapped from second values), and (ii) generating one or more larger, less specific groups based on one or more additional common criteria or characteristics shared by one or more larger, less specific groups. In some implementations, performing a second computer-executable anonymization function may include performing data cohorting by (i) generating additional groups beyond the number of existing groups, and (ii) regrouping entities and one or more second values in the standardized dataset into additional groups.
[0099] In step 310, a first data structure is generated that represents the modified standardized dataset. For example, an electronic device can format the modified standardized dataset data into a standardized data format suitable for storage in a hardware storage device.
[0100] In 312, the first data structure is stored in a hardware storage device. For example, the data structure may be stored in a data store (e.g., data store 210, database module 122 of electronic device 120, one or more storage devices 530, etc.).
[0101] In some implementations, instead of storing the first data structure in a hardware storage device, or in addition to doing so, the first data structure may be output to a user interface (for example, using user interface module 128). For example, if the first data structure is output for user display without storing the data structure, in step 312, the first data structure may be generated in another format suitable for display. For example, if the first data structure is stored and then output for user display, the first data structure can be converted from its storage format to an appropriate display format before output.
[0102] In some implementations, the first data structure can be provided to or transmitted to a computerized health analysis platform.
[0103] Figure 4 is a flowchart of an exemplary process 400 for anonymizing medical data. Specifically, the exemplary process 400 [1] performs a first anonymization on medical data, [2] generates standardized medical data, and [3] performs a second anonymization on the standardized medical data. Process 400 can be implemented by a processor-based system such as the medical data anonymization system 100 and system 500 as described herein, and in combination with exemplary implementation form 200 and exemplary process 300.
[0104] In step 402, the first anonymization operation is performed on each of the clinical research datasets (e.g., clinical research dataset 212). A processor-based device (e.g., electronic device 120 of the medical data anonymization system 100, or a computing device of system 500) can be used to access the clinical research datasets from a data store (e.g., data store 210, database module 122 of electronic device 120, one or more storage devices 530, etc.) and perform the first anonymization operation on each of the clinical datasets. For example, the technique for the first anonymization operation may be similar to the technique used in step 304 of Figure 3.
[0105] For example, each clinical research dataset (e.g., clinical research dataset 212) may include (i) one or more first data fields that store one or more first values representing identification information (e.g., information represented by direct identification data 220) associated with one or more clinical research and multiple entities participating in one or more clinical research, and (ii) one or more second data fields that store one or more second values representing clinical research information (e.g., information represented by clinical research data 230) collected during one or more clinical research.
[0106] For example, performing a first anonymization operation on each clinical dataset involves performing a first computer-executable anonymization function using a first set of rules (e.g., the first set of rules described above in relation to the description of the exemplary implementation form 200 in Figure 2) to generate an anonymized representation of the identifying information contained in the clinical research dataset (e.g., the information represented by the direct identifying data 220).
[0107] In step 404, the standardized dataset is generated using the clinical research dataset that has undergone a first anonymization operation. Processor-based devices can generate the standardized dataset based on the clinical research dataset. For example, the technique for step 404 can be similar to the technique used in step 306 in Figure 3. For example, generating the standardized dataset may include concatenating the modified clinical research dataset into the standardized dataset and at least one of the following: (i) remapping at least one of the first data fields of the standardized dataset to the respective standardized first data field; (ii) remapping at least one of the second fields of the standardized dataset to the respective standardized second data field; (iii) remapping at least one of the first values of the standardized dataset to the respective standardized first value; or (iv) remapping at least one of the second values of the standardized dataset to the respective standardized second value.
[0108] In step 406, a second anonymization operation is performed on the standardized dataset to generate a modified standardized dataset. A processor-based device can perform a second anonymization operation on the standardized dataset to generate a modified standardized dataset. For example, the technique for step 406 can be the same as the technique used in step 306 in Figure 3.
[0109] For example, performing a second anonymization operation on a standardized dataset involves performing a second computer-executable anonymization function using a second set of rules (e.g., the second set of rules described above in relation to the exemplary implementation form 200 in Figure 2). For example, performing a second computer-executable anonymization function involves at least one of the following: (i) modifying naming conventions for clinical research information; (ii) performing data aggregation; or (iii) performing data cohorting to generate an anonymized representation of clinical research information and / or other information related to a standardized dataset, as described above with respect to exemplary implementation form 200 and process 300.
[0110] In 408, after the second anonymization operation is complete, the anonymized and standardized dataset is sent to a computerized health analysis platform or output to a user interface (for example, using user interface module 128) for display. In some implementations, the anonymized and standardized dataset can be stored in a data store.
[0111] Exemplary computer system Figure 5 shows an exemplary computing system in an implementation of the present disclosure. System 500 can be used for any of the operations described with respect to the various implementations discussed herein. System 500 may be included in, used by, communicate with, or correspond to an electronic device 120. Furthermore, System 500 may include, be used by, or communicate with a sensor device 110. System 500 may include one or more processors 510, memory 520, one or more storage devices 530, and one or more input / output (I / O) devices 560 controllable via one or more I / O interfaces 540. Various components 510, 520, 530, 540, or 560 may be interconnected through at least one system bus 550, enabling data transfer between various modules and components of System 500.
[0112] The processor 510 may be configured to process instructions for execution within the system 500. The processor 510 may include a single-threaded processor, a multi-threaded processor, or both. The processor 510 may be configured to process instructions stored in memory 520 or a storage device 530. The processor 510 may include a hardware-based processor, each containing one or more cores. The processor 510 may include a general-purpose processor, a dedicated processor, or both.
[0113] Memory 520 may store information within the system 500. In some implementations, memory 520 includes one or more computer-readable media. Memory 520 may include any number of volatile memory units, any number of non-volatile memory units, or both volatile and non-volatile memory units. Memory 520 may include read-only memory, random-access memory, or both. In some examples, memory 520 may be used as active memory or physical memory by one or more executable software modules.
[0114] The storage device 530 may be configured to provide the system 500 with (for example, persistent) large-capacity storage. In some implementations, the storage device 530 may include one or more computer-readable media. For example, the storage device 530 may include a floppy disk device, a hard disk device, an optical disk device, or a tape device. The storage device 530 may include read-only memory, random-access memory, or both. The storage device 530 may include one or more internal hard drives, external hard drives, or removable drives.
[0115] One or both of the memory 520 or the storage device 530 may include one or more computer-readable storage media (CRSMs). A CRSM may include one or more electronic storage media, magnetic storage media, optical storage media, magneto-optical storage media, quantum storage media, mechanical computer storage media, etc. A CRSM may provide storage for computer-readable instructions describing data structures, processes, applications, programs, other modules, or other data for the operation of the system 500. In some implementations, a CRSM may include a datastore providing non-temporary storage of computer-readable instructions or other information. A CRSM may be integrated into the system 500 or external to the system 500. A CRSM may include read-only memory, random-access memory, or both. One or more CRSMs suitable for tangibly embodying computer program instructions and data may include, but are not limited to, any type of non-volatile memory, including, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. In some examples, the processor 510 and memory 520 may be complemented by or incorporated into one or more application-specific integrated circuits (ASICs).
[0116] System 500 may include one or more I / O devices 560. I / O devices 560 may include one or more input devices such as a keyboard, mouse, pen, game controller, touch input device, audio input device (e.g., microphone), gesture input device, haptic input device, image or video capture device (e.g., camera), or other devices. In some examples, I / O devices 560 may also include one or more output devices such as a display, LED, audio output device (e.g., speaker), printer, haptic output device. I / O devices 560 may be physically integrated into one or more computing devices of System 500, or they may be external to one or more computing devices of System 500.
[0117] System 500 may include one or more I / O interfaces 540 that allow components or modules of System 500 to control, interface with, or otherwise communicate with I / O device 560. The I / O interfaces 540 may enable the transfer of information within or outside System 500, or between components of System 500, via serial, parallel, or other types of communication. For example, an I / O interface 540 may conform to a version of the RS-232 standard for serial ports and a version of the IEEE 1284 standard for parallel ports. As another example, an I / O interface 540 may be configured to provide connectivity via Universal Serial Bus (USB) or Ethernet. In some examples, an I / O interface 540 may be configured to provide serial connectivity conforming to a version of the IEEE 1394 standard.
[0118] The I / O interface 540 may also include one or more network interfaces that enable communication between computing devices within System 500, or between System 500 and other computing systems connected to a network. The network interfaces may include one or more network interface controllers (NICs) or other types of transceiver devices configured to send and receive communications over one or more networks using any network protocol.
[0119] The computing devices of System 500 may communicate with each other and with other computing devices using one or more networks. Such networks may include public networks such as the Internet, private networks such as an organization's or individual's intranet, or any combination of private and public networks. Networks may include, but are not limited to, any type of wired or wireless network, including local area networks (LANs), wide area networks (WANs), wireless WANs (WWANs), wireless LANs (WLANs), and mobile communication networks (e.g., 3G, 4G, edge, etc.). In some implementations, communication between computing devices may be encrypted or otherwise protected. For example, communication may use one or more public or private encryption keys, cryptographics, digital certificates, or other credentials supported by a security protocol such as any version of the Secure Sockets Layer (SSL) or Transport Layer Security (TLS) protocol.
[0120] System 500 may include any number of computing devices of any kind. Computing devices may include, but are not limited to, personal computers, smartphones, tablet computers, wearable computers, embedded computers, mobile gaming devices, e-readers, automotive computers, desktop computers, laptop computers, notebook computers, game consoles, home entertainment devices, network computers, server computers, mainframe computers, distributed computing devices (e.g., cloud computing devices), microcomputers, systems on a chip (SoC), systems in a package (SiP), etc. While the examples herein may describe computing devices as physical devices, the form of implementation is not limited thereto. In some examples, computing devices may include one or more virtual computing environments, hypervisors, emulations, or virtual machines running on one or more physical computing devices. In some examples, two or more computing devices may include a cluster, cloud, farm, or other group of multiple devices that coordinate their operation to provide load balancing, failover support, parallel processing capabilities, shared storage resources, shared networking capabilities, or other aspects.
[0121] In this specification, the term “configured” is used in relation to systems and computer program components. When one or more computer systems are configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof is installed on the system that causes the system to perform that operation or action during operation. When one or more computer programs are configured to perform a particular operation or action, it means that one or more programs, when executed by a data processing device, contain instructions that cause the device to perform that operation or action.
[0122] The subject matter and functional embodiments described herein can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware, or one or more combinations thereof, including structures disclosed herein and their structural equivalents. Embodiments of the subject matter described herein can be implemented as one or more modules of computer programs, i.e., computer program instructions encoded on a tangible non-temporary storage medium for execution by or control of the operation of a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage board, a random or serial access memory device, or one or more combinations thereof. Alternatively or additionally, the program instructions can also be encoded into artificially generated propagating signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiving device for execution by a data processing device.
[0123] The term "data processing device" refers to data processing hardware and encompasses all kinds of devices, machines, and equipment for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. A device may be, or further include, a special-purpose logic circuit, such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). In addition to hardware, a device may optionally include code that creates an environment for executing computer programs, such as processor firmware, a protocol stack, a database management system, an operating system, or code comprising one or more of these.
[0124] Computer programs, also called, or described as, programs, software, software applications, apps, modules, software modules, scripts, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program may be stored in part of a file that holds other programs or data, for example, in one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple interconnected files, for example, in a file that stores one or more modules, subprograms, or parts of code. A computer program can be deployed to run on one computer, or on multiple computers located in one site or distributed across multiple sites and interconnected by a data communication network.
[0125] In this specification, the term “database” is used broadly to refer to any collection of data. The data does not need to be structured in any particular way, or even structured at all, and can be stored on one or more storage devices in different locations. Therefore, for example, an index database may contain multiple collections of data, each of which may be organized and accessed in a different way.
[0126] The processes and logic flows described herein can be executed by one or more programmable computers that run one or more computer programs to perform functions by acting on input data and producing outputs. The processes and logic flows can also be executed by dedicated logic circuits, such as FPGAs or ASICs, or by a combination of dedicated logic circuits and one or more programmed computers.
[0127] A computer suitable for running computer programs can be based on a general-purpose or dedicated microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory, random-access memory, or both. The basic elements of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be complemented by or integrated into dedicated logic circuits. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operationally coupled to receive data from, transfer data to, or both of those storage devices. However, a computer does not necessarily need to have such devices. Furthermore, a computer can also integrate portable storage devices, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or, for example, Universal Serial Bus (USB) flash drives, into another device, to name just a few examples.
[0128] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0129] The term "memory subsystem" can include one or more memories, each of which may be a computer-readable medium. A memory subsystem may include memory hardware units (e.g., hard drives or disks) that store data or instructions in software format. Alternatively, or additionally, a memory subsystem may include data or instructions hardwired to processing circuits.
[0130] To provide user interaction, embodiments of the subject matter described herein can be implemented on a computer equipped with a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, on which the user can provide input to the computer. Other types of devices can also be used to provide user interaction, for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, voice, or tactile input. Furthermore, the computer can interact with the user by sending and receiving documents to and from devices used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from a web browser. The computer can also interact with the user by sending text messages or other forms of messages to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user in return.
[0131] Embodiments of the subject matter described herein can be implemented in a computing system including, for example, a backend component as a data server, or a computing system including, for example, a middleware component as an application server, or a computing system including, for example, a frontend component of a client computer having, for example, a graphical user interface, a web browser, or an application that allows a user to interact with an implementation of the subject matter described herein, or in any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0132] A computing system can include clients and servers. Clients and servers are typically remote to each other and usually interact through a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with each other. In some embodiments, the server sends data, such as an HTML page, to a user device acting as a client, for the purpose of displaying data to the user or receiving user input, for example, a device interacting with the user. Data generated on the user device, such as the results of user interaction, can be received from the device to the server.
[0133] This specification includes many details of specific implementations, but these should not be construed as limiting the scope of any invention or claim, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable subcombination in multiple embodiments. Furthermore, even if features are described above as functioning in a particular combination and were initially claimed as such, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.
[0134] Similarly, while the drawings show operations in a specific order and the claims describe them in a specific order, this should not be understood as requiring that such operations be performed in the specific illustrated order or sequential order, or that all illustrated operations be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems should be understood as generally being able to be integrated into a single software product or packaged into multiple software products.
[0135] Specific embodiments of the subject matter have been described. Other embodiments are also included in the following claims. For example, the actions described in the claims can still achieve the desired results even if they are performed in a different order. As an example, the process shown in the accompanying drawings does not necessarily require the specific order or sequential order shown to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A system for improving data security in a computerized health analysis platform, Hardware storage devices and At least one processor, A memory subsystem that is communicatively coupled to at least one of the processors and The memory subsystem is provided, and when the memory subsystem is executed by the at least one processor, the at least one processor is configured to Reading multiple clinical research datasets relating to one or more clinical studies from the hardware storage device, wherein each of the clinical research datasets is One or more first data fields that store one or more first values representing identification information associated with one or more clinical studies and multiple entities participating in one or more clinical studies, One or more second data fields that store one or more second values representing clinical research information collected during the one or more clinical studies, To be equipped, to read, To generate a plurality of modified clinical research datasets, the method of performing a first computer-executable anonymization function on each of the plurality of clinical research datasets, wherein the method of performing the first computer-executable anonymization function involves masking one or more of the first values in the plurality of clinical research datasets, and the method of performing the first computer-executable anonymization function involves modifying the one or more of the first values in the plurality of clinical research datasets according to a first set of rules in order to generate an anonymized representation of the identification information. The process involves generating a standardized dataset based on the aforementioned multiple modified clinical research datasets, wherein the generation of the standardized dataset is The modified clinical research dataset is linked to the standardized dataset, Remapping at least one of the first data fields of the standardized dataset to each of the standardized first data fields, Remapping at least one of the second fields of the standardized dataset to each of the standardized second data fields, Remapping at least one of the first values of the standardized dataset to each of the standardized first values, or Remapping at least one of the second values of the standardized dataset to each of the standardized second values, At least one of the following, To have, to generate, To generate a modified standardized dataset, the execution of a second computer-executable anonymization function on the standardized dataset, wherein the execution of the second computer-executable anonymization function masks at least one of the research design, data collection, or treatment information of one or more clinical studies, and the execution of the second computer-executable anonymization function modifies at least a portion of the standardized dataset according to a second set of rules in order to generate an anonymized representation of the clinical study information. To generate a first data structure representing the modified standardized dataset, The first data structure is stored using the hardware storage device, A system that stores instructions for executing actions that include the following.
2. The system according to claim 1, wherein the operation comprises providing the first data structure to a computerized health analysis platform.
3. The aforementioned identification information, Population information relating to the aforementioned entity, The entity, the one or more clinical studies, or an identifier associated with at least one of the one or more parties associated with the one or more clinical studies, Location information relating to one or more clinical studies, Treatment information relating to at least one of the drugs, products, or treatments used in the aforementioned one or more clinical studies, At least one of the date information associated with one or more clinical studies, The system according to claim 1, comprising:
4. The first set of rules mentioned above is The system according to claim 3, comprising replacing at least one of the demographic information, identifiers, location information, or treatment information in the plurality of clinical study datasets with one or more alphanumeric characters or symbols.
5. The first set of rules mentioned above is Determining that one or more first identifiers are associated with one or more of the entities, Determining that one or more second identifiers are associated with the one or more parties associated with the one or more clinical studies, and that the one or more parties provide at least one of the drugs, products, or treatments used in the one or more clinical studies, Deleting one or more first identifiers from the aforementioned multiple clinical research datasets, Replacing one or more second identifiers in the aforementioned multiple clinical research datasets with one or more alphanumeric characters or symbols, The system according to claim 3, comprising:
6. The first set of rules mentioned above is The aforementioned date information is determined to include a calendar day, Replacing the calendar dates in the aforementioned multiple clinical study datasets with relative dates, wherein each of the relative dates represents a time offset for a predetermined event; The system according to claim 3, comprising:
7. The system according to claim 6, wherein the predetermined event corresponds to the start date of participation in one of the entities in the one or more clinical studies.
8. The system according to claim 1, wherein at least some of the clinical research information is represented by a naming convention of at least one of the following: (i) visit information, (ii) administration stage, (iii) trial or test, (iv) treatment name, (v) treatment group associated with the plurality of entities, (vi) product, reference to a specific point in time relating to the patient being studied in the trial or test, or (vii) category of the trial or test.
9. The second set of rules mentioned above is The system according to claim 8, comprising modifying the naming convention for at least one of the following in the standardized dataset: (i) the visit information, (ii) the administration stage, (iii) the test or examination, (iv) the treatment name, (v) the treatment group, (vi) the product, the reference to the patient being subjected to the test or examination at a specific point in time, or (vii) the category of the test or examination.
10. The standardized dataset comprises (i) a plurality of entities that share one or more common criteria or characteristics, and (ii) a plurality of groups associated with one or more second values in the standardized dataset. Executing the aforementioned computer-executable second anonymization function is The system according to claim 1, comprising (i) combining one or more second values from two or more of the plurality of groups, and (ii) generating one or more larger, less specific groups based on one or more additional common criteria or characteristics shared by one or more larger, less specific groups.
11. The standardized dataset comprises (i) a plurality of entities that share one or more common criteria or characteristics, and (ii) a plurality of groups associated with one or more second values in the standardized dataset. Executing the aforementioned computer-executable second anonymization function is The system according to claim 1, comprising (i) generating a number of additional groups exceeding the number of the aforementioned groups, and (ii) performing data cohorting by regrouping the aforementioned entities and the one or more second values in the standardized dataset into the aforementioned additional groups.
12. Reading multiple clinical research datasets relating to one or more clinical studies from a hardware storage device using an electronic device, wherein each of the clinical research datasets is One or more first data fields that store one or more first values representing identification information associated with one or more clinical studies and multiple entities participating in one or more clinical studies, One or more second data fields that store one or more second values representing clinical research information collected during the one or more clinical studies, To be equipped with, To generate a plurality of modified clinical research datasets, the electronic device performs a first computer-executable anonymization function on each of the plurality of clinical research datasets, wherein performing the first computer-executable anonymization function involves masking one or more of the first values in the plurality of clinical research datasets, and performing the first computer-executable anonymization function involves modifying the one or more of the first values in the plurality of clinical research datasets according to a first set of rules in order to generate an anonymized representation of the identification information. The electronic device generates a standardized dataset based on the plurality of modified clinical study datasets, and the generation of the standardized dataset is The modified clinical research dataset is linked to the standardized dataset, Remapping at least one of the first data fields of the standardized dataset to each of the standardized first data fields, Remapping at least one of the second fields of the standardized dataset to each of the standardized second data fields, Remapping at least one of the first values of the standardized dataset to each of the standardized first values, or Remapping at least one of the second values of the standardized dataset to each of the standardized second values, To do at least one of the following, To generate a modified standardized dataset, the electronic device performs a second computer-executable anonymization function on the standardized dataset, wherein the performance of the second computer-executable anonymization function masks at least one of the research design, data collection, or treatment information of one or more clinical studies, and the performance of the second computer-executable anonymization function modifies at least a portion of the standardized dataset according to a second set of rules to generate an anonymized representation of the clinical study information. The electronic device generates a first data structure representing the modified standardized dataset, The electronic device stores the first data structure in the hardware storage device, A method that includes [a certain feature].
13. The method according to claim 12, further comprising providing the first data structure to a computerized health analysis platform using the electronic device.
14. The aforementioned identification information, Population information relating to the aforementioned entity, The entity, the one or more clinical studies, or an identifier associated with at least one of the one or more parties associated with the one or more clinical studies, Location information relating to one or more clinical studies, Treatment information relating to at least one of the drugs, products, or treatments used in the aforementioned one or more clinical studies, At least one of the date information associated with one or more clinical studies, The method according to claim 12, comprising:
15. The first set of rules mentioned above is The method according to claim 14, further comprising replacing at least one of the demographic information, identifiers, location information, or treatment information in the plurality of clinical study datasets with one or more alphanumeric characters or symbols.
16. The first set of rules mentioned above is Determining that one or more first identifiers are associated with one or more of the entities, Determining that one or more second identifiers are associated with the one or more parties associated with the one or more clinical studies, and that the one or more parties provide at least one of the drugs, products, or treatments used in the one or more clinical studies, Deleting one or more first identifiers from the aforementioned multiple clinical research datasets, Replacing one or more second identifiers in the aforementioned multiple clinical research datasets with one or more alphanumeric characters or symbols, The method according to claim 14, comprising:
17. The first set of rules mentioned above is The aforementioned date information is determined to include a calendar day, Replacing the calendar dates in the aforementioned multiple clinical study datasets with relative dates, wherein each of the relative dates represents a time offset for a predetermined event; The method according to claim 14, comprising:
18. The method according to claim 17, wherein the predetermined event corresponds to the start date of participation in one of the entities in the one or more clinical studies.
19. The method according to claim 12, wherein at least some of the clinical research information is represented by a naming convention of at least one of the following: (i) visit information, (ii) administration stage, (iii) trial or test, (iv) treatment name, (v) treatment group associated with the plurality of entities, (vi) product, reference to a specific point in time relating to a patient subject to the trial or test, or (vii) category of the trial or test.
20. The second set of rules mentioned above is The method of claim 19, comprising modifying the naming convention for at least one of the categories of the standardized dataset, which includes (i) the visit information, (ii) the administration stage, (iii) the test or examination, (iv) the treatment name, (v) the treatment group, (vi) the product, the reference to the patient being subjected to the test or examination, or (vii) the naming convention for the test or examination.
21. The standardized dataset comprises (i) a plurality of entities that share one or more common criteria or characteristics, and (ii) a plurality of groups associated with one or more second values in the standardized dataset. Executing the aforementioned computer-executable second anonymization function is The method of claim 12, comprising (i) combining one or more second values from two or more of the plurality of groups, and (ii) generating one or more larger, less specific groups based on one or more additional common criteria or characteristics shared by one or more larger, less specific groups.
22. The standardized dataset comprises (i) a plurality of entities that share one or more common criteria or characteristics, and (ii) a plurality of groups associated with one or more second values in the standardized dataset. Executing the aforementioned computer-executable second anonymization function is The method of claim 12, comprising (i) generating a number of additional groups exceeding the number of the aforementioned groups, and (ii) performing data cohorting by regrouping the aforementioned entities and the one or more second values in the standardized dataset into the aforementioned additional groups.
23. One or more non-temporary computer-readable media that, when executed by at least one processor, store instructions causing the at least one processor to perform the method according to any one of claims 12 to 22.