Displaying and reconciling differences in data stored in heterogeneous data storage devices

The system automates data reconciliation across heterogeneous databases, reducing operational costs and errors by identifying and visually highlighting discrepancies, thus improving efficiency in pharmacovigilance operations.

JP7824329B2Active Publication Date: 2026-03-04BRISTOL MYERS SQUIBB CO
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2026-03-04

Smart Images

  • Figure 0007824329000008
    Figure 0007824329000008
  • Figure 0007824329000009
    Figure 0007824329000009
  • Figure 0007824329000010
    Figure 0007824329000010
Patent Text Reader

Abstract

Provided herein are embodiments of systems, apparatus, devices, methods, and / or computer program products for generating an output indicative of differences in data stored in heterogeneous data stores and / or for reconciling data stored in heterogeneous data stores, and / or combinations and subcombinations thereof. In an embodiment, a server loads a first subset of a first data set corresponding to one or more first columns and a second subset of a second data set corresponding to one or more second columns into a data repository. The server identifies one or more differences between the first subset of data and the second subset of data in the data repository and causes the one or more differences to be displayed. The server may generate an output including the first and second data sets and a visual indicator indicative of each of the one or more differences, and causes the output to be displayed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates to displaying and reconciling differences in data stored in heterogeneous data storage devices. [Background technology]

[0002] Multiple entities, such as companies, government agencies, educational institutions, etc., store common data in two or more data stores. However, differences or omissions in the data in one or more of the data stores may exist. As a result, the multiple entities may need to reconcile the data to identify differences in the data between one or more of the data stores. Summary of the Invention [Problem to be solved by the invention]

[0003] However, conventional systems require individually examining many files to identify differences, which can be error prone and operationally expensive. [Means for solving the problem]

[0004] Provided herein are embodiments of systems, apparatus, devices, methods, and / or computer program products, and / or combinations and subcombinations thereof, for generating output indicating differences in data stored in heterogeneous data storage devices.

[0005] Further provided herein are system, apparatus, device, method and / or computer program product embodiments, and / or combinations and subcombinations thereof, for reconciling data stored in heterogeneous data storage devices.

[0006] A given embodiment includes a computer-implemented method for generating an output indicating differences between data stored in heterogeneous data storage devices. The method includes loading a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository. The first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets. The method further includes identifying one or more differences between the first and second data sets in the data repository. The method further includes generating an output including the first and second data sets and visual indicators indicating each of the one or more differences, and displaying the output.

[0007] Another embodiment includes a system for generating an output indicating differences between data stored in heterogeneous data storage devices. The system includes a memory and a processor coupled to the memory. The processor is configured to load a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository. The first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets. The processor is further configured to identify one or more differences between the first and second data sets in the data repository. The processor is further configured to generate an output including the first and second data sets and visual indicators indicating each of the one or more differences, and to cause the output to be displayed.

[0008] A further embodiment includes a non-transitory computer-readable medium having instructions stored thereon, wherein executing the instructions by one or more processors of the device causes the one or more processors to perform the following operations: loading a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository; the first data set including one or more first rows, the second data set including one or more second rows, and the data repository including a set of columns corresponding to the first and second data sets; identifying one or more differences between the first and second data sets in the data repository; generating an output including the first and second data sets and a visual indicator indicating each of the one or more differences; and displaying the output.

[0009] Another embodiment includes a computer-implemented method for generating an output indicating differences in data stored in heterogeneous data storage devices. The method includes loading a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository. The first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets. The method further includes identifying one or more differences between the first data set and the second data set in the data repository. The method further includes generating an output including the first data set and the second data set and a visual indicator indicating each of the one or more differences, and displaying the output.

[0010] Another embodiment includes a system for generating an output indicating differences in data stored in heterogeneous data storage devices. The system includes a memory and a processor coupled to the memory. The processor is configured to load a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository. The first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets. The processor is further configured to identify one or more differences between the first data set and the second data set in the data repository. The processor is further configured to generate an output including the first data set and the second data set and a visual indicator indicating each of the one or more differences, and to cause the output to be displayed.

[0011] A further embodiment includes a non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by one or more processors of the device causes the one or more processors to perform the following operations: loading a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository; the first data set including one or more first rows, the second data set including one or more second rows, and the data repository including a set of columns corresponding to the first and second data sets; identifying one or more differences between the first and second data sets in the data repository; generating an output including the first and second data sets and visual indicators indicating each of the one or more differences; and displaying the output.

[0012] Another embodiment includes a computer-implemented method for reconciling data stored in heterogeneous storage devices. The method includes retrieving one or more first data files containing a first data set stored in a first database. The method further includes retrieving one or more second data files containing a second data set stored in a second database. The method further includes identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. The method further includes loading a first subset of the first data set and a second subset of the second data set into a data repository. The first subset of data corresponds to the one or more first columns and includes one or more first rows, and the second subset of data corresponds to the one or more second columns and includes one or more second rows. The data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the method includes identifying one or more differences between the first subset of data and the second subset of data in the data repository and displaying the one or more differences.

[0013] Another embodiment includes a system for reconciling data stored in heterogeneous storage devices. The system includes a memory and a processor coupled to the memory. The processor is configured to retrieve one or more first data files containing a first data set stored in a first database. The processor is further configured to retrieve one or more second data files containing a second data set stored in a second database. The processor is further configured to identify one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. The processor is further configured to load a first subset of the first data set and a second subset of the second data set into a data repository. The first subset of data corresponds to the one or more first columns and includes one or more first rows, and the second subset of data corresponds to the one or more second columns and includes one or more second rows. The data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the processor is configured to identify one or more differences between the first subset of data and the second subset of data in the data repository and cause the one or more differences to be displayed.

[0014] A further embodiment includes a non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by one or more processors of the device causes the one or more processors to perform the following operations: searching one or more first data files stored in a first database, the first data set included; searching one or more second data files stored in a second database, the second data set included; identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading a first subset of the first data set and a second subset of the second data set into a data repository; the first subset of data corresponding to the one or more first columns and including one or more first rows, and the second subset of data corresponding to the one or more second columns and including one or more second rows. The data repository includes a set of columns corresponding to the first and second subsets of data, and the operation further includes identifying one or more differences between the first subset of data and the second subset of data in the data repository and displaying the one or more differences.

[0015] Another embodiment includes a method for reconciling data stored in heterogeneous storage devices. The method includes retrieving one or more first data files containing a first data set stored in a clinical database and one or more second data files containing a second data set stored in a safety database. The method further includes identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. The method further includes loading a first subset of the first data set and a second subset of the second data set into a data repository. The first subset of data corresponds to the one or more first columns and includes one or more first rows, and the second subset of data corresponds to the one or more second columns and includes one or more second rows. The data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the method includes identifying one or more differences between the first subset of data and the second subset of data in the data repository and displaying the one or more differences.

[0016] A further embodiment includes a system for reconciling data stored in heterogeneous storage devices. The system includes a memory and a processor coupled to the memory. The processor is configured to search one or more first data files containing a first data set stored in a clinical database and search one or more second data files containing a second data set stored in a safety database. The processor is further configured to identify one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. The processor is further configured to load a first subset of the first data set and a second subset of the second data set into a data repository. The first subset of data corresponds to the one or more first columns and includes one or more first rows, and the second subset of data corresponds to the one or more second columns and includes one or more second rows. The data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the processor is configured to identify one or more differences between the first subset of data and the second subset of data in the data repository and cause the one or more differences to be displayed.

[0017] A further embodiment includes a non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by one or more processors of the device causes the one or more processors to perform the following operations: searching one or more first data files stored in a clinical database, the first data files including a first data set; and searching one or more second data files stored in a safety database, the second data set including a second data set. The operations further include identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. Furthermore, the operations include loading a first subset of the first data set and a second subset of the second data set into a data repository. The first subset of data corresponds to the one or more first columns and includes one or more first rows, and the second subset of data corresponds to the one or more second columns and includes one or more second rows. The data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the operations include identifying one or more differences between the first subset of data and the second subset of data in the data repository and displaying the one or more differences. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a block diagram of an exemplary system for reconciling data stored in heterogeneous data storage devices. [Figure 2] FIG. 1 is a block diagram of a system for reconciling clinical and safety data according to some embodiments. [Figure 3] 1 is a graphical user interface portion of an output generated by a system for reconciling data stored in heterogeneous data storage devices, according to some embodiments. [Figure 4] 1 is a graphical user interface portion of an output generated by a system for reconciling data stored in heterogeneous data storage devices, according to some embodiments. [Figure 5] 1 is a flowchart illustrating a process for identifying differences in data stored in disparate databases, according to some embodiments. [Figure 6] 1 is a flowchart illustrating a process for generating and outputting an output indicating differences between data in heterogeneous data storage devices, according to some embodiments. [Figure 7] FIG. 1 is a block diagram of exemplary components of an apparatus according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0019] The accompanying drawings, which are incorporated in this application and form a part of the specification, illustrate the present disclosure and, together with the description, further serve to explain the principles of the disclosure and to enable those skilled in the art to make and use the disclosure.

[0020] The drawing in which an element first appears is typically indicated by the leftmost digit(s) in the corresponding reference number. In the drawings, like reference numbers may indicate identical or functionally similar elements.

[0021] Provided herein are embodiments of systems, apparatus, devices, methods, and / or computer program products, and / or combinations and subcombinations thereof, for generating output indicating differences in data stored in heterogeneous data storage devices.

[0022] Further provided herein are system, apparatus, device, method and / or computer program product embodiments, and / or combinations and subcombinations thereof, for reconciling data stored in heterogeneous data storage devices.

[0023] As previously discussed, conventional methods of reconciling data stored in disparate data storage devices are burdensome, costly, and subject to error. For example, in the field of pharmacovigilance (PV) operations, a first database may store clinical trial data, and a second database may store drug safety data. The first and second databases may contain similar columns storing data related to a particular drug or product. Thus, the first and second databases should share the same data associated with a particular drug or product. However, often inconsistent or missing data exists in either the first or second database. As a result, the data must be reconciled so that the data discrepancies can be resolved.

[0024] Conventional systems require manual review of spreadsheets containing data stored in the first and second databases. In one example, a clinical trial may include data from 24 spreadsheets. Reconciling the data from 24 spreadsheets per study each year against the data stored in the safety database can require 200 hours per year and cost over $25,000 per year. As a result, conventional systems are expensive to operate and prone to errors.

[0025] The embodiments described herein solve technical problems encountered in conventional systems by automatically reconciling data stored in data storage devices and indicating differences between the two data storage devices. In some embodiments, the differences between the two data storage devices are visually indicated. In some embodiments, a server retrieves one or more first data files containing a first data set stored in a first database. The server retrieves one or more second data files containing a second data set stored in a second database. The server further identifies one or more first columns in the one or more first data files and one or more second columns in the one or more second data files. The server loads a first subset of the first data set corresponding to the one or more first columns and a second subset of the second data set corresponding to the one or more second columns into a data repository. The first subset of data includes one or more first rows, the second subset of data includes one or more second rows, and the data repository includes sets of columns corresponding to the first and second subsets of data. Additionally, the server identifies one or more differences between the first subset of data and the second subset of data in the data repository and causes the one or more differences to be displayed.

[0026] In some embodiments, the server loads a first data set corresponding to one or more first columns of the first database and a second data set corresponding to one or more second columns of the second database into a data repository. The first data set comprises one or more first rows. The second data set comprises one or more second rows. The data repository includes a set of columns corresponding to the first and second data sets. Further, the server identifies one or more differences between the first and second data sets in the data repository. The server generates output including the first and second data sets and visual indicators for each of the one or more differences, and causes the output to be displayed.

[0027] The embodiments described herein include automatically identifying differences between different data stores, eliminating the generation of approximately 1,320 spreadsheets for the reconciliation process. Additionally, the output visually indicates the differences between the data in the different data stores, allowing for quick correction of the data stored in each data store. As a result, the embodiments described herein eliminate the extensive time, manpower, and potential errors incurred by conventional systems when reconciling data.

[0028] 1 is a block diagram of a system for reconciling data stored in heterogeneous data storage devices. The system may include a server 100, a client device 110, one or more first subsystems 114, one or more second subsystems 116, and a data repository 124. The devices in the system may be connected via a network. For example, the devices in the system may be connected via wired connections, wireless connections, or a combination of wired and wireless connections. In an exemplary embodiment, one or more portions of the network may be an ad-hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the public switched telephone network (PSTN), a cellular telephone network, a wireless network, a WiFi network, a WiMax network, any other type of network, or a combination of two or more such networks.

[0029] The client device 110 includes an application 112. The application 112 may be configured to send a request to reconcile data to the server 100. For example, in response to the client device 110 launching the application 112, a user may enter their user credentials. The application 112 may authenticate the user based on their user credentials. As a non-limiting example, the user may log in to the application 112 using a Windows Active Directory login. In response to authenticating the user, the application 112 may allow the user to provide input associated with a request to reconcile data stored in different data stores. The request may include an identifier for the data store. The request may also include a time period for the data to be reconciled. For example, the request may include instructions to reconcile data added to each data store in the past six months.

[0030] Server 100 may include a reconciliation application 102. Reconciliation application 102 may be configured to reconcile data stored in heterogeneous data stores in response to receiving requests from applications 112. This may include performing Extract, Transform, and Load (ETL) operations on the data stores, such as loading data, extracting data, transforming data, deleting data, transferring data, etc. Reconciliation application 102 may also be configured to generate output that identifies differences between data stored in the heterogeneous data stores. In some embodiments, reconciliation application 102 may be configured to periodically reconcile data stored in the heterogeneous data stores without being directed by applications 112.

[0031] The first subsystem 114 may be a third-party system configured to store one or more first databases 118. The first databases 118 may be data storage devices configured to store structured or unstructured data. The first subsystem 114 may store multiple first databases 118. Each of the databases stored in the first subsystem 114 may store a different type of data. Additionally, the first subsystem 114 may include one or more application program interfaces (APIs) 120.

[0032] The mediation application 102 may use the APIs 120 to access data stored in the first database 118. Each API 120 may provide access to a particular type of data. As a result, a particular API may be used to access the data stored in the first database 118. The API 120 may present the data stored in the first database 118 to the mediation application 102 in the form of a spreadsheet-like file (e.g., a MICROSOFT EXCEL® file).

[0033] The second subsystem 116 may be configured to store one or more second databases 122. The second databases 122 may be data storage devices configured to store structured or unstructured data. The reconciliation application 102 may be configured to access the data stored in the second databases 122 and reconcile the data stored in the first database 118 and the second databases 122.

[0034] The data repository 124 may be a data storage device configured to store data extracted from the first database 118 and the second database 122. The data repository 124 may include a subset of the columns of the first database 118 and the second database 122.

[0035] Figure 2 is a block diagram of a system for reconciling clinical and safety data according to some embodiments. Figure 2 is described with respect to Figure 1. As a non-limiting example, first subsystem 114 may be clinical subsystem 114-1 and second subsystem 116 may be safety subsystem 116-1.

[0036] The clinical subsystem 114-1 may include a first database 118. The first database 118 may be a clinical database. Specifically, the first database 118 may store data associated with one or more drug or product clinical trials. For example, the data may include a clinical trial timeline, drugs or products included in the clinical trial, information about subjects (e.g., users), drug or product effects, etc. The first database 118 may store data associated with a single clinical trial. Alternatively, the first database 118 may store data associated with multiple clinical trials. Furthermore, the clinical subsystem 114-1 may be associated with an entity responsible for conducting the clinical trials. In this manner, the entity may be conducting multiple different clinical trials. Accordingly, the clinical subsystem 114-1 may store multiple first databases 118. Each first database 118 may be associated with a single clinical trial. Each clinical trial may be referred to as a protocol.

[0037] The safety subsystem 116-1 may include a second database 122. The second database 122 may be a safety database. The second database 122 may store safety data associated with a drug or other product. For example, the second database 122 may store data related to pharmacovigilance (PV). The data may include information about the drug or product, information about the subject (e.g., the user), and information about adverse effects reported for use of the drug or product. The second database 122 may store data associated with a single drug or product (e.g., safety reports for a particular drug or product). Alternatively, the second database 122 may store data associated with multiple different drugs or products.

[0038] The first database 118 and the second database 120 may store information associated with adverse effects experienced by subjects and caused by products or medications. For example, an adverse event is a serious adverse event if it meets one of the following requirements: ·Causing death or being life-threatening, requiring hospitalization or prolonging current hospitalization; · causing persistent or significant impairment or incapacity; · causing birth defects; Or otherwise, · Medically significant, as the treatment and / or intervention is required to prevent one of the above conditions. Additionally, when conducting clinical trials of a drug or other product, it may be determined whether the adverse effect is a serious unexpected result adverse reaction (SUSAR).

[0039] The first database 118 may share common data associated with patients, drugs, and products. The first database 118 may also include a first set of columns, and the second database 122 may include a second set of columns. The first set of columns and the second set of columns may include one or more similar columns. To that end, data associated with the same drug, product, and / or patient with respect to one or more similar columns should be the same across the first database 118 and the second database 122. For example, both the first database 118 and the second database 122 may include data associated with a site number, subject ID, AER number, case identifier, reported term, severity, co-manifestation, preferred term, adverse effect onset date, adverse effect end date, outcome, age, sex, product, and causality. Data corresponding to these columns and associated with the same drug, product, and / or patient should be the same in the first database 118 and the second database 122.

[0040] As non-limiting examples, the data repository 124 may include a site number column, a subject ID column, an AER number column, a case identifier column, a reported term column, a severity column, a concurrent condition column, a preferred term column, an adverse effect onset date column, an adverse effect end date column, an outcome column, an age column, a gender column, a product column, and a causality column.

[0041] The Site Number column may contain a data value indicating an identifier for the clinical trial site. The Subject ID column may contain a data value indicating an identifier for the subject portion of the clinical trial or the subject reporting the adverse effect. The AER Number column may contain a data value indicating an adverse effect report (AER) number. The Case ID column may contain a data value indicating an identifier for the case for which the adverse effect is being reported. The Reported Term column may contain a data value corresponding to the event being reported (e.g., headache, nausea, fever, etc.). The event being reported may be an adverse effect. The Severity column may contain a data value indicating whether the adverse effect is serious or not. The Preferred Term column may contain a data value indicating an identifier and a preferred description for the event being reported. The Adverse Effect Onset Date column may store a data value indicating the date the adverse effect began. The Adverse Effect End Date column may store a data value indicating the date the adverse effect ended. The Outcome column may store a data value indicating the outcome (resolution or recovery) of the effect or adverse event. The Age column may store a data value indicating the age of the subject. The gender column may store a data value indicating the gender of the subject. The product column may contain a data value indicating the identifier of the product (for the clinical trial or which may have caused the adverse effect). The causation column may contain a data value indicating whether the product caused the adverse effect.

[0042] However, differences may exist between the first database 118 and the second database 122. For example, the first database 118 may be updated while the second database 122 is not, or vice versa. Furthermore, there may be erroneous updates / additions / deletions of data in either the first database 118 or the second database 122. As a result, there may be data inconsistencies between the first database 118 and the second database 122. Furthermore, there may be missing data in either the first database 118 or the second database 122. The missing data may include a single missing entry or a missing row.

[0043] In some embodiments, the reconciliation application 102 may receive a request from the application 112 to reconcile data in the first database 118 and the second database 122. The request may include identifiers for the clinical subsystem 114-1, the first database 118, the safety subsystem 116-1, and / or the second database 122. Alternatively or additionally, the request may include an identifier for a particular clinical trial (e.g., a protocol ID), drug, or product. The request may also include a time period. Specifically, the request may include instructions to reconcile the data with respect to data loaded into the first database 118 and the second database 122 over a given time period.

[0044] The reconciliation application 102 may identify the first database 118 and the second database 122 using identifiers for the clinical subsystem 114-1, the first database 118, the safety subsystem 116-1, and / or the second database 122. Alternatively or additionally, the reconciliation application 102 may identify the first database 118 and the second database 122 based on an identifier for a particular clinical trial (e.g., a protocol ID), drug, or product.

[0045] The reconciliation application 102 may access data from the first database 118 by interfacing with the API 120. The API 120 may present one or more data files to the reconciliation application 102. For example, the API 120 may present the data from the first database 118 in the form of a spreadsheet-like file. The spreadsheet may include a first set of columns from the first database 118. As non-limiting examples, the first set of columns may include project ID, project, internal study ID, environment, internal subject ID, internal study site ID, subject name or identifier, SDVTier, internal site ID, site name, site number, site group, internal instance id, folder instance name, instance repeat count, internal folder id, folder OID, folder name, folder sequence number, total days since study start, internal ID for data page, eCRF page name, sequence number of eCRF page in folder, clinical record date, internal ID for record, earliest data created, timestamp of last save in clinical view, last data update time, coder hierarchical, SE site number, study environment site number, age, age character, sex, sex code, ethnicity, ethnicity code, race, race code, age unit, age unit code, enrollment date, enrollment date character, birth year, birth year character, age at onset of SAE for SG, associated adverse effect record, associated adverse effect code, reported term for the adverse effect, start date of the adverse effect, severity, end date of the adverse effect, etc.

[0046] The reconciliation application 102 may identify one or more columns of the first set of columns that correspond to the data to be loaded into the data repository 124. The data to be loaded into the data repository 124 may correspond to one or more columns of the data repository 124. For example, the reconciliation application 102 may identify an additional column in one or more data files that stores data associated with a site number, subject ID, AER number, case identifier, reported term, severity, concurrent symptoms, preferred term, adverse effect onset date, adverse effect end date, outcome, age, sex, product, and causality.

[0047] The reconciliation application 102 may perform ETL operations to extract data corresponding to data from one or more columns in the first set of columns. Additionally, the reconciliation application 102 may perform ETL operations to transform the extracted data so that it can be loaded into the data repository 124. The reconciliation application 102 may map the extracted data to one or more columns in the data repository 124. Each of the columns may be configured to receive data in a specific format. Multiple elements of the extracted data may need to be combined and transformed to be loaded into each column. Transformation operations may include cleaning, deduplication, formatting correction, key reconstruction, derivation, filtering, joining, splitting, data validation, summarization, aggregation, consolidation, etc. For example, the reconciliation application 102 may implement the following transformation and mapping rules:

[0048] [Table 1] [Table 2] [Table 3] [Table 4] [Table 5] [Table 6]

[0049] With regard to the above transformations and mappings, please note the following acronyms and definitions:

[0050] [Table 7]

[0051] The reconciliation application 102 may detect incorrect or partial dates and may convert missing days and months in incorrect or partial dates to 01 and JAN, respectively.

[0052] The reconciliation application 102 may perform ETL operations to load the transformed data into the data repository 124. The transformed data may be loaded into each column of the data repository 124.

[0053] The reconciliation application 102 may interface with the safety subsystem 116-1 to extract data from the second database 122. In some embodiments, the reconciliation application 102 may retrieve one or more data files that contain copies of the data stored in the second database 122. The one or more data files may include a second set of columns.

[0054] The reconciliation application 102 may identify one or more columns of the second set of columns that correspond to the data to be loaded into the data repository 124. For example, the reconciliation application 102 may identify one more column that stores data associated with a site number, subject ID, AER number, case identifier, term reported, severity, concurrent symptoms, preferred term, adverse effect onset date, adverse effect end date, outcome, age, sex, product, and causality.

[0055] The reconciliation application 102 may perform ETL operations to extract data corresponding to one or more columns in the second set of columns. Further, the reconciliation application 102 may perform ETL operations, as described above, to transform the extracted data so that it can be loaded into the data repository 124.

[0056] The reconciliation application 102 may perform ETL operations to load the transformed data into the data repository 124. The transformed data may be loaded into each column of the data repository 124.

[0057] The reconciliation application 102 may calculate a correlation of each row of data corresponding to the first database 118 in the data repository 124 with each row of data corresponding to the second database 122 based on identifier values ​​stored in each row of data corresponding to the first database 118 and each row of data corresponding to the second database 122. For example, each row of data corresponding to the first database 118 in the data repository 124 may include a location number, a subject ID, and an AER number. Furthermore, each row of data corresponding to the second database 122 in the data repository 124 may also include a location number, a subject ID, and an AER number. In this manner, the reconciliation application 102 may match one or more of the location number, the subject ID, and the AER number from the row corresponding to the first database 118 in the data repository 124 to each row corresponding to the second database 120 in the data repository 124.

[0058] The reconciliation application 102 may compare the data values ​​of each row corresponding to the first database 118 in the data repository 124 against the correlated row corresponding to the second database 122 in the data repository 124. The reconciliation application 102 may identify differences in the data values ​​based on the comparison. Differences may include data inconsistencies or missing data values. A data inconsistency may indicate that the first database 118 and the second database 120 have different data values ​​in entries corresponding to common columns and rows. A missing data value may indicate that a predetermined data value is missing in either the first database 118 or the second database 124.

[0059] The reconciliation application 102 may generate output indicating the identified differences. The reconciliation application 102 may also generate visual indicators to highlight the identified differences. The visual indicators may differ based on the type of difference. The types of difference may include, but are not limited to, missing data values ​​in the first database 118, missing data values ​​in the second database 122, and inconsistencies in specific columns.

[0060] The output may be a safety reconciliation report 200. The safety reconciliation report 200 may be a spreadsheet containing columns in the data repository 124, with data values ​​in each row corresponding to the first database 118 and the second database 122. The safety reconciliation report 200 may be output as a file to the application 112. Alternatively, the safety reconciliation report 200 may be displayed on a user interface in the application 112.

[0061] In some embodiments, the reconciliation application 102 may automatically perform actions on the first database 118 or the second database 120 to resolve the identified discrepancies. For example, if the reconciliation application 102 identifies a missing data value in the first database 118 that is present in the second database 120, the reconciliation application 102 may store the data value indicated in the second database 120 in the first database 118. Similarly, if the reconciliation application 102 identifies a missing data value in the second database 120 that is present in the first database 118, the reconciliation application 102 may store the data value indicated in the first database 118 in the second database 120.

[0062] Furthermore, if the reconciliation application 102 identifies an inconsistency in the data values ​​in the first database 118 and the second database 120, the reconciliation application 102 may determine which data value is likely to be the accurate data value. For example, the reconciliation application 102 may determine that the precision of other data values ​​in the row containing the data value and corresponding to the first database 118 exceeds a predetermined threshold. Furthermore, the reconciliation application 102 may determine that the precision of other data values ​​in the row containing the data value and corresponding to the second database 122 is less than a predetermined threshold. As a result, the reconciliation application 102 may determine that the data value in the first database 118 is likely accurate. Accordingly, the reconciliation application 102 may update the data value in the second database 122 to match the data value in the first database 118.

[0063] In some embodiments, the reconciliation application 102 may use the following logic to extract severity and outcome data from one or more data files received from the clinical subsystem 114-1.

[0064] Scenario 1 For a given case, one or more data files may contain multiple rows for consecutive events of that case (e.g., multiple headache events for the same case). Both events are labeled as serious. In this scenario, the reconciliation application 102 may extract the outcome (e.g., recovery / resolution) from the most recent record of the two events and the earliest adverse effect onset date.

[0065] Scenario 2 For a given case, one or more data files may contain multiple rows for consecutive events of that case (e.g., multiple headache events for the same case). The first event is labeled as serious and the second event is not labeled as serious. In this scenario, the reconciliation application 102 may extract the outcome (e.g., recovery / resolution) from the most recent record of the two events and the earliest adverse effect onset date.

[0066] Scenario 3 The one or more data files may contain multiple rows for a case for consecutive events of the case (e.g., multiple headache events for the same case) and for different events for the same case (e.g., nausea). The reconciliation application 102 may extract data for consecutive events and different events. The reconciliation application 102 extracts the earliest adverse effect onset date for consecutive events.

[0067] Scenario 4 For a given case, one or more data files may contain multiple rows for consecutive events of the case (e.g., multiple headache events in the same case) and different events of the same case (e.g., nausea). Consecutive events and different events are not labeled as severe. The reconciliation application 102 may not process these events for the safety gateway reconciliation report 200.

[0068] 3 is a graphical user interface portion of an output generated by a system for reconciling data stored in disparate data storage devices, according to some embodiments. As shown with respect to FIG. 2, the output may be a safety reconciliation report 200. The output may be a graphical user interface (GUI) displayed on an application 112 executing on the client device 110. Alternatively, the output may be a GUI displayed on an internet browser on the client device 110. In another example, the output may be a file (e.g., PDF, spreadsheet, DOC, TXT, CSV, etc.) sent to the client device 110.

[0069] The output may include a GUI 300. The GUI 300 may provide a summary of the output. Specifically, the GUI 300 may provide a summary of the safety reconciliation report. The summary may include a protocol number. The protocol number may be an identifier for the clinical trial. The safety reconciliation report may correspond to the particular clinical trial corresponding to the protocol number.

[0070] The summary may further include the number of cases identified in the clinical database (e.g., the first database 118 as shown in FIG. 1), the number of cases identified in the safety database (e.g., the second database 122 as shown in FIG. 2), the total number of events identified in the clinical database, and the total number of events identified in the safety database.

[0071] The summary may further indicate that the safety reconciliation report identifies all cases missing in the safety database, all cases missing in the clinical database, and all inconsistencies in data values ​​between the safety and clinical databases.

[0072] 4 is a graphical user interface portion of an output generated by a system for reconciling data stored in heterogeneous data storage devices, according to some embodiments. As described above, the output may be a safety reconciliation report. The safety reconciliation report may include a GUI 400. The GUI 400 may be displayed after the GUI 300 of FIG. 3.

[0073] GUI 400 may be a spreadsheet showing identified differences between a clinical database (e.g., first database 118 as shown in FIG. 1 ) and a safety database (e.g., second database 122 as shown in FIG. 1 ). GUI 400 may include columns 420 from a data repository (e.g., data repository 124 as shown in FIG. 1 ). Columns 420 may have corresponding columns in the clinical and safety databases. Additionally, columns 420 may relate to clinical trials and safety reports related to drugs and products. As non-limiting examples, columns 420 may include source, site number, subject ID, AER number, case ID, reported term, severity, concurrent symptoms, preferred term, adverse effect onset date, adverse effect end date, outcome, age, sex, product, causality, and mismatch. Columns 420 may store data related to drugs, products, subjects, and reported adverse effects for subjects associated with the drug or product.

[0074] The source column may contain data values ​​that indicate whether a row corresponds to the clinical database or the safety database, and the mismatch column may contain data values ​​that indicate whether a mismatch exists between the data values.

[0075] The rows of the spreadsheet in GUI 400 may be from a data repository. However, each row may correspond to a clinical or safety database. The source column may indicate whether the row corresponds to the clinical or safety database.

[0076] The spreadsheet in GUI 400 may include a legend 401 and a protocol ID 402. Legend 401 may indicate the type of identified difference between the clinical trial and safety databases and a corresponding visual indicator. The visual indicator may be a different color, pattern, haptic effect, animation, shape, etc. As a non-limiting example, legend 401 may indicate the following types of identified differences: gaps in the clinical database, gaps in the safety database, event / case attribute inconsistencies, and causality inconsistencies. Although not shown in legend 401, the preferred term inconsistency may be a type of difference. An event / case attribute inconsistency may be an inconsistency corresponding to any of the data values ​​corresponding to the severity, adverse effect stop date, outcome, age, or gender columns. Legend 401 may also indicate that if no identified differences exist, the visual indicator may be absent (e.g., blank or white background).

[0077] The spreadsheet in GUI 400 may be arranged such that correlated rows corresponding to safety and clinical data are grouped together. As a non-limiting example, rows corresponding to clinical data are placed before correlated rows corresponding to safety data. Additionally, a visual indicator of the difference in data values ​​may be displayed above or for each data value.

[0078] As a non-limiting example, row 403 may correspond to a clinical database, and row 404 may correspond to a safety database. Rows 403 and 404 may be correlated to one another. For example, rows 403 and 404 may be associated with the same subject, case, and AER number. GUI 400 may indicate an identified inconsistency with respect to the data values ​​in rows 403 and 404 in the severity column. Row 403 may indicate a data value of "YES" in the severity column, and row 404 may indicate a data value of "NO" in the severity column. A visual indicator may be displayed for the data values ​​in rows 403 and 404 in the severity column. The visual indicator may correspond to a visual indicator for an inconsistency in an event / case attribute, as shown in legend 401.

[0079] GUI 400 may also indicate a mismatch in the causality column for the data values ​​in rows 403 and 404. For example, row 403 may indicate a data value of "not suspected" in the causality column, and row 404 may indicate a blank data value in the causality column. Accordingly, a visual indicator may be displayed for the data values ​​in rows 403 and 404 in the causality column. The visual indicator may correspond to a visual indicator for a causality mismatch, as shown in legend 401.

[0080] Continuing with the non-limiting example, row 406 may correspond to a clinical database. Row 408 may correspond to a safety database. Rows 406 and 408 may be correlated to one another. For example, rows 406 and 408 may be associated with the same subject, case, and AER number. GUI 400 may indicate that a data value for row 408 is missing in the safety database. A visual indicator may be displayed for the entire row 408. The visual indicator may correspond to a visual indicator for a gap in the safety database, as shown in legend 401.

[0081] Continuing with the non-limiting example, row 410 may correspond to a clinical database, and row 412 may correspond to a safety database. Rows 410 and 412 may be correlated to one another. Rows 410 and 412 may be associated with the same subject, case, and AER number. GUI 400 may indicate that a data value for row 410 is missing in the clinical database. A visual indicator may be displayed for the entire row 412. The visual indicator may correspond to a visual indicator for a gap in the clinical database, as shown in legend 401.

[0082] Continuing with the non-limiting example, row 414 may correspond to a clinical database, and row 416 may correspond to a safety database. Rows 414 and 416 may be correlated to one another. For example, rows 414 and 416 may be associated with the same subject, case, and AER number. GUI 400 may indicate an identified inconsistency with respect to the data values ​​in rows 414 and 416 in the outcome column. For example, row 414 may indicate a blank data value in the outcome column, and row 416 may indicate a data value of "recovered" in the outcome column. A visual indicator may be displayed for the data values ​​in rows 414 and 416 in the outcome column. The visual indicator may correspond to a visual indicator for an inconsistency in an event / case attribute, as shown in legend 401.

[0083] GUI 400 may also indicate a mismatch in the causality column for the data values ​​in rows 414 and 416. Row 414 may indicate a "suspect" data value in the causality column, and row 416 may indicate a blank data value in the causality column. A visual indicator may be displayed for the data values ​​in rows 414 and 416 in the causality column. The visual indicator may correspond to a visual indicator for a causality mismatch, as shown in legend 401.

[0084] The spreadsheet may include several pages. A user may filter the spreadsheet in GUI 400 based on column type, difference type, data value (e.g., case ID), source (e.g., clinical or safety database), etc. Additionally, a user may export data to a commonly accepted format such as EXCEL, DOC, PDF, etc. A user may also export only filtered or selected data to a commonly accepted format. As a non-limiting example, a user may export instances of missing values ​​in a clinical database to a commonly accepted format.

[0085] The output may also include a summary of the queries that were performed on the clinical database, the safety database, and the data repository to generate the output.

[0086] As a non-limiting example, visual indicators for GUI 400 may be generated as follows for the following scenarios:

[0087] Scenario 1 – Product inconsistency. If there is an identified inconsistency in the data values ​​in the product column of the correlated rows, GUI 400 includes a row for each combination of product and reported term for both the clinical and safety databases. The reported term column, concurrent symptom column, adverse effect onset date column, and product column may be indicated as missing data values ​​for each clinical and safety database. For example, a first row corresponding to the clinical database may include a data value for product A in the product column. A second row corresponding to the safety database and correlated to the first row may include a data value for product B in the product column. In this scenario, GUI 400 includes a row corresponding to the clinical database with a data value for product A in the product column and a correlated row corresponding to the safety database. The correlated row corresponding to the safety database may indicate missing data values ​​for the reported term column, concurrent symptom column, and adverse effect onset date column. Additionally, GUI 400 includes a row corresponding to the safety database with a data value for product B in the product column and a correlated row corresponding to the clinical database. The correlated rows corresponding to the clinical database may show missing data values ​​for the reported term column, the concurrent symptoms column, and the adverse effect onset date column.

[0088] Scenario 2 – Combination of multiple products and events in safety and clinical settings, including product inconsistencies. Two correlated rows may contain data values ​​for multiple events (e.g., multiple reported terms) and multiple products. The events may be the same, but the products may be different. In this scenario, GUI 400 may contain a row for each combination of product and reported term for both the clinical and safety databases.

[0089] For example, the data value for the Product column in a first row may correspond to the clinical database and include Product A, while the data value for the Product column in a second row may correspond to the safety database and include Products B and C. Further, the data value for the Reported Terms column in the first row may include fever and cold. Similarly, the data value for the Reported Terms column in the second row may also include fever and cold. In this scenario, GUI 400 may include a row corresponding to the clinical database that includes a data value for the Product column, Product A, and a Reported Terms column, Fever. The correlated row corresponding to the safety database may indicate missing data values ​​for the Reported Terms column, Concurrent Symptom column, Adverse Effect Onset Date column, and Product column. Further, GUI 400 may include a row corresponding to the clinical database that includes a data value for the Product column, Product A, and a Reported Terms column, Cold. The correlated row corresponding to the safety database may indicate missing data values ​​for the Reported Terms column, Concurrent Symptom column, Adverse Effect Onset Date column, and Product column.

[0090] Additionally, GUI 400 may include a row corresponding to the safety database that includes a data value for the product, Product B, and a reported term column of Fever. As a result, the correlated row corresponding to the clinical database may show missing data values ​​for the reported term column, the concurrent symptom column, the adverse effect onset date column, and the product column. Additionally, GUI 400 may include a row corresponding to the safety database that includes a data value for the product column, Product B, and a reported term column of Cold. As a result, the correlated row corresponding to the clinical database may show missing data values ​​for the reported term column, the concurrent symptom column, the adverse effect onset date column, and the product column.

[0091] GUI 400 may include a row corresponding to the safety database that includes a data value for a product, Product C, and a reported term column of Fever. A correlated row corresponding to the clinical database may indicate missing data values ​​for the reported term column, the concurrent symptom column, the adverse effect onset date column, and the product column. Additionally, GUI 400 may include a row corresponding to the safety database that includes a data value for a product, Product C, and a reported term column of Cold. A correlated row corresponding to the clinical database may indicate missing data values ​​for the reported term column, the concurrent symptom column, the adverse effect onset date column, and the product column.

[0092] Scenario 3 – Blinded product If a product is blinded in a clinical trial, the row corresponding to the clinical database may include the string "masked for" before the product name in the product column. When comparing the data values ​​in the product column against the data values ​​in the product column in the correlated row corresponding to the safety database, the string "masked for" may be removed.

[0093] Scenario 4 – Concurrent symptoms mapped to the safety database A row corresponding to the safety database may include multiple data values ​​for the Reported Terms column and multiple data values ​​for the Concurrent Conditions column. For example, a row may include cancer progression and include cancer as a data value in the Reported Terms column and include an "N" for cancer progression and a "Y" for cancer in the Concurrent Conditions column. A row corresponding to the clinical database may include cancer progression as a data value in the Reported Terms column and include an "N" for the Concurrent Conditions column.

[0094] GUI 400 may include two correlated rows corresponding to the clinical and safety databases, where the data value in the Reported Term column is Cancer Progression and the data value for the Concurrent Condition is "N." Additionally, GUI 400 may include a row corresponding to the Safety database, where the data value in the Reported Term column is Cancer and the data value for the Concurrent Condition is "Y." The correlated row corresponding to the Clinical Database may indicate missing data values ​​for the Reported Term column, the Concurrent Condition column, the Adverse Effect Onset Date column, and the Product column.

[0095] Scenario 5 - Causality If there is any inconsistency between the data values ​​of the causal relationships in the corresponding rows of the clinical and safety databases, the GUI 400 may indicate the inconsistency between the data values, even if there is a blank data value in the causal relationship in the corresponding row of either the clinical and safety databases.

[0096] Scenario 6 - Multiple episodes of an event with different adverse effect onset dates A row corresponding to a safety or clinical database may show multiple events in the Reported Terms column and multiple adverse effect onset dates. For example, a first row corresponding to a clinical database may show "headache" in the Reported Terms column, an adverse effect onset date of 6 / 8 / 2019, and an adverse effect end date of 6 / 22 / 2019. A second row corresponding to a safety database may show two events of "headache" in the Reported Terms column, an adverse effect onset date of 6 / 8 / 2019 for the first headache event, an adverse effect onset date of 6 / 20 / 2019 for the second headache event, an adverse effect end date of 6 / 14 / 2019 for the first headache event, and an adverse effect end date of 6 / 22 / 2019 for the second headache event.

[0097] In this scenario, GUI 400 may include a row corresponding to the clinical database, where the data value in the Reported Term column is Headache, the data value in the Adverse Effect Onset Date column is 6 / 8 / 2019, and the Adverse Effect End Date column is 6 / 22 / 2019. Additionally, GUI 400 may include a row corresponding to the Safety database, where the data value in the Reported Term column is Headache (e.g., First Headache Event), the data value in the Adverse Effect Onset Date column is 6 / 8 / 2019, and the Adverse Effect End Date column is 6 / 14 / 2019. GUI 400 may visually indicate the inconsistency in the data values ​​for the Adverse Effect End Date column.

[0098] Additionally, GUI 400 may include a row corresponding to the safety database, where the data value for the reported term is headache (e.g., second headache event), the data value for the adverse effect onset date column is 6 / 20 / 2019, and the adverse effect end date is 6 / 22 / 2019. Additionally, a correlated row corresponding to the clinical database may indicate missing data values ​​for the reported term column, the concurrent symptoms column, the adverse effect onset date column, and the product column.

[0099] Scenario 7 - Multiple episodes of an event with different adverse effect onset dates A corresponding row in the clinical or safety database may contain incorrect or partial dates for the adverse effect onset date and adverse effect end date columns. In this scenario, GUI 400 may visually indicate the inconsistency in the data values ​​in the adverse effect onset date column.

[0100] Scenario 8 - Inconsistent adverse effect onset dates A row corresponding to a clinical or safety database may include multiple different adverse effect onset dates. As a result, these may be treated as separate adverse events. For example, a first row corresponding to a clinical database may include 6 / 7 / 2019 for the adverse effect onset date column, and a second row corresponding to a safety database may include 6 / 8 / 2019 for the adverse effect onset date column. In this scenario, GUI 400 may include a row corresponding to the safety database that indicates 6 / 8 / 2019 for the adverse effect onset date column. Additionally, the correlated row corresponding to the clinical database may indicate missing data values ​​for the reported term column, the concurrent symptom column, the adverse effect onset date column, and the product column.

[0101] Additionally, GUI 400 may include a row corresponding to the clinical database that indicates 6 / 7 / 2019 for the adverse effect onset date column. A correlated row corresponding to the safety database may indicate missing data values ​​for the reported term column, the concurrent symptoms column, the adverse effect onset date column, and the product column.

[0102] 5 is a flowchart illustrating a process for identifying differences in data stored in disparate databases, according to some embodiments. Method 500 can be performed by processing logic, which may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be appreciated that not all steps are required to practice the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 5, as will be understood by one of ordinary skill in the art.

[0103] The method 500 shall be described with reference to Figure 1. However, the method 500 is not limited to that exemplary embodiment.

[0104] At 502, the reconciliation application 102 of the server 100 receives a request from the application 112 of the client device 110 to reconcile data stored in the first database 118 and the second database 120. The request may include identifiers of the first database 118 and the second database 120. As an example, the first database 118 may store clinical trial data, and the second database 122 may store safety data associated with a drug or product. As such, the request may also include identifiers of the clinical trial, the product, and / or the drug. Thus, the reconciliation application 102 may identify the first database 118 and the second database 122 using identifiers of one or more of the first database 118, the second database 120, the clinical trial, the product, or the drug.

[0105] At 504, the reconciliation application 102 retrieves one or more first data files containing the first data set stored in the first database 118. The reconciliation application 102 may access the one or more data files by interfacing with the API 120. The API 120 may present the one or more data files to the reconciliation application 102. The one or more data files may be spreadsheets. The one or more data files may include columns of the first database 118.

[0106] At 506, the reconciliation application 102 retrieves one or more second data files containing a second data set stored in the second database 122. The one or more second data files may include columns of the second database 122.

[0107] At 508, the reconciliation application 102 identifies one or more first columns in one or more first data files and one or more second columns in one or more second data files. The reconciliation application 102 identifies the one or more first and second columns that correspond to data to be loaded into the data repository 124.

[0108] At 510, the reconciliation application 102 extracts data corresponding to one or more first columns from one or more first data files and extracts data corresponding to one or more second columns from one or more second data files. The reconciliation application 102 may extract the data by performing ETL operations.

[0109] At 512, the reconciliation application 102 transforms the data extracted from the one or more first data files and the one or more second data files. The reconciliation application 102 may transform the data by performing ETL operations. Additionally, the reconciliation application 102 may transform the extracted data so that the extracted data is compatible with and can be loaded into the data repository 124.

[0110] At 514, the reconciliation application 102 loads the extracted and transformed data from the one or more first data files and the one or more second data files into the data repository 124. The reconciliation application 102 may load the data by performing an ETL operation. The transformed data may be loaded into columns of the data repository 124. The columns of the data repository 124 may correspond to one or more first columns and one or more second columns.

[0111] At 516, the reconciliation application 102 identifies one or more differences between the data extracted from the one or more first data files and the data extracted from the one or more second data files in the data repository. The differences may be inconsistencies in data values ​​or missing data values.

[0112] At 518, the reconciliation application 102 causes one or more differences to be displayed. The output may include a visual indicator identifying the difference. The visual indicator may differ based on the type of difference. The visual indicator may include, but is not limited to, highlighting with a different color, animation, pattern, tactile output, gradient, shape, or visual effect.

[0113] 6 is a flowchart illustrating a process for generating and outputting an output indicating differences between data in heterogeneous data storage devices, according to some embodiments. Method 600 can be performed by processing logic, which may comprise hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be appreciated that not all steps are required to practice the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 6, as will be understood by one of ordinary skill in the art.

[0114] The method 600 will be described with reference to Figures 1-2. However, the method 600 is not limited to that exemplary embodiment.

[0115] At 602, the reconciliation application 102 loads a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository. The first data set includes one or more first rows and the second data set includes one or more second rows. The data repository includes a set of columns corresponding to the first and second data sets.

[0116] At 604, the reconciliation application 102 identifies one or more differences between the first data set and the second data set in the data repository. A difference may be an inconsistency in data values ​​between the first data set and the second data set. Alternatively, a difference may be a missing data value in the first data set or the second data set.

[0117] At 606, the reconciliation application 102 generates an output including the first and second data sets and a visual indicator indicating each of the one or more differences. The output may be one or more graphical user interfaces (GUIs) displayed on the client device 110. Additionally, the output may be a file such as a spreadsheet. The visual indicator may differ based on the type of difference. For example, the type of difference may be an inconsistency or a missing value.

[0118] At 608, the reconciliation application 102 causes the output to be displayed in the user interface of the application 112. The output may be a file (e.g., a spreadsheet) that is sent to the client device 110.

[0119] Various embodiments may be implemented using one or more computer systems, such as, for example, computer system 700 shown in Figure 7. Computer system 700 may be used to implement, for example, methods 500 of Figure 5 and 600 of Figure 6. Furthermore, computer system 700 may be at least a portion of server 100, client device 110, first subsystem 114, second subsystem 116, first database 118, second database 120, and data repository 124, as shown in Figure 1. For example, computer system 700 routes communications to various applications. Computer system 700 may be any computer capable of performing the functions described herein.

[0120] Computer system 700 may be any known computer capable of performing the functions described herein.

[0121] Computer system 700 includes one or more processors (also referred to as central processing units or CPUs), such as processor 704. Processor 704 is connected to a communication infrastructure or bus 706.

[0122] One or more processors 704 may each be a graphics processing unit (GPU). In some embodiments, a GPU is a processor that is a dedicated electronic circuit designed to handle mathematically intensive applications. A GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common in computer graphics applications, images, video, etc.

[0123] The computer system 700 also includes one or more user input / output devices 703, such as a monitor, keyboard, pointing device, etc., which communicate with a communications infrastructure 706 via one or more user input / output interfaces 702.

[0124] The computer system 700 also includes a main or primary memory 708, such as random access memory (RAM). The main memory 708 may include one or more levels of cache. The main memory 708 stores control logic (i.e., computer software) and / or data therein.

[0125] Computer system 700 may also include one or more secondary storage devices or memories 710. Secondary memory 710 may include, for example, a hard disk drive 712 and / or a removable storage device or drive 714. Removable storage drive 714 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.

[0126] Removable storage drive 714 may interact with removable storage device 718. Removable storage device 718 includes a computer usable or readable storage device having computer software (control logic) and / or data stored thereon. Removable storage device 718 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. Removable storage drive 714 reads from and / or writes to removable storage device 718 in a known manner.

[0127] According to an exemplary embodiment, secondary memory 710 may include other means, devices, or approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 700. Such means, devices, or approaches may include, for example, removable storage 722 and interface 720. Examples of removable storage 722 and interface 720 may include a program cartridge and cartridge interface (such as found in a video game device), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage device and associated interface.

[0128] Computer system 700 may further include a communications or network interface 724. Communications interface 724 enables computer system 700 to communicate and interoperate with any combination of remote devices, remote networks, remote entities, etc. (individually and collectively indicated by the numeral 728). For example, communications interface 724 may enable computer system 700 to communicate with remote devices 728 via communications path 726, which may be wired and / or wireless, and which may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 700 via communications path 726.

[0129] In some embodiments, tangible, non-transitory devices or articles of manufacture comprising a tangible, non-transitory computer-usable or readable medium having control logic (software) thereon are also referred to herein as computer program products or program storage devices. This includes, but is not limited to, computer system 700, main memory 708, secondary memory 710, removable storage devices 718 and 722, and tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 700), causes such data processing devices to operate as described herein.

[0130] Based on the teachings contained herein, it will be apparent to one skilled in the art how to make and use embodiments of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 7. In particular, embodiments may operate using software, hardware, and / or operating system implementations other than those described herein.

[0131] It should be recognized that the Detailed Description section and no other section is intended to interpret the claims. The other sections may set forth one or more, but not all, example embodiments contemplated by the inventors, and are therefore not intended to limit the scope of this disclosure or the appended claims in any way.

[0132] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and variations are possible and are within the scope and spirit of the disclosure. For example, without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware, and / or entities shown in the drawings and / or described in this application. Furthermore, the embodiments (whether or not explicitly described in this application) have significant utility for fields and applications beyond the examples described in this application.

[0133] Embodiments are described herein with the aid of functional building blocks that illustrate implementations of, and relationships between, specified functions. The boundaries of these functional building blocks are arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or equivalents thereof) are appropriately performed. Alternative embodiments may also implement functional blocks, steps, operations, methods, etc., using an order different from that described herein.

[0134] References herein to “one embodiment,” “embodiment,” “exemplary embodiment,” or similar phrases indicate that, while the described embodiment may include a particular feature, structure, or characteristic, all embodiments need not necessarily include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described with respect to one embodiment, incorporating such feature, structure, or characteristic into other embodiments would be within the knowledge of one of ordinary skill in the art, whether or not explicitly mentioned or described herein. Furthermore, some embodiments may be described using the terms “coupled” and “connected,” along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some embodiments may be described using the terms “connected” and / or “connected” to indicate that two or more components are in direct physical or electrical contact with each other. However, the term “coupled” may also mean that two or more components are not in direct contact with each other, but still cooperate or interoperate with each other.

[0135] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. 1. A method for reconciling data stored on heterogeneous data storage devices, the method comprising: retrieving, by a processor, one or more first data files containing the first data set stored in a first database; retrieving, by the processor, one or more second data files containing a second data set stored in a second database; identifying, by the processor, one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading, by the processor, a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The above method is identifying, by the processor, one or more differences between the first data subset and the second data subset in the data repository; causing the processor to display the one or more differences; Identifying the one or more differences includes: calculating, by the processor, a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; and matching, by the processor, a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows. method.

2. the first identifier value is a combination of two or more data elements stored in each row of the one or more first rows, and the second identifier value is a combination of two or more data elements stored in each row of the one or more second rows; The method of claim 1.

3. the one or more differences include inconsistencies or missing values; The method of claim 1.

4. The method further includes generating, by the processor, a visual indicator for each of the one or more differences; Certain types of visual indicators correspond to certain types of differences, The method of claim 1.

5. retrieving the one or more first data files includes interfacing, by the processor, with an application program interface (API) of a type corresponding to a type of first data set; The method of claim 1.

6. the first dataset comprises clinical trial data and the second dataset comprises safety data; The method of claim 1.

7. performing, by the processor, an operation on the first database or the second database that resolves at least one difference of the one or more differences. The method of claim 1.

8. at least one difference of the one or more differences is a mismatch between a first data value in a first row of the one or more first rows and a second data value in a second row of the one or more second rows; the first and second data values ​​correspond to the same column of the set of columns; The method of claim 1.

9. determining, by the processor, a remainder of the data values ​​in the one or more first rows having a precision level exceeding a threshold amount; determining, by the processor, that the second data value is incorrect based on a remaining precision level of data values ​​in the one or more first rows that exceeds the threshold amount; updating, by the processor, an entry in the second database corresponding to the second data value to the first data value. The method of claim 8.

10. determining, by the processor, a remainder of the data values ​​in the one or more second rows having a precision level exceeding a threshold amount; determining, by the processor, that the first data value is incorrect based on a remaining precision level of data values ​​in the one or more second rows that exceeds the threshold amount; updating, by the processor, an entry in the first database corresponding to the first data value to the second data value. The method of claim 8.

11. 1. A system for reconciling data stored on heterogeneous data storage devices, comprising: The system includes a memory and a processor coupled to the memory; The processor is retrieving one or more first data files containing a first data set stored in a first database; retrieving one or more second data files containing a second data set stored in a second database; identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The processor is identifying one or more differences between the first data subset and the second data subset in the data repository; and displaying the one or more differences; Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; system.

12. the first identifier value is a combination of two or more data elements stored in each row of the one or more first rows, and the second identifier value is a combination of two or more data elements stored in each row of the one or more second rows; The system of claim 11.

13. the one or more differences include inconsistencies or missing values; The system of claim 11.

14. the processor is configured to generate a visual indicator for each of the one or more differences; Certain types of visual indicators correspond to certain types of differences, The system of claim 11.

15. retrieving the one or more first data files includes interfacing with an application program interface (API) of a type corresponding to the type of first data set; The system of claim 11.

16. the first dataset comprises clinical trial data and the second dataset comprises safety data; The system of claim 11.

17. the processor is further configured to perform an operation on the first database or the second database that resolves at least one difference among the one or more differences. The system of claim 11.

18. at least one difference of the one or more differences is a mismatch between a first data value in a first row of the one or more first rows and a second data value in a second row of the one or more second rows; the first and second data values ​​correspond to the same column of the set of columns; The system of claim 11.

19. The processor is determining a remainder of the data values ​​in the one or more first rows having a precision level exceeding a threshold amount; determining that the second data value is incorrect based on a remaining accuracy level of data values ​​in the one or more first rows that exceeds the threshold amount; updating an entry in the second database corresponding to the second data value to the first data value.

20. The system of claim 18.

20. The processor is determining a remainder of the data values ​​in the one or more second rows having a precision level exceeding a threshold amount; determining that the first data value is incorrect based on a remaining precision level of data values ​​in the one or more second rows that exceeds the threshold amount; updating an entry in the first database corresponding to the first data value to the second data value.

20. The system of claim 18.

21. A non-transitory computer-readable medium having instructions stored thereon, the execution of the instructions by one or more processors of a device causing the one or more processors to perform the following operations: The above calculation is retrieving one or more first data files containing a first data set stored in a first database; retrieving one or more second data files containing a second data set stored in a second database; identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The above calculation is identifying one or more differences between the first data subset and the second data subset in the data repository; and displaying the one or more differences; Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; Non-transitory computer-readable medium.

22. the first identifier value is a combination of two or more data elements stored in each row of the one or more first rows, and the second identifier value is a combination of two or more data elements stored in each row of the one or more second rows; 22. The non-transitory computer-readable medium of claim 21.

23. the one or more differences include inconsistencies or missing values; 22. The non-transitory computer-readable medium of claim 21.

24. the operation further comprising a visual indicator for each of the one or more differences; Certain types of visual indicators correspond to certain types of differences, 22. The non-transitory computer-readable medium of claim 21.

25. retrieving the one or more first data files includes interfacing with an application program interface (API) of a type corresponding to the type of first data set; 22. The non-transitory computer-readable medium of claim 21.

26. the first dataset comprises clinical trial data and the second dataset comprises safety data; 22. The non-transitory computer-readable medium of claim 21.

27. the operations are further configured to perform an operation on the first database or the second database that resolves at least one difference of the one or more differences.

22. The non-transitory computer-readable medium of claim 21.

28. at least one difference of the one or more differences is a mismatch between a first data value in a first row of the one or more first rows and a second data value in a second row of the one or more second rows; the first and second data values ​​correspond to the same column of the set of columns; 22. The non-transitory computer-readable medium of claim 21.

29. The above calculation is determining a remainder of the data values ​​in the one or more first rows having a precision level exceeding a threshold amount; determining that the second data value is incorrect based on a remaining accuracy level of data values ​​in the one or more first rows that exceeds the threshold amount; updating an entry in the second database corresponding to the second data value to the first data value.

30. The non-transitory computer-readable medium of claim 28.

30. The above calculation is determining a remainder of the data values ​​in the one or more second rows having a precision level exceeding a threshold amount; determining that the first data value is incorrect based on a remaining precision level of data values ​​in the one or more second rows that exceeds the threshold amount; updating an entry in the first database corresponding to the first data value to the second data value.

30. The non-transitory computer-readable medium of claim 28.

31. 1. A method for reconciling data stored on heterogeneous data storage devices, the method comprising: retrieving, by a processor, one or more first data files containing a first data set stored in a clinical database; retrieving, by the processor, one or more second data files stored in a safety database, the second data files including a second set of data; identifying, by the processor, one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading, by the processor, a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The above method is identifying, by the processor, one or more differences between the first data subset and the second data subset in the data repository; causing the processor to display the one or more differences; Identifying the one or more differences includes: calculating, by the processor, a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; and matching, by the processor, a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows. method.

32. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 32. The method of claim 31.

33. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 32. The method of claim 31.

34. 1. A system for reconciling data stored on heterogeneous data storage devices, comprising: The system includes a memory and a processor coupled to the memory; The processor is retrieving one or more first data files containing a first data set stored in a clinical database; retrieving one or more second data files containing a second data set stored in the safety database; identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The processor is identifying one or more differences between the first data subset and the second data subset in the data repository; and displaying the one or more differences; Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; system.

35. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 35. The system of claim 34.

36. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 35. The system of claim 34.

37. A non-transitory computer-readable medium having instructions stored thereon, the execution of the instructions by one or more processors of a device causing the one or more processors to perform the following operations: The above calculation is retrieving one or more first data files containing a first data set stored in a clinical database; retrieving one or more second data files containing a second data set stored in the safety database; identifying one or more first columns in the one or more first data files and one or more second columns in the one or more second data files; loading a first data subset of the first data set and a second data subset of the second data set into a data repository; the first data subset corresponds to the one or more first columns and includes one or more first rows, the second data subset corresponds to the one or more second columns and includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data subsets; The above calculation is identifying one or more differences between the first data subset and the second data subset in the data repository; and displaying the one or more differences; Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; Non-transitory computer-readable medium.

38. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 38. The non-transitory computer-readable medium of claim 37.

39. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 38. The non-transitory computer-readable medium of claim 37.

40. 1. A method of generating an output indicative of differences in data stored in disparate data storage devices, the method comprising: loading, by a processor, into a data repository a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The above method is identifying, by the processor, one or more differences between the first data set and the second data set in the data repository; generating, by the processor, an output including the first data set and the second data set and a visual indicator of each of the one or more differences; causing the processor to display the output; Identifying the one or more differences includes: calculating, by the processor, a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; and matching, by the processor, a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows. method.

41. The output is a graphical user interface (GUI) or a file.

41. The method of claim 40.

42. the one or more types of differences include inconsistent data values ​​or missing data values; 41. The method of claim 40.

43. determining, by the processor, a type of visual indicator to include in the output for each difference of the one or more differences based on one or more combinations of the type of each difference, a column type, and whether the each difference corresponds to the first or second database.

43. The method of claim 42.

44. The type of visual indicator includes one or more of a color, a tactile output, a pattern, a gradient, an animation, a shape, or a visual effect; 41. The method of claim 40.

45. the output includes the one or more first rows and the one or more second rows.

41. The method of claim 40.

46. 1. A system for generating an output file showing differences in heterogeneous data storage devices, comprising: The system includes a memory and a processor coupled to the memory; the processor is configured to execute loading a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The processor is identifying one or more differences between the first data set and the second data set in the data repository; generating an output including the first data set and the second data set and a visual indicator of each of the one or more differences; and displaying the output. Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; system.

47. The output is a graphical user interface (GUI) or a file.

47. The system of claim 46.

48. the one or more types of differences include inconsistent data values ​​or missing data values; 47. The system of claim 46.

49. the processor is further configured to determine a type of visual indicator to include in the output for each difference of the one or more differences based on one or more combinations of the type of each difference, a column type, and whether the each difference corresponds to the first or second database.

49. The system of claim 48.

50. The type of visual indicator includes one or more of a color, a tactile output, a pattern, a gradient, an animation, a shape, or a visual effect; 47. The system of claim 46.

51. the output includes the one or more first rows and the one or more second rows.

47. The system of claim 46.

52. A non-transitory computer-readable medium having instructions stored thereon, the execution of the instructions by one or more processors of a device causing the one or more processors to perform the following operations: the operations include loading a first data set corresponding to one or more first columns of a first database and a second data set corresponding to one or more second columns of a second database into a data repository; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The above calculation is identifying one or more differences between the first data set and the second data set in the data repository; generating an output including the first data set and the second data set and a visual indicator of each of the one or more differences; and displaying the output. Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; Non-transitory computer-readable medium.

53. The output is a graphical user interface (GUI) or a file.

53. The non-transitory computer-readable medium of claim 52.

54. the one or more types of differences include inconsistent data values ​​or missing data values; 53. The non-transitory computer-readable medium of claim 52.

55. the computing further includes determining a type of visual indicator to include in the output for each difference of the one or more differences based on one or more combinations of the type of each difference, a column type, and whether the each difference corresponds to the first or second database.

55. The non-transitory computer-readable medium of claim 54.

56. The type of visual indicator includes one or more of a color, a tactile output, a pattern, a gradient, an animation, a shape, or a visual effect; 53. The non-transitory computer-readable medium of claim 52.

57. the output includes the one or more first rows and the one or more second rows.

53. The non-transitory computer-readable medium of claim 52.

58. 1. A method for generating an output indicative of differences in data stored in disparate data stores, the method comprising: loading, by a processor, a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The above method is identifying, by the processor, one or more differences between the first data set and the second data set in the data repository; generating, by the processor, an output including the first data set and the second data set and a visual indicator of each of the one or more differences; causing the processor to display the output; Identifying the one or more differences includes: calculating, by the processor, a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; and matching, by the processor, a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows. method.

59. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 59. The method of claim 58.

60. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 59. The method of claim 58.

61. 1. A system for generating an output file showing differences in heterogeneous data storage devices, comprising: The system includes a memory and a processor coupled to the memory; the processor is configured to execute loading a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The processor is identifying one or more differences between the first data set and the second data set in the data repository; generating an output including the first data set and the second data set and a visual indicator of each of the one or more differences; and displaying the output. Identifying the one or more differences includes: calculating, by the processor, a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; and matching, by the processor, a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows. system.

62. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 62. The system of claim 61.

63. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 62. The system of claim 61.

64. A non-transitory computer-readable medium having instructions stored thereon, the execution of the instructions by one or more processors of a device causing the one or more processors to perform the following operations: the operation includes loading a first data set corresponding to one or more first columns of a clinical database and a second data set corresponding to one or more second columns of a safety database into a data repository; the first data set includes one or more first rows, the second data set includes one or more second rows, and the data repository includes a set of columns corresponding to the first and second data sets; The above calculation is identifying one or more differences between the first data set and the second data set in the data repository; generating an output including the first data set and the second data set and a visual indicator of each of the one or more differences; and displaying the output. Identifying the one or more differences includes: calculating a correlation of each row of the one or more first rows with each row of the one or more second rows based on a comparison of a first identifier value stored in each row of the one or more first rows with a second identifier value stored in each row of the one or more second rows; matching a first data value stored in each row of the one or more first rows against a second data value stored in a correlated row of the one or more second rows; Non-transitory computer-readable medium.

65. the first data set is associated with clinical trial data and the second data set is associated with drug safety data; 65. The non-transitory computer-readable medium of claim 52 or 64.

66. the first data set and the second data set include one or more common data values ​​associated with pharmacovigilance (PV); 65. The non-transitory computer-readable medium of claim 52 or 64.

Citation Information

Patent Citations

  • Setting reflection program, setting reflection method and setting reflection device

    JP2014137650A

  • Automated method of generating reconciliation reports regarding mismatches of clinical data received from multiple sources during a clinical trial

    US20120246149A1

  • Medical information processing device

    WO2007116899A1