A method of amalgamating data in a plurality of datasets from a plurality of heterogeneous data sources and a system thereof

The system automates the integration of heterogeneous datasets by determining data authenticity, accuracy, and completeness, addressing inefficiencies in existing technologies and ensuring traceable and auditable data processes, thereby improving operational efficiency and reducing costs.

WO2026024229A1PCT designated stage Publication Date: 2026-01-29DLIGENCE GLOBAL HOLDINGS PTE LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/SG2025/050496
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-22
Filing Date
2025-07-22
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing data integration technologies struggle with integrating heterogeneous datasets from multiple sources due to non-standardized formats, requiring substantial human intervention for authenticity, accuracy, and completeness verification, and lack automated methods for attribute-level lineage and provenance tracking, leading to operational inefficiencies and high costs.

Method used

A system and method for automating the integration of heterogeneous datasets by determining authenticity, accuracy, and completeness of data attributes, maintaining attribute-level lineage and provenance, and generating master domain data without human intervention, using a processor to ingest, standardize, and resolve entities across domains and subdomains.

Benefits of technology

Enables efficient, automated data integration with traceable and auditable processes, reducing operational costs and enhancing data quality by identifying the most accurate attributes and maintaining comprehensive lineage and provenance across the data lifecycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050496_29012026_PF_FP_ABST
    Figure SG2025050496_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A method of amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data is provided. Method includes ingesting the plurality of datasets from the plurality of heterogeneous data sources, such that the plurality of datasets comprises records with attributes; identifying changed attributes in the plurality of datasets, such that the changed attributes comprise changes to existing attributes and / or new records; standardizing the changed attributes to aggregate or decompose the standardized changed attributes into a plurality of domains and subdomains; resolving one or more entities across the plurality of heterogeneous data sources within its corresponding domains and subdomains; clustering the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes; determining the most accurate overlapping attribute of the overlapping attributes to be the master attribute; and generating master domain data comprising a plurality of master attributes sourced from the plurality of heterogeneous data sources. A system of the method is also provided.
Need to check novelty before this filing date? Find Prior Art

Description

A Method of Amalgamating Data in a Plurality of Datasets from a Plurality of Heterogeneous Data Sources and A System ThereofCross-Reference to Related Applications

[0001] The present application claims the benefit of Singapore Patent Application No. 10202402168T filed on July 22, 2024, which is incorporated by reference herein.Technical Field

[0002] The present invention relates to a method of amalgamating data in a plurality of datasets from a plurality of heterogeneous data sources and a system thereof.Background

[0003] It is often necessary for enterprises to acquire data from a plurality of data sources, including but not limited to public and private repositories. However, such data sources frequently store and disseminate information in a variety of localized formats that are heterogeneous in nature. These formats are typically non-standardized, thereby presenting significant challenges to automate data ingestion, normalisation, and integration processes.

[0004] In order to standardise and integrate heterogenous datasets comprising (i) distinct datasets with no shared attributes, e.g. entity name, address, etc. (ii) related datasets with partial schema or semantic correspondence, and (iii) overlapping datasets containing variations, duplicate or conflicting data attributes across sources, organizations are required to make substantial investments in data engineering pipelines and data operations. These pipelines typically rely on rule-based or heuristic-driven processes, which in conventional implementations require human intervention to perform one or more of the following functions: verifying the authenticity of the sourced data; prioritizing more reliable data sources over less reliable ones; and determining the completeness and accuracy of the sourced data attributes.

[0005] Existing technologies provide a general-purpose suite of tools that may be employed to construct data engineering pipelines; however, such tools are not designed to address the specific challenges associated with integrating distinct, related, and overlapping datasets asdescribed above. The available toolsets, including commercially available data engineering frameworks, are intended to serve broad use cases and are not purpose-built to accommodate domain-specific requirements or to autonomously resolve structural or semantic inconsistencies across heterogeneous data sources. Absent explicit configuration and ongoing human intervention, such tools in isolation lack the capability to conclusively establish the authenticity, accuracy, or completeness of data, thereby necessitating additional human-centric validation processes.

[0006] Although certain existing tools are capable of providing column-level data lineage within data processing pipelines, conventional tools do not support lineage tracking on a perattribute basis within a dataset when individual attributes are derived from heterogeneous data sources. Achieving such fine-grained provenance tracking — wherein the origin, transformation, and flow of individual data records are traceable across multiple stages of the pipeline — necessitates the development and integration of customized software components within the data engineering architecture. Conventional platforms do not natively support such capabilities, thereby requiring organizations to implement bespoke extensions or augmentations to existing tooling.

[0007] Data providers may provide data amalgamation services with limited support for lineage and provenance through specialized data services and offerings. However, such capabilities are realized using bespoke data engineering pipelines, manual workflows that necessitate human intervention for decision-making at multiple stages across the data lifecycle, including but not limited to data sourcing, transformation, conflict resolution, and integration. Moreover, these services are not general -purpose in nature; rather, they are custom-engineered solutions tailored to address the specific requirements and operational contexts of individual organizations. As a result, they lack portability, scalability, and applicability across diverse data environments and problem domains. Certain data providers may disclose the source of a an entire record, in limited cases, the source of individual data values; however, such disclosure typically lacks complete lineage and provenance information. Consequently, the data remains insufficiently traceable and non-auditable across its lifecycle. It is noted that data providers' systems necessitate human intervention for the generation of limited lineage and provenance information.

[0008] When datasets are acquired from a plurality of heterogeneous data sources, wherein said datasets comprise overlapping records, the resolution of conflicts arising from such overlaps typically necessitates manual intervention. In particular, the overlapping records may contain both common and source-specific attributes, rendering the consolidation process nontrivial. The synthesis of such datasets into unified records that retain a maximal and non- redundant set of attributes conventionally requires human intervention.

[0009] Furthermore, datasets obtained from such sources commonly comprise a highdimensional attribute space, often encompassing several hundred columns per record, which presents interpretability and computational challenges. The absence of automated methodologies for distinguishing, identifying, and categorizing incremental changes within these datasets into semantically meaningful events further exacerbates the complexity of managing said data.

[0010] Additionally, the traceability of amalgamated records to their respective source datasets, and the auditing of modifications effected during the integration process, remains a significant operational challenge in existing systems. Consequently, the execution of the aforementioned processes at scale — particularly over datasets numbering in the thousands to millions — requires the deployment of extensive human resources, thereby resulting in substantial operational expenditure.

[0011] The objective of this patent is to define a novel system and process to source data from a plurality of heterogeneous data sources, comprising (i) distinct datasets with partial or no shared attributes, (ii) related datasets with partial schema or semantic correspondence, and (iii) overlapping datasets containing variations, duplicates or conflicting records across sources, standardise, harmonise and master the data into data domains through an automated process that operates without human intervention. Critically, the system determines the authenticity, accuracy, and completeness of sourced data attributes, thereby automating the human-centric decision-making processes. The said system creates a framework to determine the most authentic, complete, and accurate data within overlapping datasets or attributes from a plurality of data sources. The foregoing system and processes will maintain attribute-level lineage and provenance, whereby the origin, sequence of transformation steps, modification history, andsource of each individual attribute and its corresponding values are recorded and preserved to enable traceability and auditability across the data lifecycle.Summary

[0012] According to various embodiments, a method of amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data is provided. Method includes ingesting the plurality of datasets from the plurality of heterogeneous data sources, such that the plurality of datasets comprises records with attributes; identifying changed attributes in the plurality of datasets, such that the changed attributes comprise changes to existing attributes and / or new records; standardizing the changed attributes to aggregate or decompose the standardized changed attributes into a plurality of domains and subdomains; resolving one or more entities across the plurality of heterogeneous data sources within its corresponding domains and subdomains; clustering the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes; determining the most accurate overlapping attribute of the overlapping attributes to be the master attribute; and generating master domain data comprising a plurality of master attributes sourced from the plurality of heterogeneous data sources.

[0013] According to various embodiments, the method may further include identifying one or more data domains based on the changed attributes.

[0014] According to various embodiments, the change window may include a pre-determined start and end date determined based on the temporal occurrence of the change.

[0015] According to various embodiments, the method may further include determining a completeness score of an overlapping attribute by calculating the proportion of its subcomponents with values present to the total number of requisite sub-components for the overlapping attribute, such that every sub-component is assigned weights.

[0016] According to various embodiments, the composite distance may be computed by comparing and measuring the distance of one overlapping attribute retrieved from one of the plurality of heterogeneous data sources with a corresponding overlapping attribute fromanother one of the plurality of heterogeneous data sources and dividing the distance by the completeness score of the overlapping attribute.

[0017] According to various embodiments, determining the most accurate overlapping attribute in the method may include selecting the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute.

[0018] According to various embodiments, the method may further include prioritizing the plurality of heterogeneous data sources and identifying the most authoritative change attribute by applying a priority rank order to the plurality of heterogeneous data sources from which the changed attribute was sourced.

[0019] According to various embodiments, the method may further include generating a verified delta event comprising changed attributes, such that the event encapsulates metadata describing the nature and scope of the change.

[0020] According to various embodiments, the method may further include evaluating a plurality of verified delta events to determine material changes in accordance with domainspecific rules or thresholds to generate a standardized event in order to signal the occurrence a significant real-world event.

[0021] According to various embodiments, the method may further include publishing a provenance event for every persistence or data manipulation step in the data process to capture metadata describing the origin, lineage, or transformation history of a data element, record, entity, and / or change attribute.

[0022] According to various embodiments, a system for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data is provided, the system includes a processor, a memory in communication with the processor for storing instructions executable by the processor, such that the processor is configured to: ingest the plurality of datasets from the plurality of heterogeneous data sources, such that the plurality of datasets comprises records with attributes; identify changed attributes in the plurality of datasets, such that the changed attributes comprise changes to existing attributes and / or new records; standardize the changed attributes to aggregate or decompose the standardizedchanged attributes into a plurality of domains and subdomains; resolve one or more entities across the plurality of heterogeneous data sources within its corresponding domains and subdomains; cluster the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes; determine the most accurate overlapping attribute of the overlapping attributes to be the master attribute; and generate master domain data comprising a plurality of master attributes sourced from the plurality of heterogeneous data sources.

[0023] According to various embodiments, the process may be further configured to identify one or more data domains based on the changed attributes.

[0024] According to various embodiments, the change window may include a pre-determined start and end date determined based on the temporal occurrence of the change.

[0025] According to various embodiments, the processor may be further configured determine a completeness score of an overlapping attribute by calculating the proportion of its subcomponents with values present to the total number of requisite sub-components for the overlapping attribute, such that every sub-component is assigned weights.

[0026] According to various embodiments, the composite distance may be computed by comparing and measuring the distance of one overlapping attribute retrieved from one of the plurality of heterogenous data sources with a corresponding overlapping attribute from another one of the plurality of heterogenous data sources and dividing the distance by the completeness score of the overlapping attribute.

[0027] According to various embodiments, to determine the most accurate overlapping attribute, the processor may be configured select the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute.

[0028] According to various embodiments, the processor may be further configured to prioritize the plurality of heterogeneous data sources and identifying the most authoritative change attribute by applying a priority rank order to the plurality of heterogeneous data sources from which the changed attribute was sourced.

[0029] According to various embodiments, the processor may be further configured to generate a verified delta event comprising changed attributes, wherein the event encapsulates metadata describing the nature and scope of the change.

[0030] According to various embodiments, the processor may be further configured to evaluate a plurality of verified delta events to determine material changes in accordance with domainspecific rules or thresholds to generate a standardized event in order to signal the occurrence a significant real-world event.

[0031] According to various embodiments, the processor may be further configured to publish a provenance event for every persistence or data manipulation step in the data process to capture metadata describing the origin, lineage, or transformation history of a data element, record, entity, and / or change attribute.Brief Description of Drawings

[0032] Fig. 1 shows a flow diagram of a method for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data.

[0033] Fig. 2 shows a method of sourcing a plurality of datasets from a plurality of heterogeneous data sources, normalise, harmonise, and amalgamate the datasets to master the data aligned to one or more domains and subdomains.

[0034] Fig. 3 shows a flowchart of an exemplary sub-method for registering a new data source.

[0035] Fig. 4 shows an exemplary table of a data source configuration.

[0036] Fig. 5 shows an exemplary table of a prioritisation setup.

[0037] Fig. 6 shows a flowchart of the Ingestion sub-method for extracting data employing a plurality of integration patterns.

[0038] Fig 7 shows a flowchart of the Manage Entity Data sub-method for detecting one or more changes in previously stored entity records or identifying new entity records based on data ingested from a plurality of heterogeneous data sources.

[0039] Fig. 8 shows an exemplary table representing a changed attribute identified by comparing an ingested record with a previously stored record associated with a resolved source entity.

[0040] Fig. 9 shows an exemplary table representing a newly identified record.

[0041] Fig. 10 shows a flowchart of a method of amalgamating and mastering change data ingested from a plurality of heterogeneous data sources.

[0042] Fig. 11A and Fig. 11B show exemplary tables of address standardisation and enrichment response.

[0043] Fig. 12 shows an exemplary table of exemplary unverified delta event of type.

[0044] Fig. 13 shows an exemplary table representing an output of the entity resolution step.

[0045] Fig. 14 shows an exemplary table of the computation of the completeness score.

[0046] Fig. 15 shows a schematic diagram illustrating the computation of the minimum weighted incident vertex or the smallest sum of incident edges.

[0047] Fig. 16 shows an exemplary table to determine the most accurate value amongst three overlapping addresses ingested from three authorised-heterogenous data sources.

[0048] Fig. 17A and Fig. 17B show an exemplary table of verified address change event.

[0049] Fig. 18 shows an exemplary table of a canonical event generated from a plurality of potential events.

[0050] Fig. 19 shows an exemplary table with a representative list of data domains.

[0051] Fig. 20 shows an exemplary table with a representation of the master data sourced from a plurality of data sources.

[0052] Fig 21 A-Fig. 21 C show an exemplary table of provenance events that may be generated by the system.

[0053] Fig. 22A and Fig. 22B show an exemplary table of a provenance event generated on the occurrence of a plurality of persistence or data manipulation steps in the data process

[0054] Fig. 23 shows a schematic diagram of an exemplary embodiment of a system for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data.

[0055] Fig. 24 shows an exemplary embodiment of an architecture of the system in Fig. 23.Detailed Description

[0056] Fig. 1 shows a flow diagram of a method 1000 for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data. The flow diagram depicts the method 1000 to be performed by a system 100 and the execution sequence of the method 1000. Method 1000 includes ingesting the plurality of datasets from the plurality of heterogeneous data sources in Step 1010, such that the plurality of datasets includes records with attributes; identifying changed attributes in the plurality of datasets in Step 1020, such that the changed attributes includes changes to existing attributes and / or new records; standardizing the changed attributes to aggregate or decompose the standardized changed attributes into a plurality of domains and subdomains in Step 1030; resolving one or more entities across the plurality of heterogeneous data sources within its corresponding domainsand subdomains in Step 1040; clustering the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes in Step 1050; determining the most accurate overlapping attributes of the overlapping attributes to be the master attribute in Step 1060, and generating master domain data which includes a plurality of master attributes sourced from a plurality of heterogeneous data sources in Step 1070.

[0057] Method 1000 may include categorizing the plurality of datasets by its data source. Method 1000 may include determining whether each of the plurality of datasets, i.e. a sourced datasets, already exists within in the system 100. Specifically, the system 100 is configured to determine if the sourced dataset was already saved in the system 100 as an existing dataset. If the sourced dataset is determined to exist in the system 100, the system 100 is configured to identify one or more changed attributes between the sourced dataset and the existing dataset. If the sourced dataset is determined to not exist in the system 100, the system 100 is configured to mark all the attributes of the sourced dataset as changed attributes.

[0058] Method 1000 may include identifying one or more data domains based on the changed attributes Upon identification of the one or more domains, the system 100 is configured to standardize and aggregate, and / or decompose the changed attributes into one or more corresponding domains and subdomain. Data domain is a collection of values with close affinity to a domain area pertaining to the sourced entity, e.g. legal entity, association or individual as a domain. Data domain may be seen as a category of data. A domain is a problem space that business occupies and provides solutions to this. It may include a collection of laws, regulations, systems, processes, a knowledge area or relationship context.

[0059] Method 1000 may include clustering the domain data within a change window, which includes a pre-determined start and end date, determined based on the temporal occurrence of the change. The composite distance may be computed by comparing and measuring the distance of one overlapping attribute retrieved from one of the plurality of data sources with a corresponding overlapping attribute from another one of the plurality of heterogeneous data sources and dividing the distance by the completeness score of the overlapping attribute.

[0060] To generate the master domain data, the method 1000 may include selecting the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute.

[0061] Fig. 2 shows a method 1100 of sourcing a plurality of datasets from a plurality of heterogeneous data sources, normalise, harmonise, and amalgamate the datasets to master the data aligned to one or more domains and subdomains.

[0062] Method 1100 may include an Ingestion sub-method 200 to acquire a plurality of datasets from a plurality of heterogeneous data sources and enforce a pre-defined schema contract associated with the corresponding data source.

[0063] Method 1100 may include a Manage Entity Data sub-method 300 to identify modifications, additions, or deletions to the sourced datasets relative to a prior to the initiation of the ingestion sub-method 200 and enforces a pre-defined schema contract associated with the corresponding data source.

[0064] Method 1100 may include Generating Delta Event sub-method 400 to correlate the detected changed attribute to a corresponding event type by evaluating contextual attributes for the identified event types. Sub-method 400 may include generating a delta event comprising changed attributes, and the event encapsulates metadata describing the nature and scope of the change, in alignment with its designated event schema. Sub-method 400 may include grouping a plurality of delta events from the plurality of heterogeneous data sources, published within a change window, comparing the grouped change events based on the composite distance of every changed attribute, and composing a consolidated verified delta event which includes only distinct data attributes with the minimum incident vertex.

[0065] Method 1100 may include Canonical Event sub-method 500 to evaluate a plurality of verified delta events to determine material changes to the datasets in accordance with domainspecific rules or thresholds to generate a standardized event representation conforming to a canonical schema in order to signal the occurrence a significant real-world event.

[0066] Method 1100 may include Create Master Domain Record sub-method 600 to process a plurality of verified delta events and canonical events to derive a consistent and trusted master domain record, where every event corresponds to a domain or subdomain.

[0067] Fig. 3 shows a flowchart 300F of an exemplary sub-method 150 for registering a new data source. Sub-method 100 may include at least one of the following steps: (i) determining the jurisdiction of the data source, (ii) classifying the data source within a prioritized ordering framework, (iii) determining the integration pattern to be employed to source the data, (iv) deriving and assigning a unique identifier to enable entity-level disambiguation and traceability, and (v) establishing a reference schema that serves as the interpretive contract for parsing and aligning the source's data model with the system 100. Sub-method 150 may be a precondition that must be satisfied prior to the commencement of data processing operations in method 1100.

[0068] Sub-method 150 may include a Classify Data Source step 152 to determine the type or nature of the data source by examining its characteristics, such as the jurisdiction, authority, and the authenticity of the data received from a data provider. System 100 may be configured to store the data source type and data source framework in a source master database 10.

[0069] Sub-method 100 may include a Determine Source Integration step 154 to determine the integration pattern in order to invoke the corresponding architectural component Managed Data Connector 1000.2.1, and the designated data processing architectural components, including Batch Processing 1000.2.2, Event Processing 1000.2.3, or Stream Processing 1000.2.4 as shown in Fig. 24.

[0070] Sub-method 100 may include a Register Schema step 156 to register the schema for a data source in a Schema Registry 12, which is a centralized repository that stores, manages and validates schemas. System 100 may be configured to validate the structural conformance of incoming data transmitted by the registered data source by referring to the corresponding registered schema. Any changes to the schema are recognized and registered as a new schema version for the corresponding data source.

[0071] Sub-method 100 may include a Capture Provider Entity Reference step 158 to disambiguate entities and assigns a unique identifier to the sourced data. The unique identifier subsequently enables persistent referencing and entity resolution.

[0072] Fig. 4 shows an exemplary table 400T of a data source configuration. Table 400T may be configured to store essential setup data to determine (i) jurisdiction of the data source, (ii) the classification of a data source within a prioritised ordering framework, (iii) the integration pattern employ the corresponding integration pattern, (iv) the frequency at which the data provider will transmit new or updated data, and (v) the unique identifier to enable entity resolution.

[0073] Fig. 5 shows an exemplary table 500T of a prioritisation setup. Table 500T may be configured to reconcile conflicting data values originating from a plurality of data sources as and when they arise. For every sourced attribute, the prioritisation framework determines the authenticity designated to each data source relative to other sources based on predefined criteria by a distinct rank order.

[0074] Fig. 6 shows a flowchart 600F of the Ingestion sub-method 200 for extracting data employing a plurality of integration patterns, including but not restricted to (i) web scraping (ii) file download (iii) Application Programming Interface (API), and (iv) Message Streams. Sub-method 200 may subsequently validate the sourced data against the registered scheme associated with the corresponding data source. Sub-method 200 may be configured to trigger the corresponding data extraction step based on the predetermined integration pattern setup for a data source. Sub-method 200 may invoke at least one of step 202 to extract data from a predetermined web-based content, step 204 to extract a file at a predetermined location, step 206 to extract data by calling a predetermined API, and step 208 to receive a message propagated through a continuous event stream or message queue originating from a predetermined data source.

[0075] Referring to Fig. 6, the method 1100 may include a sub-method 210 to validate the sourced data against its corresponding registered schema associated with a predetermined data source. Sub-method 210 may invoke step 212 to retrieve the registered schema and verify whether the data structures of the sourced data satisfy the structural and semantic constraints defined by the applicable registered schema retrieved from the Schema Registry 12. In step214, if the system 100 is notified of any schema modifications to the schema associated with a data source, it may trigger in response a schema evolution procedure to transition the corresponding schema representation and maintain compatibility with the modified data structure while ensuring no loss of data. In step 216, if any misalignment is detected when comparing the data against its corresponding registered schema, a failure message may be published to a dead letter queue 20.

[0076] Fig 7 shows a flowchart 700F of the Manage Entity Data sub-method 300 for detecting one or more changes in previously stored entity records or identifying new entity records based on data ingested from a plurality of heterogeneous data sources, and for storing a resulting change record reflective of such identified changes, i.e. changed attributes, or new entities, new records. Sub-method 300 may invoke step 302 to persist the ingested data in its original form in a Persistent Store 30, thereby preserving the data's native attributes and structure for downstream processing or auditability. Sub-method 300 may subsequently invoke step 304 to determine whether the ingested data corresponds to a previously stored record in a Delta Store 32. To accomplish this, the sub-method 300 employs the unique identifier associated with the data source to resolve the source entities from the ingested records. Sub-method 300 may perform a lookup operation by matching one or more unique identifiers from the ingested records to the corresponding records in the Delta Store 32. If the method is unable to resolve the source entities to one or more previously stored records, it is ascertained that one or more ingested records are new and determines the corresponding registered schema by looking up a Schema Registry 12. Upon identification of the new record, the sub-method 300 may invoke step 312 to insert the new record into the Delta Store 32 in conformity with the corresponding registered schema. Upon resolving the source entities against one or more records in the Delta Store 32, it is ascertained one or more source entities exist Sub-method 300 may then execute an attribute-level comparison between the ingested records where source entities exist and the corresponding records in the Delta Store 32. Sub-method may invoke step 306 to determine the corresponding registered schema associated with the identified change attributes. Thereafter, the sub-method 300 may invoke step 310 to insert the changed attributes as a new version of the previously stored data record into the Delta Store 32 in conformity with the corresponding registered schema.

[0077] Fig. 8 shows an exemplary table 800T representing a changed attribute identified by comparing an ingested record with a previously stored record associated with a resolved sourceentity. The source entity was resolved by determining the unique identifier for the data source ACRA to be UEN. The UEN ‘ 196300306G’ from the ingested record was used to lookup previously stored records in the Delta Store 32. The previously stored record is herein referred to as the ‘Current record’ in the table. Each attribute of the ingested data record, herein referred to as the ‘New record’ in the table 800T, is systematically compared against corresponding attributes of the previously stored data record. Upon comparison, the registeredAddress:levelNo is identified to be the changed attribute, as it was modified from the value 26 in the previously stored record to ‘28’ in the ingested record.

[0078] Fig. 9 shows an exemplary table 900T representing a newly identified record, wherein the identification is based on matching a unique identifier from the ingested record with previously stored records in the Delta Store 32, and wherein no corresponding previously stored record associated with the unique identifier is found within the Delta Store 32. The source entity was resolved by determining the unique identifier for the data source GLEIF to be LEI. Upon looking up the LEI ‘254900OXRICREAQ3VU73 from the ingested record, herein referred to as New Record in the illustration,’ no corresponding record is found in the Delta Store 32.

[0079] Fig 10 shows a flowchart WOOF of a method of amalgamating and mastering change data ingested from a plurality of heterogeneous data sources. Method may include (i) Generate Delta Event sub-method 400 to produce a signal indicative of attribute-level changes between ingested and previously stored records or a new ingested record associated with a resolved entity that lacks any previously stored representation, the generated verified delta event is representative of a domain or a sub domain and comprises of the attributes with a highest probability of accuracy and completeness (ii) Generate Canonical Event sub-method 500 to correlate a plurality of associated verified delta events to generate a canonical event in order to signal the occurrence of a significant real world event (iii) Create Master Domain Record submethod 600 to create a unified authoritative record that represents the canonical state of an entity, derived from one or more contributing canonical events.

[0080] Sub-method 400 may invoke the step 402 to decompose a newly ingested record or group the associated changed attributes to its corresponding domain and its constituent subdomains, wherein every sub-domain corresponds to a single, predefined change event type.This step is performed by looking up the changed attributes against the registered delta domain schema for the corresponding event type in the Schema Registry 12

[0081] Fig. 11A and Fig. 11B show exemplary tables HOOT of address standardisation and validation response. Upon determining the corresponding delta domain schema model, the submethod 400 may invoke the step 406 to enrich attributes to augment the domain data with additional contextual data in conformance with the registered schema. The enriched contextual data may include code lists, taxonomies, lookup tables, or authoritative datasets that define the permissible values, formats, or relationships for specific data attributes within a given domain. The enriched data may be derived from pre-determined internal or external sources. For example, based on a given address, the latitude and longitude may be determined using geocoding APIs as shown in Fig. 1 IB. In this example, the addition of the latitude and longitude is considered an enrichment step.

[0082] Sub-method 400 may subsequently invoke step 408 to standardise the data, i.e. changed attribute, which may be configured to normalize the format, structure, or representation of changed attributes within an ingested record, wherein the standardization may include steps, but not restricted to, transforming values into a predefined schema, applying formatting rules and resolving inconsistencies in conformance to the registered schema. For example, an address is decomposed into its sub-domain that include Unit Number, Building Name, Block Number, Street Name, Locality, City, Postal Code, and Country as shown in Fig. 11.

[0083] Upon enriching and standardising the data in conformance with the delta domain schema model, the sub-method 400 may invoke step 410 to generate an unverified delta event and publish the generated event to an Event Store 14, designated to store all historical unverified events. The generated unverified delta event is published to Delta Event Stream 16, designated to transmit a sequence of unverified events in the order in which it was generated. Sub-method 400 may include aggregating the standardized changed attributes into a plurality of domains and subdomains.

[0084] Fig. 12 shows an exemplary table 1200T of exemplary unverified delta event of type entity. address delta unverified message. The depicted event comprises an event header and event data. The event header contains metadata describing the event itself, including but not limited to identifiers, timestamps, event type, source, versioning information, and schemareferences. The event data typically includes one or more domain-specific attributes that describe the action, state change, or context of the event.

[0085] Upon publishing an event, the sub-method 400 may invoke step 412 to resolve entities, which may be configured to associate the identified real-world entity within one or more generated unverified delta events, ingested from a plurality of heterogeneous data sources, with a unique master entity identifier. When a common unique identifier is present across a plurality of heterogeneous data sources, the sub-method 400 may employ the identifier as a key to associate a plurality of unverified delta events with the corresponding real-world entity. When such a common unique identifier does not exist, a prioritised subset of attributes is extracted from the unverified delta events that are relevant and informative for determining whether two or more unverified delta events refer to the same real-world entity, these attributes are determined to be the comparing attributes. Sub-method 400 may employ a predetermined matching decision model that compares and matches the comparing attributes to associate two or more unverified delta events with a real-world entity. Sub-method 400 may generate a unique master entity identifier to associate all the unverified delta events with a real-world entity.

[0086] Fig. 13 shows an exemplary table 1300T representing an output of the entity resolution step, wherein three unverified events, each associated with a separate unique identifier, are determined to correspond to the same real-world entity and are accordingly linked via a common master entity identifier.

[0087] Upon generating a master entity identifier, the sub-method 400 may cluster the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes. Sub-method 400 may invoke step 414 to cluster or group a plurality of unverified delta events, which comprises standardized changed attributes, within their corresponding domain and sub-domain occurring within an active event window for each of the one or more resolved entities, having an identical event type and linked to a unique master entity identifier, to determine overlapping attributes. The active event window is determined by the data refresh frequency of the originating data source and historical data evolution thresholds. Sub-method 400 may determine the thresholds by continuously evaluating previously generated unverified delta events to determine the temporal window, e.g.60 days, within which unverified delta events of a specific event type, generated from heterogeneous data sources, are considered to exhibit similarity or deviation values beyond which the unverified delta events of a specific event type, generated from heterogeneous data sources are considered to exhibit dissimilarity. Two or more unverified events are considered similar when they exhibit a sufficient degree of correspondence across one or more selected attributes, such that the events can be inferred to pertain to the same real-world occurrence, entity, or process.

[0088] Sub-method 400 may invoke step 416 to measure the completeness, i.e. determine a completeness score, of every attribute in the generated unverified delta event, wherein the attribute, domain, or subdomain may be further decomposed into its constituent subcomponents, e.g. city or country of an address. Sub-method 400 determines the completeness score by calculating the proportion of sub-components with values present to the total number of requisite sub-components for that overlapping attribute, domain, or subdomain. Every subcomponent may be assigned weights to quantify the criticality of the values.

[0089] Fig. 14 shows an exemplary table 1400T of the computation of the completeness score, wherein the first address is complete with all the components of the address, the second address is missing the building name, and the third address is missing the building name and the floor level, as a consequence, the method calculates a proportionate measure of completeness based on the available data values.

[0090] Sub-method 400 may invoke step 418 to calculate the centrality of overlapping attributes. The centrality of attributes is calculated by comparing overlapping attribute values associated with the clustered unverified delta events. Overlapping attributes are determined where the attribute names, types, or semantic meanings align such that the attributes can be used for comparison. The centrality is computed by measuring the distance between three or more overlapping attributes. The technique employed to compute the distance may vary based on the data type, which may include:String: Calculate distance using matching techniques such as KNN, N-gram,Euclidean, and Levenshtein that help determine a distance between two strings The suitability of the selected technique is determined based on the specific properties of the string being processed.Date: Calculate the distance in days.Time: Calculate distance in hours, minutes, and or seconds.Numeric: Calculate the distance in percentages, value assessed divided by value compared minus I .Percentages: Calculate the distance by subtracting the value assessed from the value compared.Currency: Calculate the distance based on equivalent currency differences, where an equivalent currency value is not determinable, identify and match based on an equivalent base currency, such as USD.Codes and unique IDs: Codes and ID columns require an exact match irrespective of the data type; a match returns a distance of 0, and a no match returns a 1 .Binary: exact match returns a 0, and no match returns a 1.

[0091] Sub-method 400 eliminates divergent values based on distance thresholds. The distance threshold is a predetermined tolerance boundary such that, when the distance between the values corresponding to two overlapping attributes exceeds the boundary, the values are considered dissimilar. Sub-method 400 may compute the composite distance between two overlapping attributes, where:DistanceComposite Distance — - - -Completeness

[0092] Sub-method 400 may determine the most accurate overlapping attribute by selecting the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute. Sub-method 400 may determine the minimum weighted incident vertex or 5W(G) by computing the smallest sum of incident edges, wherein every overlapping attribute value obtained from a trusted or authoritative data source that is recognized as a reliable source of truth within the relevant domain is represented as a vertex v E V, and the composite distance represents the weighted incident edges w(e . Sub-method 400 may eliminate the divergent or dissimilar overlapping attribute values where w(e) is greater than a predetermined distance threshold T. The weighted incident edge is determined by computing the sum of incident edges e~v w(e . Where three or more overlapping attribute values exist, the minimum incident vertex is computed by determining the minimum of all the computed weighted degrees:8w(G), the minimum incident vertex is the most central, as a result, the most accurate of all overlapping attributes.

[0093] Fig. 15 shows a schematic diagram illustrating the computation of the minimum weighted incident vertex or the smallest sum of incident edges. In the illustration, the vertices v e V are A, B, C, and D, each representing data values of overlapping attributes ingested from four authorised and heterogeneous data sources. Sub-method 400 may compute the composite distance to determine the weighted incident edges w(el, AB, AC, BC, BD, CA, CB, DA, DB, and DC. Sub-method 400 may compare the weighted incident edges w(ej with the distance threshold; vertex D is eliminated as it is greater than the predetermined distance threshold T. Sub-method may compute the sum of incident edgesSub-method may compute to determine the minimum weighted incidentvertex.

[0094] Fig. 16 shows an exemplary table 1600T to determine the most accurate value amongst three overlapping addresses ingested from three authorised-heterogenous data sources: ACRA, a business registry; GLEIF, a global issuer of unique identifiers of unique IDs for legal entities; and SGX, a stock exchange. Sub-method 400 may compute the minimum incident vertex of the identified overlapping attributes of a given entity. GLEIF. Address of value “6 Raffles Quay #25-01 048580” denoted as A, ACRA.Address of value “6 RAFFLES QUAY #25-01 JOHN HANCOCK TOWER 048580 ” denoted as B and SGX. Address of value ”6 Raffles Quay #25- 01 Singapore 048580” denoted as C. All three addresses are determined to have variations wherein C is missing the component Building Name, and .4 is missing Building Name and City. As a result, the Completeness is determined to be 49% and 67% for A and C, respectively. B is determined to have a higher Completeness of 82% as the address included the Building Nameand City. Given the addresses are strings, a predetermined technique, including but not limited to k-nearest neighbours (KNN), is employed to perform address matching and comparison in pairs to arrive at the distance.The following is the computation of the sum of incident edges £e~v w(e):Ae= 4 36 + 49% I 3.16 + 49% = 15.31Be= 4.36 + 82%+ 5.36 + 82% = 11.82Ce= 5.34 + 67% + 3.16 + 67% = 12.82Sub-method may determine the minimum weighted incident vertex min( % Bc, Cc) to be: vEV min(15.31, 11.82, 12.82) = 11.82 vEVSub-method may determine ACBA.Address or ,9cto be the most central of the three values, as a result is the most accurate of all overlapping attributes and therefore is a master attribute.

[0095] Sub-method 400 may invoke step 420 to compose and publish a verified delta event to the Event Store 14 and Delta Event Stream 16. The verified delta event is generated to include the attributes determined to be the minimum weighted incident vertex in conformance with the delta domain schema model.

[0096] Sub-method may mark every published attribute as ‘Certified’ or ‘Uncertified’ in conformance with. An attribute may be marked as ‘Certified’ only when the value is determined to be a near equivalent from two or more authorised heterogeneous data sources, where the distance is less than the predetermined distance threshold.

[0097] In cases where three or more values amongst overlapping attributes exist, the weighted incident vertex is determined to be the most central as a result, the most accurate amongst the three, and is marked as ‘Certified’ .

[0098] In cases where only two values amongst overlapping attributes exist, the comparison of the two value is determined to be near equivalent, wherein the distance is within the distancethreshold, a priority-based rank order is applied, as illustrated in the accompanying Fig. 5, to determine the precedence of the attribute ingested from the more authoritative data source and the attribute is marked as "Certified .

[0099] In cases where only two values amongst overlapping attributes exist, the comparison of the two values is determined to be divergent, wherein the distance is greater than the distance threshold, the sub-method evaluates if the value is sourced from the authoritative data source or priority rank order 1, and the attribute is marked as "Uncertified . System 100 may be configured to prioritize the plurality of data sources and identifying the most authoritative change attribute by applying a priority rank order to the plurality of heterogeneous data sources from which the changed attribute was sourced

[0100] In cases where only one value amongst overlapping attributes exists, the sub-method 400 may evaluate if the value is sourced from the authoritative data source or priority rank order 1, and the attribute is marked as "Uncertifie .

[0101] Fig. 17A and Fig. 17B show an exemplary table 1700T of verified address change event. Sub-method 400 may publish verified events to two destinations: (i) an event store 14, which acts as a durable repository for raw or structured event data to support auditing, historical analysis, or replay mechanisms, and (ii) a delta event stream 16, which comprises a standardized stream of normalized events that adhere to a unified schema format. The verified event stream enables the system 100 to consume events in a consistent format regardless of the original source or payload variations. In certain embodiments, schema evolution mechanisms may be applied to ensure backward compatibility of events published to the canonical stream Fig. 17A and 17B are illustrations of a verified address change event.

[0102] Fig. 18 shows an exemplary table 1800T of a canonical event generated from a plurality of potential events. Sub-method 500 may invoke step 502 to determine the materiality of a plurality of verified delta events published. The materiality is determined based on predetermined materiality rules. Plurality of potential events may occur under certain conditions for a given entity ‘00437f7695b89487el2da690a31e23d5 a real-world legal entity named ‘SHIN-MO JI PTE. LTD. ’ registrationStatus is updated to ‘AML - AMALGAMATED, and name is changed to ‘TATEYAMA PTE. LTD. ’. Sub-method 500 may determine that the two events are related and linked based on the materiality rule; two verified delta events publishedfor a given entity of unique master entity identifier ‘00437 695b89487el2da690a3Ie23d5 ’ are of event type ‘entity. name. delta.verified’ and ‘entity.incorporation.delta.verifted’ wherein registrationStatus was updated from ‘ACT - ACTIVE ’ to ‘AML - AMALGAMATED ' within a threshold horizon of 60 days Sub-method 500 may classify both verified delta events as material.

[0103] Sub-method 500 may invoke step 504 to enrich additional attributes wherein one or more attributes of a data record are supplemented, derived, or modified based on corresponding attribute definitions retrieved from the Schema Registry 12. Schema Registry 12 includes standardized metadata, permissible value formats, and enrichment rules that govern the structure, semantics, and augmentation logic for the attributes. The enrichment may include populating missing values, applying default values, deriving calculated fields, or aligning attribute formats with canonical definitions.

[0104] Sub-method 500 may invoke step 506 to validate the event payload or data. This validation step enforces conformance to the schema definition in Schema Registry 12 associated with the corresponding event type. Schema Registry 12 defines the expected structure, required fields, data types, and permissible value constraints for each event type. Upon receiving the event, the system 100 determines the applicable schema by identifying the event type or classification and retrieving the corresponding schema definition from the Schema Registry 12. The event is then validated against the retrieved schema to ensure conformity. This includes verifying the presence of mandatory attributes, compliance with defined data types, value format constraints, and any domain-specific validation logic defined within the schema metadata. Any validation failures are published as an error message to the dead letter queue.

[0105] Upon successful validation, the sub-method 500 proceeds to a publication step. In this step, the validated canonical event is published to two destinations: (i) an event store 14, which acts as a durable repository for raw or structured event data to support auditing, historical analysis, or replay mechanisms; and (ii) a canonical event stream 18, which comprises a standardized stream of normalized events that adhere to a unified schema format. The canonical event stream 18 enables the system 100 to consume events in a consistent format regardless of the original source or payload variations. In certain embodiments, schema evolutionmechanisms may be applied to ensure backward compatibility of events published to the canonical stream.

[0106] Fig. 19 shows an exemplary table 1900T with a representative list of data domains, which is provided for illustrative purposes and is not intended to be exhaustive. Fig. 20 shows an exemplary table 2000T with a representation of the master data sourced from a plurality of data sources. Sub-method 600 may process canonical events, and verified events are ingested from respective event streams 17,18 or event store 14,15 to distinct master domain data. Each ingested event is classified based on its domain type (e.g., legal entity, listing arrangements, associated parties) and is evaluated to determine whether it pertains to an existing master record or a new candidate. A correlation process is performed to identify matching or overlapping entity identifiers, reference attributes, or semantic keys across the events. All events are processed in the exact sequence in which it was generated, with one event overwriting another. The finalized master domain record is persisted in a master data store and optionally published to downstream consumers or services that rely on high-fidelity entity representations.

[0107] Fig 21A-Fig. 21C show an exemplary table 2100T of provenance events that are generated by the system 100. Method 1000 may include publishing a provenance event for every persistence or data manipulation step in the data process to capture metadata describing the origin, lineage, or transformation history of a data element, record, entity, and / or change attribute. A provenance event may include information such as the source, timestamp of creation or modification, the identity of the entity or process that resulted in altered the dataset, and the nature of the operation performed. Provenance events are used to enable auditability, traceability, and integrity validation within the system.

[0108] Fig. 22A and Fig. 22B show an exemplary table 2200T of a provenance event generated on the occurrence of a plurality of persistence or data manipulation steps in the data process, including but not limited to, step 306 to store raw ingested data, step 419 to publish an unverified event, step 418 to calculate centrality and step 600 to create a master domain record.

[0109] Fig. 23 shows a schematic diagram of an exemplary embodiment of a system 100 for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data. System 100 includes a processor 100P, a memory 100M in communication with the processor 100P for storing instructions executable by the processor 100P, such thatthe processor 100P is configured to execute the method 1000. System 100 may include at least one of a multimedia module 100U configured to display the scanned image and the scoring input interface and receive user input, an audio module 100A configured to input / output audio signals, an input / output (I / O) interface 100N configured to provide an interface between the processor 100P and peripheral interface modules, e.g. keyboard, and a communication module 100C configured to facilitate communication between the system 100 and other devices or server. System 100 may include a storage module 100S, e.g. a database, cloud server, configured to store data.

[0110] Fig. 24 shows an exemplary embodiment of an architecture 2000 A of the system 100 for amalgamating a plurality of datasets from a plurality of heterogeneous data sources The architectural components may be configured to perform one or more of the following operations: sourcing datasets from one or more sources, standardizing the data to predetermined formats, harmonising data in domains, amalgamating, mastering, and distributing the data.

[0111] System 100 may include a Data Orchestration Module 1000.1 configured to coordinate and execute a plurality of data processing operations of system 100. Data Orchestration Module 1000.1 may be configured to manage the execution of these data processing operations in a predetermined sequence and may be configured to manage all dependencies between data processing steps, manage failures of data processing operations, and track the lineage of data.

[0112] System 100 may include an Ingestion Module 1000.2 configured to acquire data using a plurality of heterogeneous integration methods. Ingestion Module 1000.2 may be configured to employ one or many integration patterns including but not limited to batch, event or stream processing to ingest and persist the sourced dataset in its original form.

[0113] System 100 may include a Manage Data Connectors Module 1000.2.1 configured to abstract the sourced data format and integration systems from the ingestion operations. Manage Data Connectors Module 1000.2.1 may be configured to provide reusable functions or interfaces to connect to a plurality of heterogeneous integration methods, including but not limited to file-based integration, application programming interfaces (APIs), message queues, data streams, database replication with change data capture (CDC), and web scraping techniques.

[0114] System 100 may include a Batch Processing Module 1000.2.2 configured to collect and processes sourced datasets in bulk instead of row-by-row processing. Batch Processing Module 1000.2.2 may be configured to operate at predetermined intervals, schedules, or conditions.

[0115] System 100 may include an Event Processing Module 1000.2.3 configured to ingest distinct events asynchronously emitted, where each event represents the change in data, state or condition. Each event is used to trigger one or more predetermined operations.

[0116] System 100 may include a Stream Processing Module 1000.2.4 configured to continuously process sourced datasets individually without having to accumulate data into batches.

[0117] System 100 may include a Transformation Module 1000.3 configured to standardise, harmonise, and master the data into data domains through an automated process. Transformation Module 1000.3 may be configured to align schemas of sourced datasets to predetermined domain data, generate events, resolve entities, amalgamate the data, and publish the master data

[0118] System 100 may include a Data Processing Module 1000.3.1 operable to cleanse, merge, align, and standardize sourced datasets to conform with downstream schema requirements. Data Processing Module 1000.3.1 may be configured to perform data manipulation operations, including but not limited to record filtering, dataset subsetting, table joins, entity merging, and attribute standardization.

[0119] System 100 may include an Entity Resolution Module 1000.3.2 configured to match and consolidate data pertaining to the same real-world entity. Entity Resolution Module 1000.3.2 may involve a number of techniques, including but not limited to deterministic rules, probabilistic models, machine learning algorithms, and graphing to evaluate the similarity of records and determine whether they represent the same entity across a plurality of datasets.

[0120] System 100 may include an Amalgamation Module 1000.3.3 configured to compare distinct, related, and overlapping datasets from heterogeneous data sources, and measures todetermine the most complete, accurate, and authentic data attributes. Amalgamation Module1000.3.4 may be configured to compose the identified datasets into domains and subdomains.

[0121] System 100 may include an Event Generation Module 1000 3.4 configured to generate domain-specific events based on detected dataset changes. System 100 may be configured to monitor datasets for updates and is operable to emit domain-specific events corresponding to detected changes.

[0122] System 100 may include a Serve Module 1000.4 configured to grant access to master data generated by the system 100 to one or more consuming entities that are external or internal to the system 100. Serve Module 1000.4 may be configured to expose the data through integration methods, including but not limited to Application Programming Interfaces(API), event streams, direct structured query language (SQL) employing query federation engines, a data fabric, or a large language interface to interact with the underlying master data.

[0123] System 100 may include an Event Streams Module 1004.1 configured to generate a continuous, time-ordered sequence of events that are emitted by the system 100. Every event within an event stream includes an immutable structured data generated in real time and transmitted over a messaging infrastructure. The messaging infrastructure may be configured to organize a plurality of event streams into distinct topics, wherein each topic corresponds to a particular domain or category of events.

[0124] System 100 may include a Query Federation Module 1000.4.2 configured to interpret and transform a federated query into a plurality of source-specific sub-queries, submit the subqueries to their respective data systems, retrieve partial results, and perform result normalization, reconciliation, or merging to produce a unified response consistent with the original query intent.

[0125] System 100 may include an API Module 1000.4.3 configured to expose data and functionality to external or internal consuming entities via one or more application programming interfaces (APIs). The APIs may be defined employing communication protocols including but not limited to GraphQL to request specific data structures from a system using a strongly typed schema, Google Remote Procedure Calls (gRPC) to transmit data employingprotocol buffers for efficient serialization, and Representational State Transfer (REST) to transmit as structured data formats including but not restricted to JSON and XML.

[0126] System 100 may include a Semantic Layer Module 1000.4.4 configured to abstract the technical implementation of a domain model, by mapping technical data structured including but not limited to tables, columns, and data types, to domain-specific concepts. Semantic Layer Module 1000.4.4 enables tailored views to reflect the unique business semantics, terminologies, and interpretive models of the respective consuming entity. The semantic layer thus facilitates contextualized access to data by technical schemas into domain-aligned representations specific to the consumer’s operational requirements.

[0127] System 100 may include a LLM Interface Module 1000.4.5 configured to operate as an intermediary by receiving user input expressed as a natural language query and converting such query into machine machine-readable query. LLM Interface Module 1000.4.5 may leverage pretrained or fine-tuned large language models to convert user intents into executable query constructs, retrieve the corresponding results, and optionally present them in human-readable form.

[0128] System 100 may include a Privacy and Security Layer Module 1000.5 configured to provide centralized and granular control over data access and usage within the data platform. It is operable to enforce user or system level permissions, monitor data activity, prevent unauthorised disclosure of sensitive information and apply privacy-preserving transformations.

[0129] System 100 may include a Data Governance Module 1000.6 configured to operate as a framework that includes, but is not restricted to, metadata management, data stewardship, compliance monitoring, lineage, and provenance. In some embodiments, the Data Governance module 1000.6 may include a Unified Governance Control Plane Sub-Module operable to orchestrate governance policies and workflows, along with one or more Sub-Modules comprising a Data Catalog Sub-Module for metadata management, a Data Lineage Tracking Sub-Module for monitoring data flow and transformations, and a Data Provenance Sub-Module configured to capture the origin and historical context of data assets.

[0130] Unified Governance Control Plane Sub-Module may be configured to provide a centralized interface and execution framework for governing the lifecycle and usage of dataassets, integrating metadata management, data stewardship, auditing, and policy-driven automation.

[0131] Data Catalogue Sub-Module, which is a metadata management component, may be configured to organize, classify, and index data assets within the platform. The catalogue enables discovery, search, and semantic understanding of datasets by exposing descriptive metadata such as schema definitions, data types, source information, and associated business terms.

[0132] Data Lineage Tracking Sub-Module may be configured to track and represent the end- to-end flow of data through the platform, including its origin, intermediate transformations, and final outputs.

[0133] Data Provenance Sub-Module configured to capture the source, history, and derivation of a data asset, thereby enabling auditability, trust, and validation of data authenticity.

[0134] System 100 may further include a Data Storage Module 1000.7 that may be configured to persist, manage, and retrieve structured, semi-structured, or unstructured data in various formats. Data Storage Module 1000.7 may comprise a plurality of underlying storage technologies, including but not restricted to Lakehouse architectures, Relational Databases, Vector Databases, and Graph Databases, each optimized for distinct data models and query patterns.

[0135] System 100 may include an Object Store Module 1000.8 configured to provide a nonrelational, flat address space for storing data in the form of objects, each of which may encapsulate binary content, descriptive metadata, and access control attributes.

[0136] System 100 may include a Data Ops Module 1000.9 which includes a collection of utilities and services that facilitate the collaboration, automation of deployment, versioning, and management of data engineering pipelines. Some embodiments may include utilities that provide quality assurance frameworks, including but not restricted to observability, test automation, quality metric computation and monitoring.

[0137] The present invention relates to a system and a method for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data and a system thereof generally as herein described, with reference to and / or illustrated in the accompanying drawings.

Claims

Claim1. A method of amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data, the method comprising, ingesting the plurality of datasets from the plurality of heterogeneous data sources, wherein the plurality of datasets comprises records with attributes, identifying changed attributes in the plurality of datasets, wherein the changed attributes comprise changes to existing attributes and / or new records, standardizing the changed attributes to aggregate or decompose the standardized changed attributes into a plurality of domains and subdomains, resolving one or more entities across the plurality of heterogeneous data sources within its corresponding domains and subdomains, clustering the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes, determining the most accurate overlapping attribute of the overlapping attributes to be the master attribute, and generating master domain data comprising a plurality of master attributes sourced from the plurality of heterogeneous data sources.

2. The method according to claim 1, further comprising identifying one or more data domains based on the changed attributes.

3. The method according to claim 1 or 2, wherein the change window comprises a predetermined start and end date determined based on the temporal occurrence of the change.

4. The method according to any one of claims 1 to 3, further comprising determining a completeness score of an overlapping attribute by calculating the proportion of its subcomponents with values present to the total number of requisite sub-components for the overlapping attribute, wherein every sub-component is assigned weights.

5. The method according to claim 4, wherein the composite distance is computed by comparing and measuring the distance of one overlapping attribute retrieved from one of the plurality of heterogenous data sources with a corresponding overlapping attribute from anotherone of the plurality of heterogenous data sources and dividing the distance by the completeness score of the overlapping attribute.

6. The method according to any one of claims 1 to 5, wherein determining the most accurate overlapping attribute comprises selecting the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute.

7. The method according to any one of claims 1 to 6, further comprising prioritizing the plurality of heterogeneous data sources and identifying the most authoritative change attribute by applying a priority rank order to the plurality of heterogeneous data sources from which the changed attribute was sourced.

8. The method according to any one of claims 1 to 7, further comprising generating a verified delta event comprising changed attributes, wherein the event encapsulates metadata describing the nature and scope of the change.

9. The method according to any one of claims 1 to 8, further comprising evaluating a plurality of verified delta events to determine material changes in accordance with domainspecific rules or thresholds to generate a standardized event in order to signal the occurrence a significant real world event.

10. The method according to any one of claims 1 to 9, further comprising publishing a provenance event for every persistence or data manipulation step in the data process to capture metadata describing the origin, lineage, or transformation history of a data element, record, entity, and / or change attribute.

11. A system for amalgamating a plurality of datasets from a plurality of heterogeneous data sources to generate master domain data, the system comprising, a processor, a memory in communication with the processor for storing instructions executable by the processor, wherein the processor is configured to:ingest the plurality of datasets from the plurality of heterogeneous data sources, wherein the plurality of datasets comprises records with attributes, identify changed attributes in the plurality of datasets, wherein the changed attributes comprise changes to existing attributes and / or new records, standardize the changed attributes to aggregate or decompose the standardized changed attributes into a plurality of domains and subdomains, resolve one or more entities across the plurality of heterogeneous data sources within its corresponding domains and subdomains, cluster the standardized changed attributes within their corresponding domain and subdomain occurring within the predetermined change window for each of the one or more resolved entities to determine overlapping attributes, determine the most accurate overlapping attribute of the overlapping attributes to be the master attribute, and generate master domain data comprising a plurality of master attributes sourced from the plurality of heterogeneous data sources.

12. The system according to claim 11, wherein the process is further configured to identify one or more data domains based on the changed attributes.

13. The system according to claim 11 or 12, wherein the change window comprises a predetermined start and end date determined based on the temporal occurrence of the change.

14. The system according to any one of claims 11 to 13, wherein the processor is further configured determine a completeness score of an overlapping attribute by calculating the proportion of its sub-components with values present to the total number of requisite subcomponents for the overlapping attribute, wherein every sub-component is assigned weights.

15. The system according to claim 14, wherein the composite distance is computed by comparing and measuring the distance of one overlapping attribute retrieved from one of the plurality of heterogenous data sources with a corresponding overlapping attribute from another one of the plurality of heterogenous data sources and dividing the distance by the completeness score of the overlapping attribute.

16. The system according to any one of claims 11 to 15, wherein to determine the most accurate overlapping attribute, the processor is configured select the overlapping attributes of a resolved entity with the minimum incident vertex as the master attribute.

17. The system according to any one of claims 11 to 16, wherein the processor is further configured to prioritize the plurality of heterogeneous data sources and identifying the most authoritative change attribute by applying a priority rank order to the plurality of heterogeneous data sources from which the changed attribute was sourced.

18. The system according to any one of claims 11 to 17, wherein the processor is further configured to generate a verified delta event comprising changed attributes, wherein the event encapsulates metadata describing the nature and scope of the change.

19. The system according to any one of claims 11 to 18, wherein the processor is further configured to evaluate a plurality of verified delta events to determine material changes in accordance with domain-specific rules or thresholds to generate a standardized event in order to signal the occurrence a significant real world event.

20. The system according to any one of claims 11 to 19, wherein the processor is further configured to publish a provenance event for every persistence or data manipulation step in the data process to capture metadata describing the origin, lineage, or transformation history of a data element, record, entity, and / or change attribute.

Citation Information

Patent Citations

  • Graph based resolution of matching items in data sources

    US10268735B1

  • Techniques for application data scrubbing, reporting, and analysis

    US20090240694A1

  • System and method for data provenance management

    US20100070463A1

  • Cloud based master data management architecture

    US20120198036A1

  • Systems and methods for merging source records in accordance with survivorship rules

    US20130166552A1