Reflecting distributed data access

The distributed data access system addresses privacy and security concerns by reflecting data from local stores to a unified interface using edge nodes and volatile memory, ensuring secure, efficient, and real-time access without permanent storage outside the data store.

JP2025522378APending Publication Date: 2025-07-15COLLIBRA BELGIUM BV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024572344
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-05-30
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing data profiling and quality assurance systems face challenges in accessing large and fragmented datasets while preserving privacy and security, particularly when sensitive data is stored locally, leading to high costs, regulatory issues, and increased privacy risks.

Method used

A distributed data access system that utilizes an edge node to reflect data from local data stores to a unified user interface, ensuring data is not stored permanently outside the data store, using end-to-end encryption and volatile memory for temporary storage during sessions.

Benefits of technology

Enhances security and privacy by maintaining data ownership and control, reduces latency, and optimizes memory resources through real-time, secure data access without transferring sensitive data to third-party storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522378000001_ABST
    Figure 2025522378000001_ABST
Patent Text Reader

Abstract

The present disclosure is directed to a system and method for securely reflecting data from a local data store. In one exemplary embodiment, a user may connect a local data store to an edge node, which may facilitate communication between a third-party data management hub and the local data store. A user operating the third-party data management hub may request read / write of specific data in the (user-owned) local data store. In response to the request, data from the local data store may be generated. The response data may be sent to the data management hub via an encrypted data session, and at the data management hub the response data is reflected to the user on a user interface. The response data is stored in a volatile memory store and is not moved from the persistent memory in the local data store. To preserve the privacy / security of the underlying data, the volatile memory store is erased as soon as the encrypted data session ends.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications U.S. Patent Application No. 16 / 776,293, entitled "SYSTEMS AND METHOD OF CONTEXTUAL DATA MASKING FOR PRIVATE AND SECURE DATA LINKAGE", U.S. Patent Application No. 17 / 103,751, entitled "SYSTEMS AND METHODS FOR UNIVERSAL REFERENCE SOURCE CREATION AND ACCURATE SECURE MATCHING", U.S. Patent Application No. 17 / 103,720, entitled "SYSTEMS AND METHODS FOR DATA ENRICHMENT", and U.S. Patent Application No. 17 / 219,340, entitled "SYSTEMS AND METHODS FOR AN ON - DEMAND, SECURE, AND PREDICTIVE VALUE - ADDED DATA MARKETPLACE," are hereby incorporated by reference in their entirety.

[0002] The present disclosure relates to distributed data and privacy - by - design.

Background Art

[0003] Data profiling and data quality (DQ) are very important for entities with a large number of large datasets and data sources. However, as the scale grows, data profiling and data quality assurance across such large amounts of data can be prohibitively expensive and may lack the computing resources to utilize. Also, data profiling and data quality analysis can be time-consuming when performed in external or non-native applications. Data profiling and data quality checks are often serialized and asynchronous to the data creation process.

[0004] In one example, when an entity attempts to profile or clean up certain data, it may discover that certain sensitive data is stored locally rather than on a server accessible via the network. To profile and assess the quality of this local data, the entity must either set up an on-premises program to analyze the data (which is time-consuming to set up and costly from a hardware and maintenance perspective), or upload the sensitive local data to a server so that it can be accessed on the network (in which case, the risk of privacy violation increases as the data is no longer on a local isolated computer). In some examples, it may be required that certain sensitive data (e.g., personally identifiable information) be stored only on a local storage device and not on a third-party / remote server (i.e., persistence at the disk level). To avoid these requirements, the client must rely on heavy processing / approval methods such as creating a bypass in the firewall or establishing a virtual private connection (e.g., virtual private link), which may cause specific regulatory and compliance issues depending on the nature of the sensitive data. Therefore, there is a need to access local data in a quick, low-cost, and secure manner (e.g., for profiling, quality assurance, analysis, etc.) without violating the basic privacy of the sensitive data.

[0005] In another example, when an entity attempts to perform DQ analysis on a specific dataset, it may discover that the entity's data sources are fragmented across a number of different data stores. To perform a DQ distribution on each of those different data stores, the entity would need to perform an individual analysis on each of those data stores. This is because each data store is isolated and siloed from the others. Therefore, it is also necessary to securely aggregate data from multiple different data stores so that DQ analysis can be efficiently performed across them.

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] This application is made to solve the problems in the above prior art.

MEANS FOR SOLVING THE PROBLEMS

[0007] Aspects disclosed herein are configured with respect to these and other general considerations. Also, while relatively specific problems may be described, of course, those examples should not be limited to solving the specific problems disclosed in the background art or elsewhere in this disclosure.

[0008] Non-limiting and non-exhaustive examples will be described with reference to the following drawings.

BRIEF DESCRIPTION OF THE DRAWINGS

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0010] Hereinafter, various aspects of the present disclosure will be described in detail with reference to the accompanying drawings. The accompanying drawings form a part of the present disclosure and show specific exemplary aspects. However, the various aspects of the present disclosure may be implemented in a variety of forms and should not be construed as limited to the aspects described herein. Instead, these aspects are provided so that the present disclosure is sufficient and complete and that the scope of each aspect is fully conveyed to those skilled in the art. Each aspect may be implemented as a method, system, or apparatus. Accordingly, each aspect may be in the form of a hardware implementation, a fully software implementation, or an implementation combining software aspects and hardware aspects. Therefore, the following detailed description should not be construed in a limiting sense.

[0011] Embodiments of the present application relate to improving data profiling and data quality by enabling on-demand and real-time distributed data access while preserving the privacy and security of underlying data. The underlying data, which may also be referred to as source data, can be defined as a data set (or a sample or subset of a data set) submitted by a user to a third-party service for profiling, sampling, classification, cataloging, or other forms of analysis. In the systems and methods described herein, the user controls what data is submitted / accessible from the system. As used herein, the term metadata, which may also be referred to as platform data, is defined as data or content that classifies, organizes, defines, or otherwise characterizes source data or the user's enterprise data structure (i.e., metadata), thereby establishing a comprehensive data catalog, data governance structure, business glossary, business process description, data stewardship roles and responsibilities, list of assets and domains, and similar data governance concepts within a third-party system. Platform data also includes logs, insights, statistics, or reports generated by a third-party system regarding the performance, availability, usage rate, completeness, or security of a third-party service. Platform data is not source data.

[0012] The metastore is an object store with a database or a file backing store. The metastore holds metadata associated with the actual underlying data stored in the data store. As part of the provision of "software as a service", if edge computing websites and edge / spoke websites are proliferating, some product components may no longer need to share the same network boundary for each client data source. That is, the client can do so from a single unified user interface when attempting to access each of its data sources, regardless of whether the data source is stored on a remote server (i.e., an accessible cloud) or in a local database / memory.

[0013] In one exemplary embodiment, a user (e.g., an entity) may own multiple different data stores that contain confidential data (such as PII). The user may connect those data stores to an edge node, which is used to reflect requests from the user to be executed against a base data store. The edge node may also be used to reflect specific data in the data store to the user, in which case that data is not actually stored in a permanent storage outside the data store (e.g., a disk). That is, the user can use the systems and methods described herein to access each of the user's data stores from a single unified user interface and perform analysis on the underlying data in the user's data stores, in which case the underlying data in the data store is not transferred to a separate storage location outside the data store. Instead, the accessed data is simply transmitted over the wire and temporarily stored in a volatile memory store specific to the currently running particular data session. That is, the underlying data in the data store is reflected to the user based on the user's instructions (this is transmitted to the edge node and then reflected to the base data store, whereby the data store then reflects the requested data to the user through a secure channel established by the edge node).

[0014] In another embodiment, a cloud-based metastore that backs a cloud service providing a single view of data may be composed of a data schema (e.g., platform data) that reflects the characteristics of the underlying data (e.g., data type, number of rows, number of columns, headers, etc.). The client can access this metadata through the view and request operations in real time with a single unified user interface. The cloud-based metastore may be communicatively coupled to a local data store (via the "data mirror" mechanism described herein), such that the requested data can be displayed to the user interface, but the underlying data displayed to the user is not stored on a persistent disk other than the data store (i.e., not transmitted and stored in a third-party disk location). Instead, the underlying data is transmitted over the wire and temporarily stored in a volatile memory location specific to the data session that is occurring in real time. When the data session ends (or simply when the user clicks and leaves the user interface), the reflected data from the underlying data store is removed from the volatile memory. Further, the reflected data transmitted between the data store and the edge node is end-to-end encrypted (e.g., using the TLS encryption protocol).

[0015] The mirroring distributed data access architecture described herein has several technical advantages. The advantages include: (i) enhanced security and privacy (since data is not permanently stored outside of the data store (i.e., the ownership and storage location of the underlying data remains only on the persistent disk within the client's data store)), (ii) reduced latency (since the time for parallel lookups by the distributed data access architecture is shorter than the individual sequential synchronous lookups of the actual underlying data at each unique data store location), and (iii) more efficient use of memory resources (since while the data session is continuous, the accessed data is reflected within a volatile memory store via the wire, eliminating the need to copy and store the underlying data in permanent storage), among others.

[0016] Another technical advantage of the mirroring distributed data access system described herein is an improvement in privacy-by-design. Typical systems implement privacy functions in a reactive or corrective manner. In contrast, in the systems and methods described herein, privacy is embedded within the architecture itself. Specifically, the architecture described herein enables the user to maintain ownership and manage their data under their own ownership within their data store. The underlying data within the data store is not transferred to third-party permanent storage for access outside of the user's data store. Instead, the underlying data may be viewed and accessed via an end-to-end encrypted data session, in which the user views reflected data temporarily stored in a volatile memory location reserved directly for the ongoing data session. That is, accessing the reflected data would require decrypting the transmitted data in real time while the data session is ongoing or while directly accessing a dedicated volatile memory location on the physical machine the user is operating on. Both scenarios are highly unlikely.

[0017] In another exemplary aspect, the metastore may be used to store the results of data quality and data profiling among other data analysis operations. Consider a "software as a service" application that is intended to connect various data stores containing source data that are in locations not directly accessible from the system housing the source data. To achieve access to a secure target data source, an edge node is set up at the secure location where the target data is stored. The edge node may periodically perform data quality assessment activities on the underlying data in the data store and update the local metastore with the results of those activities. If an error / problem is detected in the underlying data, the problem may be reflected (and stored in the metastore) from the location where the data is stored (e.g., in the form of an error notification) by the edge node to a cloud-based user interface. And the user can request to investigate the cause of the problem to analyze the problem in more depth. A user request to view specific data (in the cloud instance of the application) is sent to the edge node and then "reflected" in the data store. The data store receives the request, searches for the underlying data, and returns the result to the local edge node. The edge node securely transmits the data via an encrypted data session to a volatile memory location of the cloud instance of the application, thereby "reflecting" the underlying data to the user interface. The user may close the data session after viewing the reflected data, and the underlying data housed in the volatile memory is erased. Note that under the requirements of the standard TCP protocol, the underlying data provided to the unified user interface of the cloud-based data management application is provided unidirectionally, so data is not returned from the cloud-based data management application to the underlying data store.Also, reflecting data to a data store and then re - reflecting it to the volatile memory location of an application's cloud instance does not require the use of a metastore.

[0018] In another example, data virtualization may be incorporated. An edge - deployed data store can locally hold data from data sources that are inherently confidential. The underlying data of the data source can be shared on - demand with cloud - based applications (by encrypted package delivery), without exposing the unencrypted underlying data to the persistent storage of the cloud instance. In a data virtualization embodiment, the underlying data is copied to a volatile memory location isolated in an ongoing data session at the edge node. This enables a user to pose complex queries to that data set from a remote instance of the application and have the processing necessary to answer those queries be performed on the data in the volatile memory at the edge node, "reflecting" only the limited data necessary for the response to the user. The target data set is not permanently stored outside the underlying data store. That is, the data management application the user is operating cannot access the actual underlying data (source data) in the data store for the sake of a secure network perimeter (e.g., as long as the data session is live, the underlying data is encrypted and transmitted to the edge node and the temporary storage of the volatile memory store).

[0019] Figure 1 shows an example of a distributed system for securely reflecting data from a data store as described in this specification. The illustrated exemplary system 100 is a combination of interdependent components that interact to form a unified whole for integrating and enhancing data. The components of this system may be hardware components, or may be software implemented in and / or executed by the hardware components of this system. For example, system 100 includes client devices 102, 104, and 106, local databases 110, 112, and 114, network 108, and server devices 116, 118, and / or 120.

[0020] The client devices 102, 104, and 106 may be configured to receive and transmit data. For example, the client devices 102, 104, and 106 may store client-specific data. The client device may store specific client confidential data in databases 110, 112, and 114. The client device may also be communicably coupled via network 108 to a distributed metastore at an edge site, and network 108 may be communicably coupled to a centralized data management application or hub (e.g., a data governance center). The client device may also be configured to receive reflected data from a data store based on a specific request received by the data store. In some examples, the client-specific data may need to be stored in local databases 110, 112, and 114. If the client-specific data includes PII, the PII may be partitioned in local databases 110, 112, and 114. The client device may receive reflected data from one of databases 110, 112, and 114. In another example, the client devices 102, 104, and 106 may be regarded as "edge nodes" that re-reflect data and commands to other devices connected to data stores 110, 112, and 114 and / or network 108. The edge nodes and client devices may be connected to a central hub via network 108 and / or satellite 122 and may be connected to servers 116, 118, and / or 120. Servers 116, 118, and / or 120 may be third-party servers that can receive data reflected from data stores 110, 112, and 114 and transmitted from the edge nodes through network 108. An edge node (e.g., client device 102, 104, 106, or server 116, 118, and / or 120) may receive from a client, via a centralized data management application, a command to read a specific underlying data / data set in the data store to which the edge node is connected.The data may be displayed (i.e., reflected or mirrored) in the user interface at the central hub location, but the underlying data itself may remain in local storage in databases 110, 112, and 114. Specifically, the specific requested underlying data stored in the data stores (e.g., 110, 112, and 114) may be encrypted and sent via network 108 to a volatile memory location associated with the live data session, where the user can view the requested data as long as the data session is operational. For example, the user may be connected to network 108 via a device running a centralized data management application. From the application, the user may request to receive status updates regarding specific data quality tests being performed on the data stores (e.g., in database 110). Edge node 102 connected to data store 110 may receive commands via network 108, and edge node 102 may reflect that command in data store 110. Edge node 102 receives data from data store 110, and that data is re-reflected to the user device running the centralized data management application. Importantly, the actual underlying data stored in data store 110 is not sent to persistent storage outside of database 110. The underlying data is only temporarily reflected to the user device within the data management application by using volatile memory. The data stored in the volatile memory location may be erased as soon as the data session is closed. Note that the requested underlying data, encrypted and sent via network 108, is not stored in persistent storage at a third-party data location (e.g., databases 116, 118, and / or 120).

[0021] In one aspect, a client device (e.g., client devices 102, 104, and 106) is accessible to one or more datasets or data sources and / or databases (e.g., source data) that include client-specific data. In another aspect, client devices 102, 104, and 106 may be equipped to receive broadband signals and / or satellite signals from an edge node that reflect per-client read requests from a centralized data management application. The signals and information that client devices 102, 104, and 106 may receive may be transmitted from satellite 122. In addition to being able to communicate directly with client devices 102, 104, and 106, satellite 122 may be configured to communicate with network 108. In some examples, the client device may be, among other devices, a cellular phone, a laptop computer, a tablet, a smart home device, a landline phone, and a wearable (e.g., a smartwatch). In another exemplary aspect, client devices 102, 104, and 106 may be regarded as edge nodes communicatively coupled to a basic data store.

[0022] In some examples, a user may have write permissions to manipulate specific data stored locally in databases 110, 112, and / or 114. The write operation may be captured by an edge site (e.g., in client devices 102, 104, and 106) and then reflected in local data stores 110, 112, and 114 via network 108 and / or satellite 122. Such write operations may be serviced from a local metastore instance, which is to partition a third-party data management application from the basic data store. Read operations may also be serviced from the local metastore instance, and the access rights to confidential data may be mirrored (or reflected) to a centralized data management application instead of being transmitted from its basic location to a permanent storage location (e.g., third-party disk storage).

[0023] Figure 2 shows an exemplary input processing device for implementing a system and method for securely reflecting data from a data store as described herein. The input processing device 200 may be embedded in a client device (e.g., client devices 102, 104, and / or 106), a remote web server device (e.g., devices 116, 118, and / or 120), and other devices capable of implementing a system and method for securely reflecting data from a data store. The input processing system includes one or more data processing devices and is capable of executing algorithms, software routines, and / or instructions based on processing data provided from at least one client source and / or third-party source. The input processing system may be a factory-installed system or an add-on unit for a particular device. Further, the input processing system may be a general-purpose computer or a dedicated special-purpose computer. There are no restrictions on the location of the input processing system with respect to a client device or a remote web server device, etc. According to the embodiment shown in FIG. 2, the system of the present disclosure may include memories 205, one or more processors 210, a communication module 215, a data mirror 220, and a read / write handler 225. The data mirror 220 communicates with a local database (data store), generates metadata based on the base data stored locally and a user request (e.g., for DQ analysis), and is configured to reflect the data to a centralized data management application via an edge site. Another embodiment of the present technology may include some, all, or none of those modules and components, and may also include other modules, applications, data, and / or components. Furthermore, in some embodiments, two or more of these modules and components may be incorporated into a single module, and / or a portion of the functionality of one or more of these modules may be associated with another module.

[0024] Memory 205 may store instructions for executing one or more applications or modules on processor 210. For example, in one or more embodiments, memory 205 may be used to accommodate all or some of the instructions required to execute the functionality of metastore generator 220, read / write handler 225, and / or data mirror 230. Generally, memory 205 may include any device, mechanism, or data structure that is used to store information. According to some embodiments of the present disclosure, memory 205 may include, but is not limited to, any type of volatile memory, non-volatile memory, and dynamic memory. For example, memory 205 may be random access memory, a memory storage device, an optical memory device, a magnetic medium, a floppy disk, magnetic tape, a hard drive, a SIMM, SDRAM, RDRAM, DDR, RAM, a SODIMM, EPROM, EEPROM, a compact disk, a DVD, and / or others. According to some embodiments, memory 505 may include one or more disk drives, flash drives, one or more databases, one or more tables, one or more files, local cache memory, processor cache memory, a relational database, a flat database, and / or others. Further, as would be understood by one of ordinary skill in the art, many more devices and techniques for storing information may be used as memory 205.

[0025] The communication module 215 is associated with transmitting / receiving information (e.g., transmitting encrypted data from the data mirror 220, read / write commands from the read / write handler 225), and transmitting / receiving commands received via a client device or a server device, other client devices, a remote web server, etc. These communications may use any suitable type of technology, such as Bluetooth, WiFi, WiMax, cellular (e.g., 5G), single-hop communication, multi-hop communication, dedicated short-range communication (DSRC), or a proprietary communication protocol. In some embodiments, the communication module 215 transmits the information output from the data mirror 220 to a centralized data management application or hub. Specifically, the communication module 215 may facilitate "reflecting" the data from the edge node to the centralized data management application. The metadata may represent the actual underlying data (e.g., data type, number of rows, number of columns, other schemas, etc., also called platform data), but the underlying data (source data) is securely stored in the local database and is not transmitted over the network to a permanent storage location located on a third-party server. In another example, based on a user request reflected from the edge site to the underlying data store, specific data retrieved from the data store (as a response to the user request) may be transmitted via an encrypted data session to a volatile memory location and reflected in the user interface managed by the centralized data management application or hub. The data stored in the volatile memory is erased immediately when the data session is closed. In another example, the communication module 215 may transmit read / write commands from the centralized data management application. The read command may be a command to display a specific data set, and the write command may be a command to change specific data or the underlying application programming interface. Such communications may be transmitted to the client devices 102, 104, and / or 106, and the memory 205 and stored for future use.In some examples, the communication module may be built on top of the HTTP protocol through a secure REST server that uses RESTful services.

[0026] The data mirror 220 may be configured to facilitate the "mirroring" (or reflection) of data between the device and the edge node. The data store may be located between the edge computing site of the client environment and the local computer cluster where the local data is stored. The edge site may include an edge proxy, an edge controller, a data access service, and a data quality agent. The edge proxy may be defined as a component responsible for maintaining a secure channel for user instructions from a centralized data management application to act on the target location hosting the target data source. This channel is a logical channel, not a permanent channel. The edge controller may be defined as a component that constantly polls (i.e., listens) for data from a centralized data management application that accommodates user instructions using the secure channel provided by the edge proxy. As soon as a user instruction targeting a specific edge site is received, that edge site receives those instructions and facilitates communication to a data access service (or data quality service) that ultimately processes the request locally at the target data store. The data access service (or data quality service) may be defined as a service that processes user instructions. This is part of the data mirror 220. For example, the instructions may relate to data quality analysis, catalog services, policy management, data virtualization, etc. Specifically, the data access service may deserialize user instructions for local execution at the data store level. The response returned by the data store to the data access service may be re-reflected by the data mirror 220 module through the edge controller, the edge proxy, and ultimately the centralized data management application, and the data may be presented by an encrypted data session. The data quality agent may be used in a use case example of data quality analysis. The data quality agent is responsible for initiating data processing jobs and is controlled by the local DQ service on the client device.The DQ agent may execute a DQ job in accordance with user instructions processed by a data access service. The data access service may be communicatively coupled to a data store. The data quality service may also be communicatively coupled to a centralized data management application. The data mirror 220 can correctly send user instructions received via the edge proxy to the underlying data store and correctly reflect the data received from the underlying data store in the centralized data management application. For example, the edge site may be located within a virtual LAN that can communicate with both a data quality module (e.g., a data quality metastore) and a centralized data management application. The data mirror 220 may also be configured to receive from the communication module 215 information regarding the type of data reflected in the metadata. For example, the data mirror 220 may be configured not to display certain confidential information to the centralized data management application.

[0027] The read / write handler 225 may be configured to receive commands from the centralized data management application, locally replicate those commands on the local computer, and perform read / write of specific data stored locally. For example, a read command from the centralized data management application may include a request to access new data. The read / write handler 225 may include a data access service, which receives user commands from the edge proxy and processes those commands to be executed at the local data store level. The read / write handler 225 may first receive a user command via the data mirror 220 and pass the command to the edge site (via the edge proxy and the edge controller). Then the edge site may pass this command to the data store (via the data access service) (the data store is communicatively coupled to the local computer), whereby the read command is locally executed against the data stored in the local database. The read / write handler 225 may receive a response from the DQ service, which includes the specific data requested based on the user command. Thereafter, the data may be encrypted and sent to a volatile memory location, where the user may be able to view the reflected data in the centralized data management application.

[0028] The data mirror 220 is also configured to prevent specific data from the local data store from flowing to a permanent storage location other than the approved storage location. This is particularly important when the user is accessing the data store in the reflection mirror process (via the data mirror 220) and / or receiving an update regarding the data quality of a specific data store and / or issuing a command that acts on the underlying data in the data store. The underlying data is encrypted and is not sent to permanent storage but is temporarily reflected to the user by means of volatile memory.

[0029] Similarly, write commands from the read / write handler 225 may send the write commands to the local database via the edge proxy, the edge controller, and the data access service. From the user's perspective, those commands may appear to occur synchronously, but may occur asynchronously due to the edge site and edge site deployment. This is to ensure that each command is accurately captured and sent as long as the security and privacy of the underlying local data are protected.

[0030] In some examples, the read / write handler 225 receives commands from the centralized data management application via the communication module 215.

[0031] Figure 3 shows an exemplary method 300 for securely mirroring data from a data store. The method 300 first, at step 302, receives a request to access remote data. At step 302, a request to access data in a database (e.g., the local databases 110, 112, and 114 of FIG. 1) is received. This request may be received from a centralized data management application communicatively coupled to an edge node communicatively coupled to the data store. At step 304, the request is analyzed to determine which data store the request is targeting, and the result is used to identify the correct target edge node communicatively coupled to that data store. Some pre-loaded requests may already exist, such that when the system receives a particular request, that request is automatically associated with a particular edge node without further analysis.

[0032] In step 306, the correct edge node receives the request. Another sub-step may be performed between step 304 and step 306. For example, after the correct edge node is identified in step 304, the request may be transferred to the edge gateway, where the request is tagged for the target edge node. While the request is at the gateway, the channel through which the request arrives remains open and the response is pending. Further, the target edge node may include a controller that polls the edge gateway (e.g., every 5 seconds or at other time intervals). The target edge node controller polls the gateway and, when it finds a request for that edge node, transfers the request to that edge node.

[0033] The reason the target edge node polls the edge gateway is that the edge gateway cannot directly transfer the request to the target edge node because there is no network path between the target edge node and the edge gateway. The edge gateway cannot access secure, non-publicly addressable endpoints within the system. Instead, the user instruction is sent to a queue associated with the target edge node. And the target edge node polls the queue until the user instruction is received in the queue at the edge gateway. When the user instruction enters the queue and the target edge node polls the queue, the user instruction is sent to the target edge node.

[0034] In step 308, the edge node executes the request and obtains data from the data store to which the edge node is connected. In some examples, the data may be obtained from the connected data store, and in some other examples, the data may be obtained from the connected metastore. The data is encrypted while it remains at the edge node.

[0035] In step 310, the data obtained by the edge node is reflected back to the user device on which the centralized data management application is operating. Before step 310 and after step 308, the edge node may query an edge controller capable of opening a secure data channel between the gateway and the edge node. The data reflected from the edge node to the user device may be a status report regarding the data quality of the underlying data store. If the status report indicates a problem, the user may issue a subsequent request to obtain the raw data in the data store to investigate the problem. The data may be reflected from the edge node to the user device, but is stored only in a volatile memory location and not in permanent storage other than the data store. Therefore, when the data channel is closed, the reflected data is immediately erased.

[0036] In some examples, from the client's perspective, the client may first receive a summary of each data store accessible to the client (e.g., a broadcast request), regardless of whether the data store is local or external. Then the client may execute read commands and / or write commands via the user interface of the centralized data management application. A read command for a local data store may pass through the distributed metastore at the customer's edge site, and then obtain the requested metadata and reflect it back to the user viewing the user interface of the centralized data management application. When this is done, the underlying local data is not actually sent at all to the permanent memory storage location hosted by a third party. Instead, the data is sent in an end-to-end encrypted data session and stored in the volatile memory storage location associated with that encrypted data session. When the data session ends, the volatile memory storage location is immediately erased. In some examples, the end of the data session may occur when the user exits the centralized data management application, when the user clicks outside the current screen of the user interface, and / or when a specific expiration time until the data session automatically ends has elapsed.

[0037] FIG. 4 shows an exemplary distributed environment 400 for securely reflecting data from a data store. Environment 400 shows a centralized data management application (e.g., a data intelligence cloud, a data governance center, etc.) 402. Within this data management application, there are an edge management application 404, a data quality web application 406, and a data quality metastore 408. The DQ web application 406 is the main interaction point for the user. This cannot access the underlying data source. The DQ metastore 408 stores data necessary for the DQ cloud application 402 to function (catalog functions, user functions, management functions, etc.), and to cache metadata (aggregates, schemas, platform data, etc.) that is not considered confidential data originating from edge nodes / sites. Storing such data in the DQ metastore 408 reduces read / write requests to the underlying local data store. Edge management 404 is the main gateway for user interaction with remote data sources. For example, when a user clicks on a data source or data in the user interface of the DQ web application 406, the DQ web application 406 sends a request to edge management. Edge management queues the request for pickup by a target edge node / site (e.g., an edge node marked as owning / being close to the target data source).

[0038] The wide area network (WAN) layer 410 enables communication to flow into the data management application 402 from the edge sites 416 and 418. The customer network 412 includes at least one local computer cluster having a local data store and at least one edge site including a metastore. In this exemplary environment 400, the customer network 412 includes a local computer cluster 414 that includes a local database. The network 412 also includes edge site A 416 and edge site B 418, both of which are located on a virtual local area network (LAN). Each edge site houses at least one edge controller, a data quality agent, a data quality service, and a metastore. The metastore is configured to communicate with the local data store to obtain metadata representing the actual underlying data stored in the local database. And this metadata may be returned to the data management application 402 via the WAN layer 410. However, the underlying local data remains at the local computer cluster 414. The metastore is merely a reflection of the underlying local data.

[0039] The edge nodes / sites 416 and 418 are where user instructions initiated within the DQ cloud application 402 are actually executed. Each component operating here has been described in detail with reference to FIG. 2. The difference between edge node 416 and node 418 is which data the node is servicing based on the scoping requirements and / or constraints set by the user. There may be one or multiple edge nodes depending on the number and location of each user data store.

[0040] Local computer clusters 414A and 414B represent computers that perform the task of actually scanning the basic data (source data) with respect to quality. In some examples, as shown in the figure, the local computer cluster may exist for each edge node. In another example, two different edge nodes may access the same local computer cluster, but the access is permitted only to a specific portion of the basic data stored in the local data store (this is, for example, due to the scoping constraints of each edge node).

[0041] The edge site metastore located within the edge nodes 416 / 418 may be configured to store the data necessary to operate the edge DQ! service. Additionally, the metastore located within the edge node is a place where confidential custom data extracted from the basic data source as a result of DQ checks and / or user instructions is stored. For example, as a fully faithful example of a record for which the DQ check fails, specifically, there are confidential personal values (e.g., account number, account balance, social security number, etc.), and these may be stored in the edge metastore.

[0042] The metastore of the edge node / site is isolated from the associated data store served by the edge node, while the DQ metastore 408 serves all available data sources connected to the centralized data management application 402.

[0043] The data quality stack, along with the metastore, represents both synchronous and asynchronous data flows. The data flows may be divided into three different classes, namely, the user class, the job class, and the agent class. The user class may enable synchronous read / write operations performed by the user. As an example of a synchronous read / write operation, it may be possible to enhance the configuration via a user interface or API. The job class may enable asynchronous read / write operations. For example, when the user needs to view persistent dataset profiles and statistics, the read / write operations of the job class will be used. The agent class may enable asynchronous read / write operations for managing job details such as status updates.

[0044] In this exemplary environment 400, job input / output operations and agent input / output operations may be performed in the context of the edge site. For some jobs such as DQ check jobs, configurations such as connection details and model configurations are required in the metastore. That is, in order to establish a specific class of data flow, information may be required regarding how the metastore is communicatively coupled to the local database and what type of metadata is being retrieved.

[0045] In another example, non-confidential data may have a two-way flow between the data management application 402 and the customer's network 412.

[0046] The metastore in the exemplary environment 400 is used to serve the data management application 402. This is because the metastore accommodates metadata described according to the underlying local data. For non-confidential data, read operations may be served directly from the local metastore housed in the local computer cluster 414. For confidential customer data, read operations may be delegated to the corresponding edge site depending on the data and the scope characteristics of the edge site. Write operations may be atomically mirrored to the corresponding edge site. At that edge site, write operations may be processed locally.

[0047] The DQ metastore 408 mirrors instances of the data management application 402, with the exception of certain data sets that do not reside at a given edge site. The local metastore on the local computer cluster 414 may accommodate confidential customer data as a result of certain data processing jobs and data profiling jobs. At the edge site, read operations may be served from the local metastore instance in the computer 414. Write operations may also be performed locally. Write operations may be mirrored to the persistent write-ahead log (WAL) backed by the local metastore instance when targeting non-confidential data. The published WAL may asynchronously publish entries to the data quality web application 406 operating in an instance of the data management application 402 (i.e., the data management cloud or the data intelligence cloud).

[0048] FIG. 5 shows an example of a suitable operating environment in which one or more embodiments of the present embodiment can be implemented. This is merely an example of a suitable operating environment and does not imply any limitations regarding the scope of use or functionality. Other well-known computing systems, environments, and / or configurations that are considered suitable for use include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronics (such as smartphones), network PCs, minicomputers, mainframe computers, distributed computing environments that optionally include any of the foregoing systems or devices, and the like.

[0049] In its most basic configuration, the operating environment 500 typically includes at least one processing unit 502 and memory 504. Depending on the exact configuration and type of the computing device, the memory 504 (which stores, among other things, information related to the detected device, association information, per-person gateway settings, and instructions for implementing the methods disclosed herein) may be volatile memory (e.g., RAM), non-volatile memory (e.g., ROM, flash memory, etc.), or some combination of the two. This most basic configuration is shown in FIG. 5 by dashed line 506. Additionally, the environment 500 may further include a storage device (removable 508 and / or non-removable 510), such as a magnetic or optical disk or tape, among others, and is not limited thereto. Similarly, the environment 500 may also have an input device 514 (e.g., keyboard, mouse, pen, voice input, etc.) and / or an output device 516 (e.g., display, speaker, printer, etc.). The environment 500 may also include one or more communication connections 512 (e.g., LAN, WAN, point-to-point, etc.).

[0050] The operating environment 500 typically includes at least some form of computer-readable medium. The computer-readable medium can be any available medium that is accessible from the processing unit 502 or other devices that make up the operating environment. By way of example and not limitation, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and nonvolatile, removable and nonremovable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other tangible medium that can be used to store the desired information. Computer storage media does not include communication media.

[0051] Communication media embodies non-transitory computer-readable instructions, data structures, program modules, or other data. The computer-readable instructions may be carried in the form of a modulated data signal, for example, a carrier wave or other transport mechanism, and may be carried on any information delivery medium. The term "modulated data signal" means a signal in which one or more characteristics of the signal are set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media (such as a wired network or direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above media are of course included within the scope of computer-readable media.

[0052] The operating environment 500 may be a single computer operating in a network environment that uses a logical connection to one or more remote computers. The remote computers may be personal computers, servers, routers, network PCs, peer devices, or other common network nodes, and typically include many or all of the elements described above, as well as other elements not so mentioned. The logical connection may include any method supported by an available communication medium. Such network environments are common in offices, enterprise-scale computer networks, intranets, and the Internet.

[0053] Aspects of the present disclosure have been described above with reference to, for example, block diagrams and / or operational diagrams of methods, systems, and computer program products according to aspects of the present disclosure. The functions / operations described in each block may be performed in an order different from the order shown in any flowchart. For example, two blocks shown as consecutive may actually be executed substantially simultaneously, or depending on the required functionality / operation, may be executed in the reverse order.

[0054] The description and illustration of one or more aspects shown in this application are not in any way intended to limit or restrict the scope of the claimed disclosure. The aspects, examples, and details shown in this application have been considered in sufficient detail to convey ownership and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail shown in this application. Various features (both structural and methodological) may be selectively included or omitted to generate embodiments having a particular set of features, whether combined and illustrated or individually illustrated and described. Those skilled in the art provided with the description and illustration of this application will envision variations, modifications, and alternative aspects within the scope of the broad spirit of the general inventive concept embodied in this application that do not depart from the broad scope of the claimed disclosure.

[0055] As can be understood from the foregoing, specific embodiments of the present invention are shown herein for illustrative purposes, but various modifications may be made without departing from the scope of the present invention. Therefore, the present invention is not limited except as defined by the appended claims.

Claims

1. A system that reflects data from at least one data store, comprising: a memory configured to store non-transitory computer-readable instructions; a processor communicatively coupled to the memory, the processor, when executing the non-transitory computer-readable instructions, receives at least one user instruction targeting the at least one data store; sends the at least one user instruction to a queue associated with an edge node, the edge node being associated with the at least one data store, the edge node polling the queue to receive the at least one user instruction, the step of sending to the queue; initiates an encrypted data session with the edge node; receives reflected data retrieved from the at least one data store via the encrypted data session, the reflected data being stored in at least one volatile memory location associated with the encrypted data session, the step of receiving the reflected data; displays the reflected data during the progress of the encrypted data session; the processor configured to perform; A system comprising.

2. The system of claim 1, wherein the processor is further configured to erase the reflected data from the at least one volatile memory location when the encrypted data session is disconnected.

3. The system of claim 1, wherein the processor is further configured to receive at least one data quality result from a metastore.

4. The system of claim 3, wherein the processor is further configured to send a second user instruction to the edge node based on the at least one data quality result from the metastore.

5. The system of claim 3, wherein the at least one data quality result is an error notification associated with the at least one data store.

6. The system of claim 1, wherein the at least one user instruction is deserialized at the edge node and then executed at the at least one data store.

7. The system according to claim 1, wherein the reflected data is transmitted through the encrypted data session after being encrypted.

8. The system according to claim 1, wherein the at least one user command is at least one of a read command and a write command.

9. The system according to claim 1, wherein the processor is further configured to generate a metastore.

10. The system according to claim 9, wherein the metastore includes at least one scoping requirement associated with the basic data stored in the at least one data store.

11. The system according to claim 9, wherein the metastore is associated with a second data store.

12. The system according to claim 1, wherein the at least one user command is a view command for viewing information associated with a plurality of data stores in a unified user interface.

13. The system according to claim 12, wherein the processor is further configured to display platform data associated with the plurality of data stores in the unified user interface based on the view command.

14. The system according to claim 9, wherein the metastore includes at least one of a schema, a user command, a DQ analysis result, and platform data.

15. A method for reflecting data from a data store, comprising: connecting at least one edge node to the data store; transmitting a first user command for processing the data stored in the data store with respect to data quality; receiving at least one data quality result from the at least one edge node; transmitting a second user command for reading a part of the data stored in the data store based on the at least one data quality result received from the at least one edge node; starting an encrypted data session; receiving reflected data from the at least one edge node through the encrypted data session; displaying the reflected data on a user interface. A method comprising the above steps.

16. The method according to claim 15, wherein the reflected data is stored in a volatile memory store while being displayed on the user interface.

17. The method according to claim 16, further comprising the step of immediately deleting the reflected data from the volatile memory store when the encrypted data session is disconnected.

18. The method according to claim 16, further comprising the step of immediately deleting the reflected data from the volatile memory store when the user interface is disconnected.

19. A computer-readable medium storing non-transitory computer-executable instructions that, when executed, cause a computing system to perform steps of reflecting data from a data store, the steps comprising: connecting at least one edge node to the data store; sending a first user instruction to process data stored in the data store with respect to data quality; receiving at least one data quality result from the at least one edge node; sending a second user instruction to read a portion of the data stored in the data store based on the at least one data quality result received from the at least one edge node; starting an encrypted data session; receiving reflected data from the at least one edge node via the encrypted data session; displaying the reflected data on a user interface; comprising a computer-readable medium.

20. The computer-readable medium according to claim 19, wherein the at least one data quality result from the at least one edge node is an error notification.