Computer-implemented system, and computer-implemented method

The shared network of data nodes addresses inefficiencies in data silo environments by enabling version-managed, securely accessible data sharing, thereby reducing deployment time and enhancing data management within organizations.

JP7682880B2Active Publication Date: 2025-05-26SHINCHI INK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022529877
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-22
Filing Date
2020-11-23
Publication Date
2025-05-26
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

Conventional data silo environments lead to inefficiencies such as wasted time, storage space, inaccurate data, security vulnerabilities, and friction within organizations due to isolated data access and management.

Method used

A system for creating a shared network of data nodes, where each node contains version-managed data, an access control layer, and a metadata layer, allowing nodes to connect and form a network that enables data sharing and collaboration while maintaining security and control.

Benefits of technology

This solution reduces the time and effort required to build data management solutions, prevents data silo environments, enables flexible enterprise alignment, and simplifies data integration, leading to faster deployment of solutions and improved data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682880000001
    Figure 0007682880000001
  • Figure 0007682880000002
    Figure 0007682880000002
  • Figure 0007682880000003
    Figure 0007682880000003
Patent Text Reader

Abstract

The following provides a platform for creating a shared network of data nodes. Each data node is self-describing, self-connecting, and self-protecting. The data network taught herein can prevent a data silo environment. Each node has a dataset containing versioned data, an access control layer that restricts user access to the dataset, and a metadata layer that defines the characteristics of the dataset and connects it to other nodes. One or more links are created to associate a node with a subsequent node to create a network of data nodes, such that changes in the dataset affect changes in the network of data nodes, and the network of data nodes includes a query layer for interacting with the dataset and subsequent datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following relates to a network of data sets such as a data fabric, a platform for providing such a network, and a method of constructing and using such a network, data set, and system.

Background Art

[0002] Conventionally, enterprises (meaning relatively large corporations such as large enterprises, medium-sized enterprises, public offices, etc.) have operated in a data silo environment. A data silo is, for example, a group of data sets that are accessible through one application but isolated from the rest of the organization. Data silos are usually the result of data being collected by analytics tools or data being generated by business applications. There are many drawbacks to a data silo environment.

[0003] For example, data silos can result in a large amount of wasted time within an organization. Data cannot be automatically streamlined between applications, and the data is isolated within each application. This means that a certain team may recognize that they need data they do not have, search for where that data is within the organization, manually access that data, and then wait until they can analyze that data for their own purposes. The data may no longer be valid by the time it is collected.

[0004] Data silos can also result in wasted storage space. For example, if each employee of an organization needs a copy of the data and saves it in their company's storage folder, a large amount of storage space will be wasted, which can be very costly.

[0005] Another drawback of a data silo environment is that it cannot maintain data accuracy. Isolated data becomes old and thus inaccurate and less suitable for use over time.

[0006] Data silos can also create security vulnerabilities. For example, once data is copied, the owner of the data set can no longer guarantee its confidentiality without a difficult and costly process that depends on the approval of other teams. If copies of the data set are stored on each team member's computer, there is a greater potential for it to be hacked.

[0007] Data silos also create friction within an organization. This is because in a data silo environment, each team has access only to its own data and thus it is the only data they deal with. For example, each team may work independently rather than collaboratively, creating a fragmented organization.

[0008] Figures 1 and 2 illustrate the creation of each conventional new custom application, involving the creation of a new data silo 24 and a data replica 26. As shown in Figure 1, an application 10 developed conventionally would be programmed to include a respective unique user interface 14, an API 15 (not shown in Figure 1), security and control functions 16, a data integration function 18, a data persistence function 20, and a data exposure function 22.

[0009] In addition, each application 10 requires its own database 24, and linked or related data 26 is created, imported, updated, maintained, etc. in that database 24. Enterprise legacy data 28 will also need to be separately imported into each application 10. Users 12 will be permitted access to each application 10 using separate security and controls 16. The data can then be exposed to a larger enterprise data lake 30, for example, to perform analytics and other data processing operations. These applications 10 may also be highly dependent on desktop tools, requiring hundreds or even thousands of instances of these tools to effectively deploy new solutions. Furthermore, it has been found that data management is often lacking in traditional data silos, and the management implemented in one application or data silo is often not reusable in another application or data silo.

[0010] Due to these inefficiencies and redundancies, it often takes weeks to create new applications 10 and solutions, and even more months to deploy them. In environments where these solutions are needed quickly, the time for deployment can be considered a competitive disadvantage or, at best, a waste of resources. SUMMARY OF THE INVENTION

[0011] One aspect relates to a system for creating a shared network of data nodes. The system comprises at least two nodes, each node having a dataset containing version-managed data, an access control layer for restricting user access to the dataset, and a metadata layer for defining the characteristics of the dataset and connecting to another node. One or more links are created to associate the nodes with subsequent nodes so as to create a network of data nodes such that changes in the dataset affect changes in the network of data nodes, and the network of data nodes comprises a query layer for communicating information with the dataset and subsequent datasets. Each data node is self-describing, self-connecting, and self-protecting.

Brief Description of the Drawings

[0012] Next, embodiments will be described with reference to the accompanying drawings.

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8a

Figure 8b

Figure 8c

Figure 8d

Figure 8e

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13a

Figure 13b

Figure 13c

Figure 14

Figure 15

Figure 16A

Figure 16B

Best Mode for Carrying Out the Invention

[0014] The following provides a platform for creating a shared network of data nodes. Each data node is self-describing, self-connecting, and self-protecting. The data network taught in this specification can reduce the time and effort required to build large-scale (e.g., enterprise-level) data management solutions. Some advantages include prevention of data silo environments, flexible enterprise alignment through a compliant application programming interface (API) integration layer, and simplified data integration with minimal effort based on data reuse between applications, effectively reducing or eliminating the need to develop "applications" in the traditional sense. Each node has a dataset containing version-managed data, an access control layer that restricts user access to the dataset, and a metadata layer that defines the characteristics of the dataset and connects to another node. One or more links are created to associate the nodes in order to create a network of data nodes such that changes in the dataset affect changes in the network of data nodes, and the network of data nodes includes a query layer for exchanging information with the dataset and subsequent datasets.

[0015] As will be described below, this system virtually eliminates the need for conventional application development, and thus, the data collaboration network and its individual data nodes eliminate the need for applications by enabling entities having access to the data to utilize the same data without the significant effort associated with conventional application development and data silo creation. Accordingly, rather than writing code once for each solution or application, the system enables more rapid delivery of solutions by separating the data from the UI and enabling the configuration of security and control layers around the data. Also, data collaboration networks within separate organizations may be linked together to create a network of networks, creating a supernetwork of linked data rather than isolated data silos.

[0016] FIG. 2 shows the conventional creation of each of a new custom application involving the creation of a new database 24 and a replication 26 of the data. Each of the conventionally developed applications 10 requires its own database 24, and linked or related data 26 will be created, imported, updated, maintained, etc. in that database 24. Each of the databases 24 stores a copy of the data 26 accessed by the application 10. Conventional applications embed data sets. Conventionally, data is stored as a copy of the data in a data silo 24 behind individual applications 10. This creates an environment such as a data silo that can be disadvantageous.

[0017] Figure 3 shows a schematic diagram of a network of data nodes. The data network 34 is created by connecting individual nodes 36 so as to form a network of data nodes 34. Raw data or data records 35 are accessible via a query interface 33, which is connected to the network of data nodes 34. For example, Figure 3 shows four different application experiences 32a, 32b, 32c, and 32d, all using data from the same data network. This not only eliminates the need to create, import, update, and maintain separate databases, but also eliminates the need to manage the security and control systems 16, data integration systems 18, data persistence systems 20, and data publishing systems 22 for each application.

[0018] Therefore, it is relatively easy to add a new application experience 32. Figure 4 shows where a new application experience 32e has been added to the existing architecture shown in Figure 3. The addition of the new application experience 32e does not require additional security and control functions 16, data integration functions 18, data persistence functions 20, and data publishing functions 22. The application experience 32 acts as a custom user interface for exchanging data and information.

[0019] Thus, it will be understood that any number of nodes 36 as well as application experiences 32 can be added to the network 34. As newer application experiences 32 are added, the network 34 of data sets grows and new links are formed between the newer data sets 26.

[0020] Data Set Nodes, Network of Data Nodes, and Platform Implementation

[0021] Next, looking at FIGS. 5 and 6, the platform described herein is configured to manage data 35 as a network 34 of data nodes. FIG. 5 shows an example of a dataset node 36, and FIG. 6 shows an example of a network 34 of nodes 36. Each dataset 26 comprises data 35. The datasets are version managed 38 and include versions of data such as a first version 35a, a second version 35b, and a third version 35c. It will be understood that any number of versions are possible. An access control layer 39 is built on top of the dataset 26. A metadata layer 37 is built on top of the access control layer 39. As seen in the single node 36 shown in FIG. 5, there is a record of data 35 in the core of the node. The record of data 35 cannot be accessed without first passing through the metadata layer 37 and the access control layer 39. The node 36 comprises a metadata layer 37 that makes the node 37 self-describing and self-connecting. Thus, each dataset 26 comprises its own metadata layer 37, which includes information about the dataset, such as characteristics and relationships (i.e., links) with other nodes, along with data that associates any record of the current node with one or more records of other nodes.

[0022] The node 36 can be self-governing by having a built-in control layer 39 to ensure data integrity and to provide governance management such as data version management 38 and change approval. Data version management 38 is shown in FIG. 11 and is described in more detail below. That the node 36 can be self-protecting means that the node has a built-in security layer for managing entitlements. The node is accessible through either the platform's metadata-driven user interface or API.

[0023] Typically, the security and control of data 16 resided in the individual applications 10 shown in the prior art example of FIG. 1. This can be dangerous. Because the security and control are linked to the individual applications and thus will not track the data when it is copied. Thus, the data will be vulnerable or insecure when it is copied.

[0024] In the system shown by the present invention, the security and control are incorporated into each data set 26. This is understood to be much more efficient and secure as it ensures a universal implementation regardless of how the data is used. In FIG. 1, if data is copied between applications as part of performing data integration, the security and control should not track the data and should be re-implemented by each application. This is avoided because the access control layer 39 is constructed as a layer on top of the data set 26.

[0025] Looking next at FIG. 6, it can be seen that a first node 36a can be connected to a second node 36b and a third node 36c. This forms a network 34 of data nodes. The nodes 36 are connected via links 40 which are relationships between the nodes 36. The links 40 are defined within a metadata layer 37 of the nodes which form the basis of the network 34 of data nodes. The links 40 relate data records 35 within the data nodes and an individual link 40 has the ability to relate a data record of a first node 36a to one or more data records of another node 36b, 36c.

[0026] Since each node 36 is self - describing, self - connecting, and self - protecting, the dataset 26 within the network 34 is not restricted by application boundaries. Data is no longer siloed as in the prior - art illustration shown in FIG. 1. This eliminates the need for databases managed by the applications shown in FIG. 1 and enables a simpler and more effective use of the dataset.

[0027] For any individual, their network is shaped by the datasets that have access for information exchange. Even the links connecting the nodes are only made public / accessible if the individual has access to the target of the link. Any user may only have access to a very small part of the entire network, as shown in FIG. 10.

[0028] By enabling data nodes to be connected to other data nodes while still remaining self - describing, self - connecting, and self - protecting, the need for the data - disclosure component and data - integration component shown in FIG. 1 is eliminated when exchanging information within the data network. There is no need to distribute copies of the data. Instead, links are established to enable access and the use of data that can be managed by different owners.

[0029] A platform configured to create, modify, and exchange information with the network 34 of data nodes is shown in FIG. 7A. The system has a query interface 33 through which data can be queried using the API 41 or the application interface 32. The user 12 exchanges information with the user interface 42 or the application interface 32 to manage both the configuration of the nodes and the data of the nodes. The API 41 and the UI 42 can be dynamically provided by the platform itself.

[0030] The query interface 33 provides a query engine for exchanging information with the network 34 and all datasets 26 within the network. This enables the exchange of mutual information beyond the scope of a single node 36, i.e., beyond what is possible using the APIs available at each node 36. The queries are written in a platform script language that is built on SQL in one example and designed to utilize the available links within the network, allowing the user to traverse the relationships between nodes when executing the query. FIG. 9 provides an example of a query that traverses relationships by leveraging the dot notation of the script language when querying nodes. The platform's script language allows the user to read / create / modify / delete data and manage the nodes on the network. Queries written in the platform's script language may support ACID transactions.

[0031] In one embodiment, legacy data 28 may be imported into the network 34. A connector may be used to bridge the gap between the legacy data 28 and the data network 34. The connector may enable the synchronization of data from outside the platform to self-describing, self-connecting, and self-protecting datasets, as well as the reverse flow for pushing data to legacy applications or the enterprise data lake shown in FIG. 1.

[0032] FIG. 7B shows one embodiment in which multiple networks 34 are connected to each other. For example, networks 34 within separate organizations may be linked to each other to create a network of networks of linked data rather than isolated data silos. The network of networks of data is referred to as a supernetwork.

[0033] Next, looking at FIGS. 8a through 8e, sample screenshots are provided to explain the user experience components shown in FIG. 7. The user experience 42 may use a web form to create and manage nodes.

[0034] FIG. 8 shows a sample screen of an example of a node. FIG. 8a is a UI for managing the data of the node. Note that the column of Primary Client is a link, and FIG. 8c shows the definition. This experience enables the user to click through and traverse the links to the relevant records of the linked data set.

[0035] FIG. 8b shows a sample screen for designing a data set, and the screenshot on the right end includes the definition of metadata.

[0036] FIG. 8c provides a sample screen for defining the link between the current data set and another data set. As shown in FIG. 8d, all data changes in the node are automatically version - managed. FIG. 8d is a sample screen of a collaboration log that displays the version of the data.

[0037] FIG. 8e provides a sample screenshot of controls that can be configured on an individual node.

[0038] Thus, the platform provides a native metadata - driven user interface for the user to interact with the data network and information. This, combined with the ability to create a custom application experience, provides an alternative to the application requirements in the sense illustrated in FIG. 1.

[0039] As shown above, FIG. 9 provides a sample of a platform script language query that utilizes the links existing between nodes to traverse the data network 34 and obtain data. Specifically, by using the query interface 33 provided by the platform, query results can be obtained using dot notation.

[0040] Data version management Data version management 38 in the platform can be performed on the data. FIG. 11 shows the data version management of a single record. The first version may represent the initial creation. Thereafter, each time the dataset is changed or modified, a new version can be created. For example, version 18 approved the changes made by a user named Dan Demers on November 13, 2020. Each version captures the details of the user who made the change, along with the timestamp of that change.

[0041] For example, in one version, a user may delete a record of data, which will be shown grayed out as it is a deletion. In another example, a user may restore a record of data from the trash, which will be shown as a restoration made by the user. In another example, an operation to revert can be performed by a user to restore the first version of the dataset. This will create a new version and will not affect the version history.

[0042] Data-level access control Figure 10 provides an illustration of how two different users (User 12a and User 12b) view the network of the same data nodes. Here, there are two different users 12a and 12b who are exchanging information with the data network 34 via the query interface 33. The black links 40 and nodes 36 represent the data sets 26 that the users have access to. The gray nodes and dashed links are not available to any individual and seem as if they do not exist at all. That is, each user has a unique perspective when viewing the data network 34, and the data network can appear very different to each of them.

[0043] The partially filled nodes indicate that, according to the set rules, the user can only have access to a subset of the data within the node 36. The access control layer 39 built on the node 36 can be very fine-grained and can define rules that allow the user 12 to view / edit / approve the data under specific conditions. Figure 12, which will be described later, provides an example of data-driven access entitlement.

[0044] Looking at Figure 12 next, the data set can support very fine-grained access control. Figure 12 is an example and shows two independent permissions that define what a user can edit. One of these permissions has conditions based on the data of the current node. These conditions also extend to the nodes and utilize the links to traverse the related databases to determine whether the user has access. Similar fine-grained control is also available for what a user can view or approve. In this example, the grayed-out cells will not be editable by the user.

[0045] The network of nodes can be linked through any application configured to utilize the platform and to exchange information with the platform. The network is an exchange of mutual information of the relationships between data and does not necessarily affect where the data in the underlying persistence is stored or what devices are used to store the data. In this way, existing technologies within an enterprise can be used without the need to adopt one or more new databases while executing the platform on top of this technology.

[0046] Network 34 can be constructed from a series of data nodes 36. Dataset 26 can be linked to other datasets, and queries can be constructed by applying a scripting language as described herein. This enables users, such as developers, to build application experiences using existing datasets and by creating new data, thus leveraging and evolving the existing network of nodes for future application development.

[0047] Note that newly created nodes can have data added and manipulated by the user and / or import existing data, such as legacy data. In this way, new nodes can be added to the network using existing enterprise data, for example, from legacy applications or data storage components. Nodes may be user-managed, synchronized, and / or application-managed, and any particular node may have individual attributes or sets of attributes that are user-managed, or synchronized, or application-managed. That is, nodes can be controlled and managed on a per-attribute basis. More specifically, the records of one node can be linked to one or more records of another node.

[0048] The enterprise environment can not only quickly build solutions by reusing existing queries and nodes, but also continuously strengthen and enhance the data network as new data is created or imported for the newly built solutions.

[0049] The configuration of the platform and various components enable several unique features that improve the way enterprises and other users of the data build solutions. By providing a network 34 of data nodes as shown in FIG. 6 and providing data layer control and interfaces to that data network, solutions can be built with less effort and faster than the conventional approach of replicating these functions in data silos.

[0050] In the conventional approach, each solution is implemented as a separate application 10 that requires a separate database 24 for persistence (see FIG. 1). In contrast, the system described herein provides a single platform for managing data for multiple solutions. Using that platform, persistence can be provided to the solutions via an API. Due to this structure, data can be reused across solutions without requiring individual data integration for each solution. This reduces the infrastructure burden, especially when creating many solutions.

[0051] Conventional databases are recognized as being designed to be used by code written by a development team. This code generally runs under the account of the application 10 rather than under separate accounts of individual users 12. That is, access control in a conventional application development environment is typically not as robust as to allow a single database to be used for multiple applications 10 by multiple teams each having a plurality of users 12. To overcome this limitation, the system described herein provides data access control that restricts what all users (including developers) can view and edit. The security layer of the platform applies these controls, and thus each user 12 sees a portion of the data network according to what access they have themselves.

[0052] In a conventional approach, data change auditing could theoretically be implemented in application code as a general feature for all data within an application, but this is in fact extremely rare. Typically, developers of the application 10 build separate "audit log" tables to record changes to specific datasets of interest. However, changes are not universally captured for all datasets, and often cannot be recovered through a systematic approach when the audit logs are available. In the system described herein, the platform is configured to perform automatic data versioning management of individual records by virtue of the ability to roll back to a previous version. This not only reduces the effort of application development and is necessarily applied to all data, without compromising on effort cost / impact. Automatic data versioning management also simplifies the data model by avoiding user-defined control attributes (such as creation time, creator, etc.).

[0053] The automatic data version management applied by the platform's data version management module can be done by remembering all data changes so that both previous versions and differences between versions can be displayed, and by having the ability to roll back to a specific version by reapplying the changes even when the table schema has changed.

[0054] In the conventional approach, the ability to limit who can view and edit data is implemented in application-specific code by each application 10 (see Figure 1). This application-specific code is written at the application / function / feature level, not at the data layer. As described above and shown in Figures 5 and 6, the platform has data access control defined at the data node level, and the solution is forced to automatically comply with these controls. As shown in Figure 9, since the execution of the script language queries used by the platform utilizes access control metadata for executing the query engine, the query results pulled from the data network are limited to what that particular user has access to.

[0055] This data layer access control reduces the effort of application development by eliminating the need to create access control for each individual application. Also, there is consistent control enforcement across all access channels (e.g., APIs, UIs, etc.) (e.g., a single user accessing the same data through multiple applications).

[0056] In the conventional approach, the links between records in database 24 used by application 10 are implemented by copying column values and / or using surrogate keys. In the system described herein, the platform provides the ability to link records of one node to one or more records of another node, regardless of attributes and attribute values. The platform also provides the ability to use links in queries. This simplifies the data model by unifying the physical and logical models and avoids dependence on manually defined surrogate keys. Linking done using the platform also simplifies queries by avoiding what would normally be complex joins.

[0057] Linking can be done by a platform that stores links between records separately from user-defined columns. Separate tables containing mappings of these relationships may be used in a manner not tied to user-defined columns. It is understood that this linking mechanism is for illustrative purposes only and that the implementation will vary depending on the underlying type of persistence.

[0058] The query engine can execute the platform's scripting language by generating an underlying persistence language, such as SQL, and, if applicable, decompose the "dot" notation (shown in FIG. 9) to convert from model to logic. The data can then be converted from logic to physical, and access control is applied so that only approved data is pulled from the underlying persistence. Thereafter, the native query of the underlying persistence can be performed and returned.

[0059] Conventional approaches to filtering data access by user privileges rely on application-specific insulation layers for security and control on the data integration interface, the persistence interface, and the public interface on the physical database layer. Thus, the ability to limit who can view and edit data is implemented within each application by application-specific code written at the application function / feature level. Such conventional approaches are not applicable to every application and require a huge amount of time-consuming application development effort.

[0060] To address this conventional approach, the platform described herein provides consistent control enforcement across all access channels (e.g., APIs and UIs) (e.g., a single user accessing the same data through multiple applications) and defines data access control in the data layer so that application development time can be significantly reduced while eliminating the risk of inappropriate access.

[0061] It has been found that inserting caches of metadata and entitlement data and incorporating dedicated processing modules and tables to capture entitlements of separate tables within the information exchange layer can accelerate performance. The cache protocol, which will be described in more detail later, can execute the cached transformed query or regenerate the query in response to changes related to the access rights or the query itself, and the changes are reflected immediately or almost instantaneously. Thus, the possibility of improper data access can be eliminated while new usage privileges are being applied.

[0062] The following describes the process of data-level access control by managing entitlements within the information exchange layer and applying them via a query engine to rewrite queries "on the fly" taking into account access rights. This can be done while caching metadata and entitlement data and incorporating dedicated processing modules and tables for capturing entitlements into separate tables within the information exchange layer.

[0063] Change Approval Another issue addressed by the platform is to design an effective underlying data structure that enables the platform to seamlessly implement the change approval process without affecting performance. It was recognized that application queries should calculate results based only on approved database changes while tracking and versioning changes awaiting approval. Existing change approval techniques in conventional approaches have been found not to be designed to handle the very unpredictable usage patterns that the platform faces and are therefore determined to be too complex or too low-performing for the operation of the platform described herein. Also, such existing techniques have been found to be more suitable for fixed database schema designs, but the platform and its one or more data collaboration networks are constantly evolving.

[0064] It has been found that managing two separate tables, a table specialized for tracking unapproved changes and another table that is the master table itself which is approved, can be used for change approval. The processes within this two-table architecture were executed to compare and identify specific fields where changes were made between the master table and the unapproved change table, including calculating on-the-fly changes in the application layer and changes outside the information exchange layer. This led to a process based on a persisting flag for identifying changes for each master data column, delivering acceptable performance and storage requirement characteristics and being built as part of the overall platform architecture.

[0065] Thus, herein, a two-table data structure is provided for managing two separate tables, a table specialized for tracking unapproved changes and another table that is the master table itself which is approved, in order to effectively version manage and track unapproved database changes. This may include the aforementioned persisting flag for identifying changes for each master data column without adversely affecting performance and storage requirement characteristics.

[0066] Data-driven Entitlement In addition to user / group-based column-level entitlement that can be consistently executed by the database across all access methods (e.g., APIs or UIs), the platform described herein may also be configured to enable data-driven entitlement. It is unconditionally understood that the data will still need to be fragmented and replicated. For example, even if a person wants to see the job titles and names of all employees in a company, that person can only see their own phone number and address, and only that person and their supervisor can see that person's salary. For this information to exist in the same table, the platform would apply separate conditions that would enable it to control access based on the data within the table.

[0067] Column-level entitlements can be applied by creating an information exchange layer and rewriting the query before sending it to the backend. This is a combination of the user's permissions with the query the user is executing. While the "where clause" of the user's query can be enriched to include any additional conditions, in some cases, this may mean completely reworking the query to include a where clause, for example when performing an update. However, this may be insufficient for data-driven entitlement. This is because the user's entitlement may be based on data that the user does not have access to. The platform can execute the entire query within the context of that user. For example, if a user only has access to Name and Title in tables of Name, Title, Phone Number, and Address, the platform can apply that user's permissions and restrict the data the user retrieves to Name and Title. According to data-driven entitlement, when a user does not have access to the address field, the platform can allow the user to view all employees whose address is in a specific state or region.

[0068] The platform can also be configured to combine and hierarchically organize multiple entitlements that affect different columns (e.g., can see one's own name, title, phone number, and address, but can only see the names and titles of other employees). By dynamically rewriting the where clause, the platform may not be able to separate the conditions into individual columns. Therefore, the platform's rewrite logic can be extended to adapt the positioning of the conditions when the conditions fall into this category.

[0069] When there are links between tables, it was also recognized that in addition to controlling the current dataset, control of the linked tables should be applied. Since a link can point to other links and thus potentially contain more than one additional set of conditions, this can become quite complex. Also, the platform supports multiple selections within the platform to enable a one-to-many relationship. In an environment where the user has access to a subset of the data, the platform can be configured to ensure that when there are multiple selections, the user can only see what they are permitted to view. This can further complicate the rewrite logic for dynamically accounting for these conditions.

[0070] To enable data-driven scenarios, the platform can also be configured to allow users within entitlement conditions to access information about the current user and which groups that user is a member of. This can be done by extending the query language to support such functions.

[0071] To address potential performance issues (due to the complexity added by these controls to the parsing of each request and the final queries issued against the underlying database), the platform can also implement a custom cache layer that can reduce the number of times statements are processed by the platform's query engine.

[0072] Thus, the platform provides a process for data-level access control to enable data-driven entitlements by operating rewritten queries through system users rather than the current user's credentials, and can hierarchically combine multiple entitlements.

[0073] Data Synchronization / Connector Architecture Extract Transform Load (ETL) tools generally provide components for inserting new data or executing scripts to clear existing data, but the platform described in this specification attempts to create data synchronization that preserves the version history by applying deltas while leaving the existing data intact. The platform is intended to do this as if the data synchronization architecture were independent of the source or target. In this way, simple connectors can be constructed to enable the creation of new connectors for an interface in a shorter time, e.g., within a period of 1-2 days, regardless of the platform, and to make the synchronization operate in a consistent manner and have consistency in the features it exposes.

[0074] The first step in implementing synchronization is to establish reconciliation logic. The platform can be configured to implement partitioning, which has been found to be relatively effective by creating a custom algorithm that depends on custom indexing and sorting strategies. Then, it was found that the bottleneck shifted to the serialization and deserialization of data during transfer between the platform's synchronization utility and the web application. Various serialization protocols can be used, such as Protocol Buffers provided by Google. Protocol Buffers are typically recognized to be aligned with a fixed payload structure, but according to the platform, the data being exchanged does not conform to a fixed schema. Therefore, the platform was configured to adapt the dynamic structure to its type of model.

[0075] Therefore, the platform can extract the source and target from the synchronization engine. That is, when adding a new connector, the platform only needs to have code written to convert the data into a standard intermediate format, rather than implementing a solution for each combination of source and target.

[0076] Network growth Figures 13a through 13c illustrate the development of the dataset network 34 over time. Specifically, Figure 13a shows the 34 dataset network at an initial stage where data is still being developed. At the initial stage, there are fewer nodes 36 and fewer links 40 between the nodes 36. As the dataset is developed, the user adds additional data to the dataset network 34, growing the network. Figure 13b shows the dataset network 34 at a stage later than the initial stage. At this stage, the dataset network has more nodes 36 and more links 40 between the nodes. Figure 13c shows the dataset network at a stage even later than the stage shown in Figure 13b. At this stage, the dataset network has a large number of nodes 36 and a large number of links 40 between the nodes 36.

[0077] Figures 14 and 15 provide schematic views of a conventional application versus the dataset network 34. Figures 14 and 15 show a conventional application (or "app") that includes a UI 14, an API 15, Logic, control 16, and persistence 20 and is integrated with the operating system 18 of the computing device. On the other hand, the dataset network 34 includes nodes 36 that are self-describing, self-connecting, and self-protecting. Thus, there is no need for integration with the UI 14, API, or operating system 18. The API 41 or application experience 32 directly retrieves data 25 from the data network 34 via a query interface 33. This eliminates the need to create, import, update, and manage separate databases. This also eliminates the need to manage the security and control system 16, data integration system 18, data persistence system 20, and data publishing system 22 for each application.

[0078] The dataset network 34 does not require additional security and control functions 16, data integration functions 18, data persistence functions 20, and data publication functions 22. Since the query interface 33 is available, it is also optional to create a custom user interface for exchanging data and information.

[0079] Therefore, it will be understood that any number of dataset nodes 36 as well as application experiences can be added to the network 34. As newer applications are added, the data network grows and newer links are formed between the new datasets.

[0080] FIG. 16A is a schematic diagram showing the security and control 16 of conventional application development using siloed databases. FIG. 16B is a schematic diagram showing the security and control via the access layer 39 of application development using the data network 34. In FIG. 16B, the data 26 is accessed through the user's credentials rather than through the application's service account as shown in FIG. 16A. The security and control 39 of the data network 34 are defined for each dataset 26 rather than in each application 10 as has been conventionally done. This enables true, application - spanning security and control over the data and eliminates data replication. The user 12 can also optionally directly exchange data and information through the data network user interface 42 rather than always through the application.

[0081] For simplicity and clarity of explanation, where appropriate, reference numerals may be repeated in multiple figures to indicate corresponding or similar elements. Also, many specific details are set forth in order to provide a thorough understanding of the examples described herein. However, one of ordinary skill in the art will understand that the examples described herein may be practiced without these specific details. In other instances, well-known methods, procedures, and components are not described in detail so as not to obscure the examples described herein. Also, the description should not be considered as limiting the scope of the examples described herein.

[0082] It will be understood that the examples and corresponding figures used herein are for illustrative purposes only. Different configurations and terminology may be used without departing from the principles set forth herein. For example, components and modules may be added, deleted, modified, or arranged with different connections without departing from these principles.

[0083] Also, any module or component that executes instructions, as exemplified herein, may include or have access to a computer-readable medium, such as a storage medium, a computer storage medium, or a (removable and / or non-removable) data storage device such as a magnetic disk, an optical disk, or a tape. A computer storage medium may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by an application, a module, or both. Any such computer storage media may be part of a platform, may be any or any component related to a platform, or may be accessible or connectable to a platform. Any application or module described herein may be implemented using computer-readable / executable instructions that may be stored or held by such a computer-readable medium.

[0084] The steps or acts in the flowcharts and diagrams described herein are merely examples. There may be many variations to these steps or acts without departing from the principles described above. For example, the steps may be performed in a different order, or steps may be added, deleted, or modified.

[0085] Although the above principles have been described with reference to specific examples, various modifications will be apparent to those skilled in the art, as set forth in the appended claims.

Claims

1. A computer-implemented system for creating, managing, and storing a network of data nodes, comprising: a first dataset comprising at least one data record; access control rules defining, for each of a plurality of users, the right to view or modify each of the data records within the first dataset; a dataset of first structured metadata defining the characteristics of each of the data records within the first dataset; a first node having the same; a subsequent dataset comprising at least one additional data record; subsequent access control rules defining, for each of a plurality of users, the right to view or modify each of at least one data record within the subsequent dataset; a dataset of structured metadata defining the characteristics of each of at least one data record within the subsequent data; at least one subsequent node having the same; comprising: To create the network of data nodes, one or more links are created that associate the first structured metadata within the first node with the structured metadata within the subsequent node, The link associates the first structured metadata of at least one data record of the first node with the subsequent structured metadata of at least one additional data record and is stored as structured metadata, At least one of a plurality of applications communicates with a query layer to request data and additional data within the network, The query layer facilitates the retrieval of data from the first dataset and at least one of the subsequent datasets for use by a plurality of applications, Thereby eliminating the need for data silos and data access control by multiple applications, A computer-implemented system.

2. The system of claim 1, wherein the access control rules comprise data level entitlements such that the rules are used to specify a particular level of entitlement for the data given to the user.

3. The system of claim 2, further comprising a graphical user interface for directly interacting with the data.

4. The system of claim 2, further comprising one or more connectors for linking legacy data to the network of the data nodes.

5. The system of claim 1, wherein the query layer exchanges information with the network of the data nodes via a metadata-driven API and a metadata-driven UI.

6. The system of claim 1, configured to perform automatic data version management of the source of the data and having the ability to roll back to a previous version of the data set.

7. The system of claim 1, wherein the data and the additional data are version-managed.

8. The system of claim 1, wherein the access control rules are defined at the source of the data such that the API and the UI are forced to automatically comply with the access control rules.

9. The system of claim 1, wherein the network of the data nodes can be connected to the network of at least one other data node to create a super network.

10. A computer-implemented method of creating, managing, and storing a network of data nodes, comprising: a first data set including at least one data record; access control rules defining, for each of a plurality of users, the right to view or modify each of the data records in the first data set; a data set of first structured metadata defining the characteristics of each of the data records in the first data set; a first node having the above; a subsequent data set including at least one additional data record; subsequent access control rules defining, for each of a plurality of users, the right to view or modify each of at least one data record in the subsequent data set; a data set of structured metadata defining the characteristics of each of at least one data record in the subsequent data; at least one subsequent node having the above; providing; creating one or more links associating the first structured metadata in the first node with the structured metadata in the subsequent node to create the network of the data nodes; comprising. The link associates the first structured metadata of at least one data record of the first node with the subsequent structured metadata of at least one further data record, and is stored as structured metadata, provides a query layer for exchanging information between the first data set, at least one of the subsequent data sets, and a plurality of applications, at least one of the plurality of applications communicates with the query layer to request data and further data within the network, the query layer facilitates the acquisition of data from the first data set and at least one of the subsequent data sets for use by a plurality of applications, thereby eliminating the need for data silos and data access control by a plurality of applications, A computer-implemented method.

11. The method of claim 10, wherein the access control rules comprise data level entitlements such that the data is used to specify a particular level of entitlement provided to the user.

12. The method of claim 11, wherein the data and the further data are version-managed.

13. The method of claim 11, further comprising one or more connectors for linking legacy data to the network of data nodes.

14. The method of claim 10, wherein the query layer exchanges information with the network of data nodes via a metadata-driven API and a metadata-driven UI.

15. The method of claim 10, configured to perform automatic data version management of the source of the data and having the ability to roll back to a previous version of the data set.

16. The method of claim 10, comprising a two-table architecture comprising a first table for tracking unapproved changes and a second table for managing approved changes.

17. The method of claim 10, wherein the access control rules are defined such that at the source of the data, the API and UI are forced to automatically comply with the access control rules.

18. The method of claim 10, wherein the network of data nodes can be connected to the network of at least one other data node to create a supernetwork.

Citation Information

Patent Citations

  • Server device, information offering method, and program

    JP2005122493A

  • Platform for data service between different application frameworks

    JP2006244488A

  • System and method for aggregating information asset metadata from multiple heterogeneous data management systems

    JP2017514259A

  • A system and method for supporting multiple partition editing sessions in a multi-tenant application server environment.

    JP2017519307A