Sharing replicas between remote deployments
By using a data exchange system on a cloud computing platform and leveraging the permission metadata and shared refresh mechanism of shared objects, the problem of low data sharing efficiency in existing technologies is solved, enabling secure, fast, and scalable data sharing and reducing the sharing costs for small and medium-sized entities.
Patent Information
- Application Number
- CN202180001590.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-05
- Filing Date
- 2021-04-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-04-27
AI Technical Summary
Existing data sharing methods are inefficient and make it difficult to achieve secure and rapid sharing of large-scale datasets. They are particularly costly and complex for small and medium-sized entities, and traditional methods cannot achieve real-time access and control of data.
The data exchange system on the cloud computing platform provides centralized data storage and management, allows data providers to control data sharing permissions, and enables data access across remote deployments through permission metadata of shared objects, avoiding data copying and transmission, and achieving efficient permission synchronization through a shared refresh mechanism.
It enables secure, fast, and scalable data sharing, reduces the data sharing costs for small and medium-sized entities, ensures real-time access and control of data, and improves the efficiency and reliability of data sharing.
Smart Images

Figure CN113994319B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefit under 35 U.S.C. §119(e) of U.S. patent application No. 17 / 193,192, filed on March 5, 2021, which is a continuation of U.S. patent application No. 16 / 900,840, filed on June 12, 2020, and now U.S. Patent No. 10,949,402, issued on March 16, 2021, which claims the benefit of U.S. Provisional Patent Application No. 63 / 030,267, filed on May 26, 2020, the disclosures of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to a data sharing platform, and more particularly, to sharing data between remote deployments thereof.
[0004] background
[0005] Databases are widely used for data storage and access in computing applications. A database may include one or more tables that contain or reference data that can be read, modified, or deleted using queries. Databases can be used to store and / or access personal or other sensitive information. Secure storage and access to database data can be provided by encrypting and / or storing data in encrypted form to prevent unauthorized access. In some cases, data sharing may be useful to enable other parties to perform queries on a set of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The embodiments described and their advantages may be best understood by referring to the following description in conjunction with the accompanying drawings, which in no way limit any changes in form and detail that may be made to the embodiments described by those skilled in the art without departing from the spirit and scope of the embodiments described.
[0008] Figure 1A is a block diagram depicting an example computing environment in which the methods disclosed herein may be implemented.
[0009] Figure 1B is a block diagram illustrating an example virtual warehouse.
[0010] Figure 2 is a schematic block diagram of data that can be used to implement a public or private data exchange according to an embodiment of the present invention.
[0011] Figure 3 is a schematic block diagram of components for implementing a data clearinghouse according to an embodiment of the present invention.
[0012] Figure 4 is a block diagram of remote deployment in a data clearinghouse according to some embodiments of the present invention.
[0013] Figure 5 is a diagram illustrating shared permissions refresh according to some embodiments of the present invention.
[0014] Figure 6A and 6B is a diagram illustrating subsequent sharing permission refresh according to some embodiments of the present invention.
[0015] Figure 7 is a flow diagram of a method for copying a shared object to a remote deployment according to some embodiments of the present invention.
[0016] Figure 8 is a flow chart of a method for performing a shared license refresh operation according to some embodiments of the present invention.
[0017] Figure 9 is a block diagram of an example computing device that can perform one or more operations described herein, according to some embodiments of the present invention.
[0018] Detailed description
[0019] Data providers often possess data assets that are difficult to share. These data assets can be data of interest to another entity. For example, a large online retail company might have a dataset that includes the purchasing habits of millions of customers over the past decade. This dataset can be enormous. If the online retailer wishes to share all or part of this data with another entity, it may need to use older, slower methods to transfer the data, such as File Transfer Protocol (FTP), or even copy the data onto physical media and mail the media to the other entity. This has several disadvantages. First, it's slow. Copying terabytes or petabytes of data can take days. Second, once the data is delivered, the sharer has no control over what happens to the data. The recipient can change the data, make copies, or share it with other parties. Third, the only entities interested in accessing such large datasets in this way are large companies that can afford the complex logistics of transferring and processing the data, as well as the high price of such cumbersome data transfers. Consequently, smaller entities (such as mom and pop stores) or even smaller, more nimble cloud-centric startups are often priced out of accessing this data, despite the potential value of the data to their business. This may be because raw data assets are often too crude and filled with potentially sensitive data to be sold directly to other companies. Before data can be shared with another party, data owners need to perform data cleansing, de-identification, aggregation, joins, and other forms of data enrichment. This is both time-consuming and costly. Finally, for the reasons mentioned above, traditional data sharing methods do not allow for scalable sharing, making it difficult to share data assets with many entities. Traditional sharing methods also introduce wait times and delays in accessing recently updated data for all parties.
[0020] Private and public data clearinghouses can make it easier and more secure for data providers to share their data assets with other entities. Public data clearinghouses (also referred to herein as the "Snowflake Data Marketplace" or "Data Marketplace") can provide a centralized repository with open access where data providers can publish and control real-time and read-only datasets to thousands of customers. Private data clearinghouses (also referred to herein as "Data Clearinghouses") can be branded by the data provider, and the data provider can control who has access to them. Data clearinghouses may be for internal use only or open to customers, partners, suppliers, or others. Data providers can control which data assets are listed and who has access to which datasets. This allows for seamless discovery and sharing of data within the data provider's organization and with its business partners.
[0021] Data clearinghouses can be accessed through cloud computing services such as SNOWFLAKE TM) to facilitate and allow data providers to offer data assets in a private online marketplace under their own brand directly from their own online domain (such as a website). A data clearinghouse can provide entities with a centralized, managed hub to list data assets shared internally or externally, stimulate data collaboration, and also maintain data governance and audit access. With a data clearinghouse, data providers can share data without duplicating data between companies. Data providers can invite other entities to view their data listings, control which data listings appear in their private online marketplace, control who can access the data listings and how others can interact with the data assets connected to the listings. This may be thought of as a "walled garden" marketplace, where visitors to the garden must be approved and access to certain listings may be restricted.
[0022] For example, Company A may be a consumer data company that collects and analyzes the spending habits of millions of individuals across several different categories. Their datasets may include data on online shopping, video streaming, electricity consumption, car usage, internet usage, clothing purchases, mobile app purchases, club memberships, and online subscription services. Company A may wish to make these datasets (or subsets or derivatives of these datasets) available to other entities. For example, a new clothing brand may want access to datasets related to consumer clothing purchases and online shopping habits. Company A can host a page on its website that is substantially similar to, or functions substantially similar to, a data clearinghouse, where data consumers (e.g., new clothing brands) can browse, explore, discover, access, and potentially purchase datasets directly from Company A. Furthermore, Company A can control who can access the data clearinghouse, the entities that can view specific listings, the actions that entities can take with respect to listings (e.g., view only), and any other appropriate actions. Furthermore, data providers can combine their own data with other datasets from, for example, a public data clearinghouse (also known as the "Snowflake Data Marketplace" or "Data Marketplace") and create new listings using the combined data.
[0023] Data clearinghouses may be a suitable venue for discovering, compiling, cleaning, and enriching data to make it more monetizable. A large company in a data clearinghouse might pool data from its various departments and agencies, which could be valuable to another company. Furthermore, participants in a private ecosystem data clearinghouse can work together to connect their datasets and collectively create a useful data product that no single participant could produce alone. Once these connected datasets are created, they can be listed on a data clearinghouse or data marketplace.
[0024] Data sharing can be performed when a data provider creates a shared object of a database (hereinafter referred to as a share) in the data provider's account and grants shared access rights to specific objects of the database (e.g., tables, secure views, and secure user-defined functions (UDFs)). A read-only database can then be created using the information provided in the share. Access to this database can be controlled by the data provider. A "share" encapsulates all the information required to share data in a database. A share can include at least three pieces of information: (1) privileges granting access rights to the database and schema containing the objects to be shared, (2) privileges granting access rights to specific objects (e.g., tables, secure views, and secure UDFs), and (3) the consumer account of the shared database and its objects. When data is shared, no data is copied or transferred between users. Shares are created through cloud computing service providers (such as SNOWFLAKE TM )’s cloud computing services.
[0025] The data shared by a provider can be described by a manifest defined by the provider in a data exchange or data marketplace. Access control, management, and governance of manifests can be similar for data marketplaces and data exchanges. Manifests can include metadata describing the shared data.
[0026] Shared data can be used to process SQL queries, which may include joins, aggregations, or other analysis. In some cases, data providers can define shares that allow "secure joins" to be performed on the shared data. Secure joins can be performed so that analysis can be performed against the shared data, but the data consumer (e.g., the recipient of the share) cannot access the actual shared data.
[0027] In a data clearinghouse, many requests for inventory come from remote deployments. To grant data access to an inventory, the inventory owner needs to have another account or partner account in the remote deployment and create an identical share. This identical share needs to have the same privileged access to the same database (newly created or copied) that was granted access in the original share in the inventory owner's deployment.
[0028] Recreating and maintaining shares across one or more remote deployments is tedious and error-prone. Users need to manually log in to their accounts in each deployment and perform a set of operations to create local shares. Users must provide all permissions of the original share to each local share. If the original share has a large number of permissions on various objects on a database containing hundreds or thousands of objects, users will need to perform hundreds or thousands of operations to recreate the original share. Any edits to the share (such as adding an account) also require a set of operations (such as ALTER operations) to be performed on each copy in each deployment.
[0029] The systems and methods described herein provide an efficient way to replicate shared objects in a data clearinghouse. For example, the method may include modifying a shared object in a first account of the data clearinghouse to a global object, wherein the shared object includes permission metadata indicating sharing permissions for a set of objects in a database. The method may also include creating a local copy of the shared object on a remote deployment based on the global object in a second account of the data clearinghouse, wherein the second account is located in the remote deployment. The set of objects in the database can be replicated to the local database copy on the remote deployment, and the sharing permissions can be replicated to the local copy of the shared object.
[0030] Figure 1A is a block diagram of an example computing environment 100 in which the systems and methods disclosed herein may be implemented. Specifically, a cloud computing platform 110, such as AMAZON WEB SERVICES TM (AWS), Microsoft Azure TM 、GOOGLECLOUD TM As is known in the art, the cloud computing platform 110 provides computing resources and storage resources that can be acquired (purchased) or leased and configured to execute applications and store data.
[0031] The cloud computing platform 110 may host cloud computing services 112 that facilitate data storage (e.g., data management and access) and analytical functions (e.g., SQL queries, analysis), as well as other computing capabilities (e.g., secure data sharing between users of the cloud computing platform 110) on the cloud computing platform 110. The cloud computing platform 110 may include a three-tier architecture: data storage 140, query processing 130, and cloud services 120.
[0032] The data store 140 can facilitate the storage of data on the cloud computing platform 110 in one or more cloud databases 141. The data store 140 can use a storage service such as AMAZON S3 to store data and query results on the cloud computing platform 110. In certain embodiments, to load data into the cloud computing platform 110, data tables can be horizontally partitioned into large, immutable files that can be similar to blocks or pages in traditional database systems. Within each file, the values of each attribute or column are grouped together and compressed using a scheme sometimes referred to as mixed columnar. Each table has a table header that contains, among other metadata, the offset of each column in the file.
[0033] In addition to storing table data, data store 140 facilitates storage of temporary data generated by query operations (e.g., joins), as well as data included in large query results. This may allow the system to compute large queries without encountering out-of-memory or out-of-disk errors. Storing query results in this way can simplify query processing because it eliminates the need for server-side cursors in traditional database systems.
[0034] Query processing 130 can handle query execution within an elastic cluster of virtual machines (referred to herein as a virtual warehouse or data warehouse). Therefore, query processing 130 can include one or more virtual warehouses 131, which can also be referred to herein as a data warehouse. Virtual warehouse 131 can be one or more virtual machines running on cloud computing platform 110. Virtual warehouse 131 can be a computing resource that can be created, destroyed, or resized at any time as needed. This functionality can create an "elastic" virtual warehouse that can be expanded, shrunk, or shut down based on user needs. Expanding a virtual warehouse includes generating one or more computing nodes 132 of virtual warehouse 131. Shrinking a virtual warehouse includes removing one or more computing nodes 132 from virtual warehouse 131. More computing nodes 132 can result in faster computing times. For example, on a system with four nodes, data loading takes fifteen hours, while on a system with thirty-two nodes, it may only take two hours.
[0035] Cloud services 120 may be a collection of services that coordinate activities across cloud computing services 110. These services tie together all the different components of cloud computing services 110 to process user requests, from login to query dispatch. Cloud services 120 may operate on computing instances provided by cloud computing services 110 from cloud computing platform 110. Cloud services 120 may include a collection of services that manage virtual repositories, queries, transactions, data exchange, and metadata associated with these services (e.g., database schemas, access control information, encryption keys, and usage statistics). Cloud services 120 may include, but are not limited to, an authentication engine 121, a fabric manager 122, an optimizer 123, an exchange manager 124, a security engine 125, and a metadata repository 126.
[0036] Figure 1Bis a block diagram illustrating an example virtual repository 131. An exchange manager 124 can facilitate data sharing between data providers and data consumers using, for example, a data clearinghouse. For example, cloud computing service 112 can manage the storage and access of database 108. Database 108 can include various instances of user data 150 for different users, such as different businesses or individuals. User data can include a user database 152 of data stored and accessed by the user. User database 152 can be access-controlled, allowing only the owner of the data to modify and access database 112 upon authentication to cloud computing service 112. For example, data can be encrypted so that it can only be decrypted using decryption information possessed by the data owner. Using exchange manager 124, specific data from user database 152 subject to these access controls can be shared with other users in a controlled manner according to the methods disclosed herein. In particular, a user can designate a share 154, which can be shared in an uncontrolled manner publicly or in a data clearinghouse, or in a controlled manner with specific other users as described above. A "share" encapsulates all the information required to share data in a database. Sharing can include at least three pieces of information: (1) privileges granting access to the database and schema containing the objects to be shared, (2) privileges granting access to specific objects (e.g., tables, secure views, and secure UDFs), and (3) the consumer account of the shared database and its objects. When sharing data, no data is copied or transferred between users. Sharing is achieved through cloud service 120 of cloud computing service 110.
[0037] Data sharing is performed when a data provider creates a share of a database in the data provider's account and grants access to specific objects, such as tables, secure views, and secure user-defined functions (UDFs). A read-only database can then be created using the information provided in the share. Access to this database can be controlled by the data provider.
[0038] The shared data can be used to process SQL queries, which may include joins, aggregations, or other analytics. In some cases, the data provider can define a share that allows a "secure join" to be performed on the shared data. A secure join can be performed so that analytics can be performed on the shared data, but the data consumer (e.g., the recipient of the share) cannot access the actual shared data. A secure join can be performed as described in U.S. application serial number 16 / 368,339, filed on March 18, 2019.
[0039] User devices 101-104, such as laptops, desktop computers, mobile phones, tablet computers, cloud-hosted computers, cloud-hosted serverless processes, or other computing processes or devices, can be used to access the virtual warehouse 131 or cloud service 120 via a network 105, such as the Internet or a private network.
[0040] In the following description, actions are attributed to users, specifically consumers and providers. Such actions should be understood to be performed on devices 101-104 operated by such users. For example, notifications to users can be understood as notifications sent to devices 101-104, input or instructions from users can be understood as received through the user's devices 101-104, and user interactions with interfaces should be understood as interactions with interfaces on the user's devices 101-104. In addition, database operations (connections, aggregations, analysis, etc.) attributed to users (consumers or providers) should be understood to include cloud computing services 110 performing such actions in response to instructions from the user.
[0041] Figure 2 2 is a schematic block diagram of data that can be used to implement a public or data clearinghouse, according to an embodiment of the present invention. Exchange manager 124 can operate on some or all of the illustrated exchange data 200, which can be stored on a platform executing exchange manager 124 (e.g., cloud computing platform 110) or at some other location. Exchange data 200 can include a plurality of manifests 202 describing data shared by a first user ("provider"). Manifests 202 can be manifests in a data clearinghouse or a data marketplace. Access control, management, and governance of manifests can be similar for data marketplaces and data clearinghouses.
[0042] The listing 202 may include metadata 204 describing the shared data. The metadata 204 may include some or all of the following information: an identifier of the sharer of the shared data, a URL associated with the sharer, the name of the share, the name of the table, the category to which the shared data belongs, how often the shared data is updated, a table of contents, the number of columns and rows in each table, and the names of the columns. The metadata 204 may also include examples to help users use the data. Such examples may include an example table including examples of rows and columns of the example table, example queries that can be run against the table, example views of the example table, and example visualizations (e.g., graphs, dashboards) based on the table data. Other information included in the metadata 204 may include metadata for use by business intelligence tools, textual descriptions of the data contained in the table, keywords associated with the table to facilitate searching, links to documents related to the shared data (e.g., URLs), and a refresh interval indicating how often the shared data is updated and the date the data was last updated.
[0043] Manifest 202 may include access controls 206, which may be configured as any suitable access configuration. For example, access controls 206 may indicate that shared data is available to any member of the private clearinghouse without restriction ("any share" as used elsewhere herein). Access controls 206 may specify a class of users (members of a particular group or organization) that are permitted to access data and / or view the manifest. Access controls 206 may specify "peer-to-peer" sharing (see Figure 4 ), where a user can request access but is only granted access after the provider approves it. Access control 206 can specify a set of user identifiers of users that are excluded from being able to access the data referenced by manifest 202.
[0044] Note that some listings 202 can be discovered by the user without further authentication or access rights, with actual access being allowed only after a subsequent authentication step (see Figure 4 and 6). Access control 206 may specify that manifest 202 is discoverable only by certain users or classes of users.
[0045] Also note that the default functionality of manifest 202 is that the data referenced by a share cannot be exported by consumers. Optionally, access control 206 can specify that this is not permitted. For example, access control 206 can specify that security operations (such as secure connections and security functions described below) can be performed on the shared data, such that viewing and exporting the shared data are not permitted.
[0046] In some embodiments, once a user is authenticated with respect to manifest 202, a reference to the user (e.g., a user identifier for the user's account in virtual warehouse 131) is added to access control 206, enabling the user to subsequently access the data referenced by manifest 202 without further authentication.
[0047] The list 202 can define one or more filters 208. For example, the filter 208 can define a specific user identifier 214 of a user who can view references to the list 202 when browsing the directory 220. The filter 208 can define a class of users (users of a specific profession, users associated with a specific company or organization, users within a specific geographic region or country) who can view references to the list 202 when browsing the directory 220. In this way, a private exchange can be implemented by the exchange manager 124 using the same components. In some embodiments, an excluded user who is excluded from accessing the list 202 (i.e., adding the list 202 to the excluded user's consumption share 156) can still be allowed to view representations of the list when browsing the directory 220, and can also be allowed to request access to the list 202, as described below. Requests to access the list by these excluded users and other users can be listed in the interface presented to the provider of the list 202. The provider of the manifest 202 can then review the requirements for access to the manifest and choose to extend the filter 208 to allow access to excluded users or excluded categories of users (eg, users in excluded geographic regions or countries).
[0048] Filters 208 can further define what data a user can view. In particular, filters 208 can indicate that a user who selects a listing 202 to add to the user's consumption share 156 is permitted to access the data referenced by the listing, but only a filtered version of the data, the filtered version including only data associated with the user's identifier 214, associated with the user's organization, or some other classification specific to the user. In some embodiments, a private exchange is by invitation: a user invited by a provider to view a listing 202 for a private exchange can do so through the exchange manager 124 upon communicating acceptance of the invitation received from the provider.
[0049] In some embodiments, the manifest 202 may be addressed to a single user. Thus, a reference to the manifest 202 may be added to a set of "pending shares" visible to the user. Once the user communicates approval to the exchange manager 124, the manifest 202 may be added to the user's set of shares.
[0050] The inventory 202 may further include usage data 210. For example, the cloud computing service 112 may implement a credit system in which credits are purchased by a user and consumed each time the user runs a query, stores data, or uses other services implemented by the cloud computing service 112. Thus, the usage data 210 may record the amount of credits consumed by accessing shared data. The usage data 210 may include other data, such as multiple queries, multiple aggregations of each type of data executed against the shared data, or other usage statistics. In some embodiments, the usage data for the user's inventory 202 or multiple inventories 202 is provided to the user in the form of a shared database, i.e., a reference to the database containing the usage data is added to the user's consumption share by the exchange manager 124.
[0051] Listing 202 may also include a heat map 211, which may indicate the geographic location of users clicking on that particular listing. Cloud computing service 110 may use the heat map to make replication or other decisions about the listing. For example, a data exchange may display a listing containing weather data for Georgia, USA. Heat map 211 may indicate that many users in California are selecting the listing to learn more about the weather in Georgia. Given this information, cloud computing service 110 may replicate the listing and make it available in a database whose server is physically located in the western United States, so that consumers in California can access the data. In some embodiments, an entity may store its data on a server located in the western United States. A particular listing may be very popular with consumers. Cloud computing service 110 may replicate the data and store it in a server located in the eastern United States, so that consumers in the Midwest and the East Coast can also access the data.
[0052] Manifest 202 may also include one or more tags 213. Tags 213 may facilitate easier sharing of data contained in one or more manifests. For example, a large company may have a human resources (HR) manifest in a data clearinghouse that contains HR data for its internal employees. The HR data may contain ten types of HR data (e.g., employee number, selected health insurance, current retirement plan, job title, etc.). The HR manifest may be accessed by 100 people in the company (e.g., everyone in the HR department). Management in the HR department may want to add an eleventh type of HR data (e.g., employee stock option plans). Instead of manually adding the new data set to the HR manifest and granting the 100 people access to the new data, management may simply apply the HR tag to the new data set, classify the new data as HR data, list it with the HR manifest, and grant the 100 people permission to view the new data set.
[0053] The manifest 202 may also include version metadata 215. The version metadata 215 may provide a way to track how a dataset has changed. This may be helpful in ensuring that data that one entity is viewing does not change prematurely. For example, if a company has an original dataset and then releases an updated version of that dataset, the update may interfere with another user's processing of the dataset because the update may have a different format, new columns, and other changes that may be incompatible with the receiving user's current processing mechanisms. To remedy this, the cloud computing service 112 may use version metadata 215 to track version updates. The cloud computing service 112 may ensure that each data consumer accesses the same version of the data until they accept the updated version that does not interfere with the current processing of the dataset.
[0054] Exchange data 200 may further include user record 212. User record 212 may include data identifying a user associated with user record 212, such as an identifier of a user having user data 150 in service database 128 and managed by virtual repository 131 (eg, a repository identifier).
[0055] A user record 212 may list shares associated with the user, such as a reference list 202 created by the user. A user record 212 may list shares consumed by the user, such as a reference list 202 created by another user and associated with the user's account according to the methods described herein. For example, a list 202 may have an identifier that will be used to reference it in a share or consumed share in a user record 212.
[0056] Exchange data 200 may further include a directory 220. Directory 220 may include a list of all available manifests 202 and may include an index of data from metadata 204 to facilitate browsing and searching according to the methods described herein. In some embodiments, manifests 202 are stored in the directory in the form of JavaScript Object Notation (JSON) objects.
[0057] Note that, in the case where there are multiple instances of a virtual warehouse 131 on different cloud computing platforms, the catalog 220 of one instance of a virtual warehouse 131 may store manifests or references to manifests from other instances on one or more other cloud computing platforms 110. Thus, each manifest 202 may be globally unique (e.g., assigned a globally unique identifier across all instances of a virtual warehouse 131). For example, instances of a virtual warehouse 131 may synchronize their copies of the catalog 220 so that each copy indicates a manifest 202 that is available from all instances of the virtual warehouse 131. In some cases, the provider of a manifest 202 may specify that it is only available on a specified one or more computing platforms 110.
[0058] In some embodiments, the catalog 220 is available on the Internet so that it can be searched by search engines such as BING or GOOGLE. The catalog may be subject to search engine optimization (SEO) algorithms to increase its visibility. Potential consumers can therefore browse the catalog 220 from any web browser. The exchange manager 124 can display a uniform resource locator (URL) linked to each listing 202. The URL can be searchable and can be shared outside of any interface implemented by the exchange manager 124. For example, a provider of a listing 202 can publish the URL of its listing 202 to promote the use of its listing 202 and its brand.
[0059] Figure 3 Various components 300-310 that may be included in the exchange manager 124 are shown. The creation module 300 may provide an interface for creating the manifest 202. For example, a web interface to the virtual warehouse 131 enables a user on a device 101-104 to select data, such as a specific table in the user's user data 150, to share and enter values that define some or all of the metadata 204, access controls 206, and filters 208. In some embodiments, the creation may be performed by the user through SQL commands in a SQL interpreter that is executed on the cloud computing platform 110 and accessed through a web interface on the user device 101-104.
[0060] When attempting to create a manifest 202, a verification module 302 can verify the information provided by the provider. Note that in some embodiments, the actions attributed to the verification module 302 can be performed by a human inspecting the information provided by the provider. In other embodiments, these actions are performed automatically. The verification module 302 can perform or facilitate the operator to perform various functions. These functions can include verifying that the metadata 204 is consistent with the shared data it references, verifying that the shared data referenced by the metadata 204 is not pirated data, personally identifiable information (PII), personal health information (PHI), or other data that is not desirable or illegal to share. The verification module 302 can also facilitate verifying that the data has been updated within a threshold time period (e.g., within the last twenty-four hours). The verification module 302 can also facilitate verifying that the data is not static or cannot be obtained from other static public sources. The verification module 302 can also facilitate verifying that the data is not merely a sample (e.g., that the data is complete enough to be useful). For example, geographically restricted data may be undesirable, while otherwise unrestricted data aggregation may still be useful.
[0061] The exchange manager 124 may include a search module 304. The search module 304 may implement a web interface accessible by users on user devices 101-104 to invoke a search for a search string on metadata in the catalog 220, receive responses to the search, and select references to the inventory 202 in the search results to add to the consumption share 156 of the user record 212 of the user performing the search. In some embodiments, the search may be performed by the user via SQL commands in a SQL interpreter executed on the cloud computing platform 102 and accessed via a web interface on the user devices 101-104. For example, the search may be performed via an SQL query against the catalog 220 within the SQL engine 310 discussed below.
[0062] The search module 304 can further implement a recommendation algorithm. For example, the recommendation algorithm can recommend other lists 202 to the user based on other lists in the user's consumption share 156 or in the consumption share of previous users. Recommendations can be based on logical similarity: one weather data source leads to a recommendation for a second weather data source. Recommendations can also be based on dissimilarity: one list targeting one area (geographical region, technical field, etc.) leads to lists in a different area to facilitate comprehensive coverage of the user's analysis (different geographic areas, related technical fields, etc.).
[0063] The exchange manager 124 may include an access management module 306. As described above, a user may add a listing 202. This may require authentication of the provider of the listing 202. Once a listing 202 is added to the consumption share 156 of the user's user record 212, the user may (a) be required to authenticate each time the user accesses the data referenced by the listing 202, or (b) be automatically authenticated and allowed access to the data once the listing 202 is added. The access management module 306 may manage automatic authentication for subsequent access to data in the user's consumption share 156 in order to provide seamless access to the shared data as if it were part of the user's user data 150. To do this, the access management module 306 may access the control 206, certificate, token, or other authentication material of the listing 202 in order to authenticate the user when performing access to the shared data.
[0064] The exchange manager 124 may include a connection module 308. The connection module 308 manages the integration of shared data referenced by the user's consumption share 156 (i.e., shared data from different providers) with each other and with the user database 152 of data owned by the user. In particular, the connection module 308 can manage the execution of queries and other computational functions on these various data sources so that their access is transparent to the user. The connection module 308 can further manage access to the data to enforce restrictions on the shared data, for example, so that analysis can be performed and the results of the analysis displayed without exposing the underlying data to the consumer of the data, where the restrictions are indicated by the access control 206 of the manifest 202.
[0065] The exchange manager 124 may also include a standard query language (SQL) engine 310 that is programmed to receive queries from users and execute the queries against the data referenced by the queries, which may include the user's consumption shares 156 and the user data 112 owned by the user. The SQL engine 310 may perform any query processing functions known in the art. The SQL engine 310 may additionally or alternatively include any other database management tools or data analysis tools known in the art. The SQL engine 310 may define a web interface executed on the cloud computing platform 102, through which SQL queries are entered and responses to the SQL queries are presented.
[0066] Figure 4 A cloud environment 400 is shown that includes multiple remote cloud deployments 401, 402, and 403. Each of the remote deployments 401, 402, and 403 may include a remote cloud deployment 401, 402, and 403. Figure 1A ). Remote deployments 401, 402, and 403 may all be physically located in separate remote geographic regions (hence the term "remote deployments"), but may all be deployments of a single data clearinghouse or data marketplace. In cloud environment 400, a request for an inventory on remote deployment 401 may originate from an account on remote deployment 402. To grant data access to an inventory, the inventory owner needs to have another account or a partner account in remote deployment 402 and create the same share. The same share needs to have the same access privileges to the database (newly created or copied) that was granted the original share in the inventory owner's deployment.
[0067] Although embodiments of the present disclosure are described with respect to a data clearinghouse for ease of description, the systems and methods for shared replication described herein need not be coupled to a data clearinghouse.For example, remote deployments 401-403 may each include user accounts and databases in independent settings.
[0068] As mentioned above, the process of recreating and maintaining shares is tedious and error-prone. Users need to manually log in to their accounts in each deployment and perform a series of operations to create the local share and provide permissions from the original share to the local share. Any edits to the local share (such as adding a table) also require a set of ALTER operations to each copy in each deployment.
[0069] Existing database replication mechanisms do not require the performance of such an extensive set of operations. For example, if account A resides on a remote deployment 401 located in region 1 and has a database DB1 on remote deployment 401 that he wants to share with account B who resides in a remote deployment 402 located in region 2, account A can alter database DB1 so that it is a global type database (as opposed to region specific) and replicate the metadata of DB1 to remote deployment 402 (e.g., by using the command "alter database DB1 enable replication to accounts Reg_2.B"). Account B can obtain a list of databases that they have access to (e.g., using the command "show replication databases"), which will return an identifier "Reg_1.A.DB1 (primary)" indicating DB1. Account B can create a local replica of DB1 on remote deployment 402 (in Figure 4 Account B is displayed as DB1R in the example above) (for example, by using the command "create database DB1R as replica of Reg_1.A.DB1"), which creates a global type database because it is created as a replica. It should be noted that data replication has not yet started. At this point, the "show replication databases" command will return the identifiers "Reg_1.A.DB1 (primary)" and "Reg_2.B.DB1 (secondary)". Account B can start data replication by using a command (for example, "alter database DB1 refresh"), which is a synchronization operation whose duration depends on the amount of data to be synchronized.
[0070] Embodiments of the present disclosure can mimic aspects of the above-described database replication process to implement shared replication without the need to manually create local shares in each deployment and manually provide all permissions on the original share S1 to each local share. For example, account A (residing in a remote deployment 401 located in region 1) has a shared object S1 that has been granted certain privileges ("share grants" or "grants") on database DB1. For example, S1 may have use and selection permissions that allow access to and use of DB1 and the schemas, tables, views, functions, and formats within DB1. S1's shared permissions can be stored in the form of permission metadata. Remote deployment 401 can utilize any appropriate metadata storage, such as FoundationDB, to store S1's permission metadata as well as S1's share metadata and DB1's database metadata. If account A wants to replicate S1 to account B (located in a remote deployment 402 in region 2), account A can change S1 to a global object and replicate S1 to remote deployment 402 (e.g., using the command "alter share S1 enable replication to accounts Reg_2.B"). More specifically, remote deployment 401 can modify S1, which is an object specific to region 1, to include a global representation present in region 2. Remote deployment 401 can set the authorization attribute of S1 to store a list of global accounts that can create local replicas of S1 and include it in list account B. S1 can include an existing "primaryOn" field, which can be set to the timestamp value when account A changed S1 to a global object and replicated S1 to remote deployment 402. In this way, remote deployment 401 can identify S1 as a primary and distinguish it from any secondary replicas. Account B can obtain a list of shared objects owned by him / herself or others in the same replication group for which access rights are enabled (e.g., using the command "show replication shares"), and the list can include an identifier for S1 ("Reg_1.A.S1 (primary)"). In some embodiments, S1 may not need to be primary, and any replica objects created from it can be designated as primary.
[0071] Since there is a global representation of S1 in Region 2, Account B can create a local replica of S1 (for example, using the command "create share S1R as replica of Reg_1.A.S1(primary)"), which creates a global shared object that is a replica of S1 (in Figure 4), where the "primaryOn" field is set to "0" because it is a replica. The local replica of DB1 (DB1R) can be created on the remote deployment 402 using the same process discussed above regarding database replication, which creates DB1R as a global type database because it is created as a replica of DB1, which has already become global. As described above, account B can also obtain a list of databases that it is allowed to access. In some embodiments, account A can use a communication protocol such as SSH to establish a secure connection to account B, through which account A can "remote login" to account B, or initiate a remote execution context in account B to create a replica shared object S1R and a replica database DB1R. More specifically, a global identity can be created that is linked to (e.g., has access to) multiple accounts (in this case, account A and account B). When this global identity is created, a user logged into account A can establish a secure connection to account B, through which account A can "remote login" to account B or initiate a remote execution context in account B. At this point, account A can initiate functions on remote deployment 401 to create a replica shared object S1R and a replica database DB1R, and these functions can then be executed in account B on remote deployment 402. S1R and DB1R can be copied to any suitable object, such as a dictionary object in the local metadata store (e.g., FoundationDB) of the remote deployment 402.
[0072] It should be noted that at this point, the replication of shared permissions (e.g., permission metadata) between S1 and the original database DB1 or the underlying data of database DB1 (e.g., DB1's underlying schema, tables, views, functions, formats, etc.) has not yet occurred. To start replicating S1's shared permissions, account B can initiate a share refresh operation (e.g., using the command "alter shareS1R refresh"). Remote deployment 402 can send a refresh message to remote deployment 401 (in Figure 5 , the remote deployment 401 will copy the necessary information to the remote deployment 402. More specifically, upon receiving the refresh message, the remote deployment 401 can copy the license metadata of S1 to S1R, as shown below. Figure 5Detailed discussion further below. Additionally, in response to a shared refresh message (e.g., which may be part of a shared refresh operation), the database replication process discussed above may be automatically triggered to initiate replication of DB1's underlying data (e.g., DB1's underlying schemas, tables, views, functions, formats, etc.) to DB1R. In some embodiments, the database replication process is separate from the shared refresh operation and is not automatically triggered by a refresh message and may be triggered at the user's discretion. For example, the owner of Account B may choose to trigger a shared refresh as a downstream task of the database replication. In some embodiments, Account A may use any appropriate communication protocol to "remotely log in" to Account B, or initiate a remote execution context in Account B as described herein, and trigger a refresh operation.
[0073] Figure 5 The sharing permissions (e.g., permissions metadata) of S1 are copied to S1R (also referred to as a "sharing refresh operation"). In response to receiving a refresh message ("refresh S1R") from the remote deployment 402, the remote deployment 401 can retrieve all of S1's sharing permissions, serialize all of the sharing permissions, place the serialized sharing permissions into a file, and transmit a message including the file to the remote deployment 402. In some embodiments, the remote deployment 401 can utilize a global messaging framework and utilize a special message type (herein referred to as a "snapshot message") specifically for implementing sharing replication. For each global message type, there is a corresponding processing function suitable for processing messages of that type. Therefore, a specific type of global message will include customized logic for the processing required for that specific message type. The snapshot message type can include information that notifies the deployment that the message includes permissions and allows the remote deployment to appropriately deserialize the permissions for the purpose of sharing replication. More specifically, the body of the message can include a global object reference for S1 and a list of modified permission objects representing the permissions that S1 has on DB1 and its underlying objects. The global object reference may include a tuple of the original deployment ID (e.g., the ID of remote deployment 401) and the ID of the object to be copied (in this example, the ID of shared object S1). Each permission object may be an internal representation of a specific permission for S1 and may be associated with S1 to allow consumers to use / select the underlying database objects defined by the permission. For example, each permission object may have several fields, including: a global object reference to DB1 or one of its underlying objects that is being shared, an object type (e.g., security type), a privilege / permission type (e.g., use or select), and a timestamp when the permission occurred.
[0074] Upon receiving the special message, remote deployment 402 can deserialize the message and apply the sharing permissions to S1R. As can be seen, the sharing refresh described here is based on a pull model (as opposed to a push model). In other words, S1R will only be updated when a sharing refresh is requested, and there will be a difference between S1R and S1 until a sharing refresh of S1R is performed.
[0075] In some embodiments, for example, database replication and share refresh operations can be separate, and database replication can be managed by the owner of Account B. The owner of Account B can control whether / when database replication is initiated. A dedicated API can be used to perform a share refresh. In other embodiments, database replication is hidden from the user and performed automatically as part of a share refresh. In either case, the replication of the share permissions of Share S1 (i.e., the share refresh operation) is performed before the database replication.
[0076] A user of account B in remote deployment 402 can initiate subsequent sharing refresh operations of S1 to S1R to update sharing permissions. Subsequent sharing refreshes may be initiated periodically in response to modifications to the sharing permissions of S1, at the discretion of the owner of account B, or based on any other appropriate criteria. As described above, the sharing refresh operation essentially requires refreshing the sharing permissions (also known as "permission metadata") on S1 (e.g., permissions on DB1's tables, views, functions, formats, etc.) to S1R. This requires pre-copying the underlying objects of DB1. Therefore, when the sharing permissions of S1 are modified (e.g., in response to modifications to DB1), there may be situations where some of the new sharing permissions cannot be copied to S1R during the sharing refresh because the underlying objects corresponding to those new sharing permissions have not yet been copied to DB1R.
[0077] For example, objects may be lost due to time delays between database replication and sharing refresh, and inconsistent database replication and sharing refresh frequencies. This is particularly evident in cases where the sharing refresh and database replication operations are separate. In some embodiments, a partial sharing refresh may be performed if the database replication has not yet completed. For example, it may be important that the sharing refresh task executes normally within a specific time frame, in which case the sharing refresh cannot wait for the database replication to complete. Sharing permissions (e.g., permission metadata) for DB1 objects detected in DB1R at the time of the sharing refresh may be updated, and for objects in DB1R that are no longer in DB1, any permissions in S1R for those objects may be revoked. For those objects in DB1 that are not in DB1R / missing in DB1R at the time of the sharing refresh, or that are included in DB1R but not in DB1 at the time of the sharing refresh, a log may be created indicating the DB1 objects that are missing in DB1R / outside of DB1R, and the updated sharing permissions that cannot be replicated to S1R or should be deleted from S1R.
[0078] In some embodiments, the shared refresh task can follow an "all or nothing" approach. If there is one or more missing objects, the new permissions will not be applied. In these cases, a log can be created that indicates the database objects whose underlying data was not located and for which the updated permissions metadata cannot be replicated.
[0079] By applying the share permissions of the shared object S1 to the replica shared object in the remote deployment, users can avoid manually logging into the account in each remote deployment to which they want to replicate the shared object S1 and performing a large set of operations to create a local share and grant permissions from the original share to the local share. In this way, a more resource-efficient share replication method is provided.
[0080] Figure 6A Some embodiments according to the present disclosure are shown. Figure 4 and Figure 5 After the shared copy operation in question, S1 and S1R and their associated shared permissions. Figure 6A Account A and Account B are shown, along with their associated databases and shares, and the share permissions applied to their associated shares. Figure 6A As shown, S1 has use permissions on DB1, use permissions on schemas A and B, and select permissions on tables T1, T2, T3, and T4 of DB1. S1R has similar use permissions on DB1R after the shared copy operation discussed here.
[0081] like Figure 6BAs shown, table T3 can then be deleted from DB1 and table T5 can be added to DB1, so the select permission corresponding to T3 no longer exists in S1's shared permissions, while T5's select permission has been added to S1's shared permissions. Account B can perform a subsequent shared refresh operation and trigger a subsequent database replication to replicate the new table T5 to DB1R and delete T3 from DB1R. As described above, database replication can be triggered at the discretion of the owner of Account B or can be triggered automatically when a shared refresh operation is initiated.
[0082] exist Figure 6B In the example of , the shared refresh operation may be required to execute normally within a certain time frame, and there may be a time delay between the initiation of the shared refresh and the database copy operation. Therefore, when the shared refresh begins, neither the delete of T3 nor the add of T5 has been replicated to DB1R. During the shared refresh, the remote deployment 402 can determine that T3 is in DB1R but not in DB1, and therefore any permissions for T3 can be revoked from S1R. The remote deployment 402 can then check the select permissions for T5 and determine that object T5 has not yet been replicated to DB1R, and therefore can add T5 and the select permissions for T5 to the log to notify the owner of Account B that they cannot be replicated due to the missing objects. The remote deployment 402 can also issue a notification / message to the owner of Account B indicating that their database DB1R must be refreshed. In this way, once table T5 is created during the database copy, the select permissions for T5 can be replicated to S1R during the subsequent shared refresh operation.
[0083] Figure 7 7 is a flow chart of a method 700 for copying a shared object to a remote deployment according to some embodiments. The method 700 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, a processor, a processing device, a central processing unit (CPU), a system on a chip (SoC), etc.), software (e.g., instructions running / executing on a processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, the method 700 may be performed by remote deployments 401 and 402 (on Figure 4 and 5 ) is executed by the corresponding processing device shown in .
[0084] Also refer to Figure 4, if account A wants to replicate S1 to account B in remote deployment 402 in region 2, account A can change S1 to a global object and replicate S1 to remote deployment 402 (e.g., using the command "alter shareS1 enable replication to accounts Reg_2.B"). More specifically, in block 705, remote deployment 401 can modify S1 (which is an object specific to region 1) to include a global representation present in region 2. Remote deployment 401 can set the authorization attribute of S1 to store a list of global accounts that can create a local replica of S1 and include it in list account B. S1 can include an existing "primaryOn" field that is set to the timestamp of when account A changed S1 to a global object and replicated S1 to remote deployment 402. This identifies S1 as the primary and distinguishes it from any secondary replicas. Account B can obtain a list of shared objects owned by him / herself or others in the same replication group to which access permissions are enabled (e.g., using the command "show replication shares"), and the list can include an identifier for S1 ("Reg_1.A.S1 (primary)". In some embodiments, S1 and any replica objects created thereby can be designated as primary.
[0085] Because a global representation of S1 exists in Region 2, Account B can then create a local replica of S1 in block 710 (e.g., using the command “create share S1R as replica of Reg_1.A.S1(primary)”), which creates a global shared object that is a replica of S1 (in Figure 4, where the "primary" field is set to "0" because it is a replica. The local replica of DB1 (DB1R) can be created on the remote deployment 402 using the same process discussed above regarding database replication, which creates DB1R as a global type database because it is created as a replica of DB1, which has already become global. As described above, account B can also obtain a list of databases that it is allowed to access. In some embodiments, account A can use a communication protocol such as SSH to establish a secure connection to account B, through which account A can "remotely log in" to account B, or initiate a remote execution context in account B to create a replica shared object S1R and a replica database DB1R. More specifically, a global identity can be created that is linked to (e.g., has access to) multiple accounts (in this case, account A and account B). When this global identity is created, a user logged into account A can establish a secure connection to account B, through which account A can "remotely log in" to account B or initiate a remote execution context in account B. At this point, account A can initiate functions on remote deployment 401 to create a replica shared object S1R and a replica database DB1R, and these functions can then be executed in account B on remote deployment 402. S1R and DB1R can be copied to any suitable object, such as a dictionary object in the local metadata store (e.g., FoundationDB) of the remote deployment 402.
[0086] It should be noted that at this point, the replication of shared permissions (e.g., permissions metadata) between S1 and the original database DB1 or the underlying data of database DB1 (e.g., DB1's underlying schema, tables, views, functions, formats, etc.) has not yet occurred. To begin replicating S1's shared permissions, in block 715, account B may initiate a replication of DB1's underlying data (e.g., DB1's underlying schema, tables, views, functions, formats, etc.) to DB1R. In block 720, account B may initiate a share refresh operation (e.g., using the command "alter share S1R refresh"). Remote deployment 402 may send a refresh message (in Figure 5 , remote deployment 401 will copy the necessary information to remote deployment 402. More specifically, upon receiving the refresh message, remote deployment 401 can copy the license metadata of S1 to S1R, as shown below. Figure 8Discussed in further detail. In some embodiments, in response to a shared refresh message (e.g., which can be part of a shared refresh operation), the database replication process discussed above can be automatically triggered to initiate replication of DB1's underlying data (e.g., DB1's underlying schema, tables, views, functions, formats, etc.) to DB1R. In some embodiments, the database replication process is separate from the shared refresh operation and is not automatically triggered by the refresh message, and can be triggered at the user's discretion. For example, the owner of Account B can choose to trigger a shared refresh as a downstream task of the database replication. In some embodiments, Account A can use any appropriate communication protocol to "remotely log in" to Account B, or initiate a remote execution context in Account B as described herein, and trigger the refresh operation.
[0087] Figure 8 800 is a flow chart of a method 800 for refreshing a shared object in a remote deployment according to some embodiments of the present disclosure. The method 800 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, a processor, a processing device, a central processing unit (CPU), a system on a chip (SoC), etc.), software (e.g., instructions running / executing on a processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, the method 800 may be performed by remote deployments 401 and 402 (on Figure 4 and 5 ) is executed by the corresponding processing device shown in ).
[0088] Also refer to Figure 5In response to receiving a refresh message ("Refresh S1R") from the remote deployment 402, in box 805, the remote deployment 401 can retrieve all shared permissions for S1 and serialize all shared permissions. In box 810, the remote deployment 401 can place the serialized shared permissions into a file and send a message including the file to the remote deployment 402. In some embodiments, the remote deployment 401 can utilize a global messaging framework and utilize special message types (referred to herein as "snapshot messages") that are specifically designed to implement shared replication. For each global message type, there is a corresponding processing function suitable for processing messages of that type. Therefore, a specific type of global message will include customized logic for the processing required for that specific message type. The snapshot message type can include information that notifies the deployment that the message includes permissions and allows the remote deployment to appropriately deserialize the permissions for the purpose of shared replication. More specifically, the body of the message can include a global object reference for S1 and a list of modified permission objects representing the permissions that S1 has on DB1 and its underlying objects. The global object reference may include a tuple of the original deployment ID (e.g., the ID of remote deployment 401) and the ID of the object to be copied (in this example, the ID of shared object S1). Each permission object may be an internal representation of a specific permission for S1 and may be associated with S1 to allow consumers to use / select the underlying database objects defined by the permission. For example, each permission object may have several fields, including: a global object reference to DB1 or one of its underlying objects that is being shared, an object type (e.g., security type), a privilege / permission type (e.g., use or select), and a timestamp when the permission occurred.
[0089] At block 815, the remote deployment 402 can deserialize the message and apply the sharing permissions to S1R. As can be seen, the sharing refresh described here is based on a pull model (as opposed to a push model). In other words, S1R will only be updated when a sharing refresh is requested, and there will be a difference between S1R and S1 until a sharing refresh of S1R is performed.
[0090] Figure 9 A diagrammatic representation of a machine is shown in the example form of a computer system 900, wherein a set of instructions is provided for causing the machine to perform any one or more of the methods discussed herein for copying a shared object to a remote deployment. More specifically, the machine may modify a shared object in a first account to a global object, wherein the shared object includes permission metadata indicating sharing permissions for a set of objects in a database. The machine may, in a second account located in the remote deployment, create a local copy of the shared object on the remote deployment based on the global object, copy the set of objects in the database to the local copy of the database on the remote deployment, and refresh the sharing permissions to the local copy of the shared object.
[0091] In an alternative embodiment, the machine can be connected (e.g., networked) to other machines in a local area network (LAN), an intranet, an extranet, or the Internet. The machine can operate with the ability of a server or client machine in a client-server network environment, or operate as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web device, a server, a network router, a switch or a bridge, a hub, an access point, a network access control device, or any machine that can (sequentially or otherwise) perform a set of instructions specifying the action to be taken by the machine. In addition, although only a single machine is shown, the term "machine" should also be understood to include any machine set that performs one set (or multiple sets) of instructions to perform any one or more methods discussed herein, either individually or jointly. In one embodiment, computer system 900 can represent a server.
[0092] The exemplary computer system 900 includes a processing device 902, a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 918, which communicate with each other via a bus 930. Any of the signals provided on the various buses described herein may be time-multiplexed with other signals and provided over one or more common buses. Furthermore, the interconnections between circuit components or blocks may be shown as buses or as single signal lines. Each of the buses may alternatively be one or more single signal lines, and each of the single signal lines may alternatively be a bus.
[0093] The computing device 900 may also include a network interface device 908 that can communicate with a network 920. The computing device 900 may also include a video display unit 910 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse), and a sound signal generating device 916 (e.g., a speaker). In one embodiment, the video display unit 910, the alphanumeric input device 912, and the cursor control device 914 may be combined into a single component or device (e.g., an LCD touch screen).
[0094] The processing device 902 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 902 can also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 902 is configured to execute function compilation instructions 925 for performing the operations and steps discussed herein.
[0095] Data storage device 918 may include a machine-readable storage medium 928 having stored thereon one or more sets of shared object copy instructions 925 (e.g., software) embodying one or more methods of functionality described herein. Shared object copy instructions 925 may also reside, completely or at least partially, within main memory 904 or within processing device 902 during execution by computer system 900; main memory 904 and processing device 902 also constitute machine-readable storage media. Shared object copy instructions 925 may also be transmitted or received over network 920 via network interface device 908.
[0096] The machine-readable storage medium 928 can also be used to store instructions to execute a method for determining a function to be compiled, as described herein. Although the machine-readable storage medium 928 is shown as a single medium in an exemplary embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) that store one or more sets of instructions. Machine-readable media include any mechanism for storing information in a form readable by a machine (e.g., a computer) (e.g., such as software, processing applications). Machine-readable media may include, but are not limited to, magnetic storage media (e.g., floppy disks); optical storage media (e.g., CD-ROMs), magneto-optical storage media; read-only memory (ROM); random access memory (RAM); erasable programmable memory (e.g., EPROMs and EEPROMs); flash memory; or another type of medium suitable for storing electronic instructions.
[0097] Unless otherwise specified, terms such as "receive," "route," "permit," "determine," "issue," "provide," "specify," "encode," and the like refer to actions and processes performed or implemented by a computing device that manipulate and transform data represented as physical (electronic) quantities in registers and memories of the computing device into other data similarly represented as physical quantities in the computing device memories or registers or other such information storage, transmission, or display devices. In addition, the terms "first," "second," "third," "fourth," and the like, as used herein, refer to labels used to distinguish different elements and do not necessarily have ordinal meanings based on their numerical names.
[0098] The examples described herein also relate to apparatus for performing the operations described herein. The apparatus may be specially constructed for the desired purpose, or it may comprise a general-purpose computing device selectively programmed by a computer program stored in the computing device. Such a computer program may be stored in a computer-readable, non-transitory storage medium.
[0099] The methods and illustrative examples described herein are not inherently related to any particular computer or other device. Various general-purpose systems may be used according to the teachings described herein, or it may prove convenient to construct more specialized devices to perform the required method steps. The structures required for various of these systems will be as described above.
[0100] The above description is illustrative rather than restrictive. Although the present disclosure has been described with reference to specific illustrative examples, it will be appreciated that the present disclosure is not limited to the described examples. The scope of the present disclosure should be determined with reference to the following claims and the full scope of equivalents to which the claims are entitled.
[0101] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprises," "comprising," "includes," and / or "including," when used herein, recite the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Therefore, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0102] It should also be noted that in some alternative embodiments, the functions / actions described may not occur in the order shown in the figures. For example, two features shown in succession may actually be performed substantially simultaneously, or may sometimes be performed in the reverse order, depending on the functions / actions involved.
[0103] Although the method operations are described in a particular order, it should be understood that other operations may be performed between the described operations, that the described operations may be adjusted so that they occur at slightly different times, or that the described operations may be distributed in a system that allows the processing operations to occur at various intervals associated with the processing.
[0104] Various units, circuits, or other components may be described or claimed as being “configured to” or “configurable to” perform one or more tasks. In such contexts, the phrases “configured to” or “configurable to” are used to imply structure by indicating that the unit / circuit / component includes structure (e.g., circuitry) that performs one or more tasks during operation. Likewise, a unit / circuit / component may be said to be configured to perform a task, or to be configurable to perform a task, even when the specified unit / circuit / component is not currently operational (e.g., not turned on). Units / circuits / components used with the “configured to” or “configurable to” language include hardware—e.g., circuitry, memory storing program instructions executable to implement an operation, etc. Stating that a unit / circuit / component is “configured” to perform one or more tasks, or “configurable” to perform one or more tasks, is expressly not intended to invoke 35 U.S.C. 112, paragraph 6, with respect to that unit / circuit / component. Additionally, "configured to" or "configurable to" may include general structures (e.g., general circuits) that are manipulated by software and / or firmware (e.g., an FPGA or a general-purpose processor executing software) to operate in a manner capable of performing the tasks in question. "Configured to" may also include adjusting a manufacturing process (e.g., a semiconductor fabrication facility) to manufacture a device (e.g., an integrated circuit) adapted to implement or perform one or more tasks. "Configurable to" is expressly not intended to apply to blank media, an unprogrammed processor or unprogrammed general-purpose computer, or an unprogrammed programmable logic device, programmable gate array, or other unprogrammed device, unless accompanied by programmed media that imparts to the unprogrammed device the ability to be configured to perform the disclosed functionality.
[0105] Any combination of one or more computer-usable or computer-readable media can be used. For example, the computer-readable medium may include one or more of the following: a portable computer disk, a hard disk, a random access memory (RAM) device, a read-only memory (ROM) device, an erasable programmable read-only memory (EPROM or flash memory) device, a portable compact disc read-only memory (CDROM), an optical storage device, a magnetic storage device. The computer program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages. Such code can be compiled from source code into a computer-readable assembly language or machine code suitable for the device or computer on which the code will be executed.
[0106] Embodiments may also be implemented in a cloud computing environment. In this specification and the appended claims, "cloud computing" may be defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned (including through virtualization) and released, and then scaled accordingly, with minimal management effort or service provider interaction. Cloud models may consist of various features (e.g., on-demand self-service, broad network access, resource pools, rapid elasticity, and measurable services), service models (e.g., Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"), and deployment models (e.g., private cloud, community cloud, public cloud, and hybrid cloud).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing a specified logical function. It will also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, may be implemented by a dedicated hardware-based system or a combination of dedicated hardware and computer instructions that performs a specified function or action. These computer program instructions may also be stored in a computer-readable medium that can direct a computer or other programmable data processing device to function in a specific manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instruction means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0108] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. In light of the above teachings, many modifications and variations are possible. These embodiments were chosen and described in order to best explain the principles of these embodiments and their practical application, thereby enabling others skilled in the art to best utilize these embodiments and various modifications that may be suitable for the specific use contemplated. Therefore, the present embodiments are to be considered illustrative and not restrictive, and the invention is not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.
Claims
1. A method for sharing replication between remote deployments, comprising: modifying, by a processing device, a shared object of a first account to a global representation of the shared object in a region associated with a second account, the first account residing in a first deployment, and wherein the shared object includes sharing permissions for a set of objects in a database; creating a local copy of the shared object in a second account based on the global representation of the shared object, wherein the second account resides in a second deployment remote from the first deployment; copying the set of objects of the database to a local database replica on the second deployment; and The sharing permissions are refreshed to the local copy of the shared object.
2. The method according to claim 1, further comprising: performing a subsequent refresh of the sharing permission to the local copy of the shared object; and A subsequent copy of the set of objects of the database to the local database copy is performed, wherein one or more of the sharing permissions and the set of objects are modified.
3. The method of claim 2 , wherein in response to determining during the subsequent refresh that one or more objects in the set of objects are not present in the local database copy: refreshing sharing permissions associated with an object present in the local database copy to the local copy of the shared object; and A log file is created indicating: the one or more objects not present in the local database copy; and Sharing permissions associated with the one or more objects not present in the local database copy.
4. The method of claim 1 , wherein refreshing the sharing permission to the local copy of the shared object comprises: Receive sharing refresh messages; serializing the sharing permission of the shared object; storing the serialized shared permission in a message, wherein the message includes information to deserialize the serialized shared permission; and The snapshot message is transmitted to the second deployment.
5. The method of claim 4, wherein refreshing the sharing permission to the local copy of the shared object further comprises: deserializing the snapshot message based on the information to deserialize the serialized sharing permission; and The sharing permission is applied to the local copy of the shared object.
6. The method of claim 1 , wherein creating the local copy comprises: establishing a secure connection between the first account and the second account; and A remote execution context is established in the second account, wherein the first account can use the remote execution context to create the local copy in the second account.
7. A system for sharing replication between remote deployments, comprising: Memory; and a processing device operatively coupled to the memory, the processing device configured to: modifying a shared object of a first account to a global representation of the shared object in a region associated with a second account, the first account residing in a first deployment, and wherein the shared object includes sharing permissions for a set of objects of a database; creating a local copy of the shared object in a second account based on the global representation of the shared object, wherein the second account resides in a second deployment remote from the first deployment; copying the set of objects of the database to a local database replica on the second deployment; and The sharing permissions are refreshed to the local copy of the shared object.
8. The system of claim 7, wherein the processing device is further configured to: performing a subsequent refresh of the sharing permission to the local copy of the shared object; and A subsequent copy of the set of objects of the database to the local database copy is performed, wherein one or more of the sharing permissions and the set of objects are modified.
9. The system according to claim 8, wherein: In response to determining during the subsequent refresh that one or more objects in the set of objects are not present in the local database copy, the processing device: refreshing sharing permissions associated with an object present in the local database copy to the local copy of the shared object; and A log file is created indicating: the one or more objects not present in the local database copy; and Sharing permissions associated with the one or more objects not present in the local database copy.
10. The system according to claim 7, wherein: To refresh the sharing permission to the local copy of the shared object, the processing device: Receive sharing refresh messages; serializing the sharing permission of the shared object; storing the serialized shared permission in a message, wherein the message includes information to deserialize the serialized shared permission; and The snapshot message is transmitted to the second deployment.
11. The system according to claim 10, wherein: In order to refresh the sharing permission to the local copy of the shared object, the processing device is further configured to: deserializing the snapshot message based on the information to deserialize the serialized sharing permission; and The sharing permission is applied to the local copy of the shared object.
12. The system of claim 7, wherein to create the local copy, the processing device: establishing a secure connection between the first account and the second account; and A remote execution context is established in the second account, wherein the first account can use the remote execution context to create the local copy in the second account.
13. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by a processing device, cause the processing device to: modifying, by the processing device, a shared object of a first account to a global representation of the shared object located in a region associated with a second account, the first account residing in a first deployment, and wherein the shared object includes sharing permissions for a set of objects in a database; creating a local copy of the shared object in a second account based on the global representation of the shared object, wherein the second account resides in a second deployment remote from the first deployment; copying the set of objects of the database to a local database replica on the second deployment; and The sharing permissions are refreshed to the local copy of the shared object.
14. The non-transitory computer-readable storage medium of claim 13, wherein: The processing equipment is also used for: performing a subsequent refresh of the sharing permission to the local copy of the shared object; and A subsequent copy of the set of objects of the database to the local database copy is performed, wherein one or more of the sharing permissions and the set of objects are modified.
15. The non-transitory computer-readable storage medium of claim 14, wherein: In response to determining during the subsequent refresh that one or more objects in the set of objects are not present in the local database copy, the processing device: refreshing sharing permissions associated with an object present in the local database copy to the local copy of the shared object; and A log file is created indicating: the one or more objects not present in the local database copy; and Sharing permissions associated with the one or more objects not present in the local database copy.
16. The non-transitory computer-readable storage medium of claim 13, wherein: To refresh the sharing permission to the local copy of the shared object, the processing device: Receive sharing refresh messages; serializing the sharing permission of the shared object; storing the serialized shared permission in a message, wherein the message includes information to deserialize the serialized shared permission; and The snapshot message is transmitted to the second deployment.
17. The non-transitory computer-readable storage medium of claim 16, wherein: In order to refresh the sharing permission to the local copy of the shared object, the processing device is further configured to: deserializing the snapshot message based on the information to deserialize the serialized sharing permission; and The sharing permission is applied to the local copy of the shared object.
Citation Information
Patent Citations
Secure Data Joins In A Multiple Tenant Database System
US20200311297A1
Distributed data management system and method
CN110019102A