Data exchange availability, inventory visibility and inventory implementation

By using a data exchange manager and a global messaging framework on a cloud computing platform, the low efficiency and control challenges of data sharing in existing technologies are solved, enabling secure and controllable cross-regional data sharing and meeting the data providers' need for flexible control over data access.

CN116097243BActive Publication Date: 2026-07-31SNOWFLAKE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SNOWFLAKE INC
Filing Date
2021-07-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing data sharing methods are inefficient, costly, and difficult to control the availability and visibility of data exchange sites across regions, making it difficult for data providers to share data securely and controllably with multiple entities.

Method used

Through the data exchange manager on the cloud computing platform, data providers can control data sharing in their private online marketplaces under their own brand. By utilizing the exchange manager and the global messaging framework, the availability and inventory visibility of data exchanges can be managed, allowing data providers to specify the availability and visibility of data in specific regions.

Benefits of technology

It enables secure and controllable cross-regional data sharing, improves data exchange efficiency, reduces costs, and meets the data providers' need for flexible control over data access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116097243B_ABST
    Figure CN116097243B_ABST
Patent Text Reader

Abstract

This document provides systems and methods for providing a secure and efficient way to manage the availability of data exchanges and the visibility of data inventories within data exchanges. For example, the method may include a set of areas available to the data exchange, specified by the exchange administrator, each of which includes one or more remote deployments. The method may also include a data provider specifying one or more areas within a set of areas where a data inventory owned by the data provider is visible. Upon receiving a request to access the data inventory from one or more remote deployments in one or more areas, the data provider can determine whether to deny or grant the request. In response to determining that the request should be granted, the data inventory data is copied to the remote deployments.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of U.S. Patent Application No. 16 / 994,325, filed August 14, 2020, pursuant to 35 U.S. SC § 119(e), the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to data sharing platforms, and more particularly to the availability of data sharing platforms between remote deployments and the visibility / implementation of listings between remote deployments.

[0004] background

[0005] Data sharing platforms, including databases, are widely used for data storage and access in computing applications. A database may include one or more tables that contain or reference data that can be read, modified, or deleted using queries. Databases can be used to store and / or access personal information or other sensitive information. Secure storage and access to database data can be provided by encrypting and / or storing data in encrypted form to prevent unauthorized access. In some cases, data sharing may be necessary to allow other parties to perform queries against a set of data. Brief description of the attached diagram

[0007] The described embodiments and their advantages can be best understood by referring to the following description taken in conjunction with the accompanying drawings. These drawings are in no way intended to limit any changes in form and detail that may be made to the described embodiments by those skilled in the art without departing from the spirit and scope of the described embodiments.

[0008] Figure 1A This is a block diagram depicting an example computing environment in which the methods disclosed herein can be implemented.

[0009] Figure 1B This is a block diagram illustrating an example virtual warehouse.

[0010] Figure 2 This is a schematic block diagram illustrating data that can be used to implement public or private data exchanges according to embodiments of the present invention.

[0011] Figure 3 This is a schematic block diagram of components for implementing a data exchange according to an embodiment of the present invention.

[0012] Figure 4A This is a block diagram illustrating remote deployment in a data exchange according to some embodiments of the present invention.

[0013] Figure 4BThis is a block diagram of remote deployment in a data exchange according to some embodiments of the present invention.

[0014] Figure 5 This is a block diagram of remote deployment in a data exchange according to some embodiments of the present invention.

[0015] Figure 6 This is a block diagram of remote deployment in a data exchange according to some embodiments of the present invention.

[0016] Figure 7 This is a flowchart of a method for managing the availability of a data exchange and the visibility of a data list according to some embodiments of the present invention.

[0017] Figure 8 This is a flowchart of a method for managing list approval requests according to some embodiments of the present invention.

[0018] Figure 9 This is a block diagram of an example computing device according to some embodiments of the present invention, which can perform one or more of the operations described herein.

[0019] Detailed description

[0020] Data providers often possess data assets that are difficult to share. These data assets can be data that another entity is interested in. For example, a large online retailer might have a dataset containing the purchasing habits of millions of customers over the past decade. This dataset can be enormous. If the online retailer wants to share all or part of this data with another entity (anonymized and / or aggregated under applicable privacy laws and contractual obligations), the online retailer might need to use outdated and slow methods to transfer the data, such as File Transfer Protocol (FTP), or even copy the data onto physical media and mail the physical media to the other entity. This has several drawbacks. First, it is slow. Copying terabytes or petabytes of data can take days. Second, once the data is transferred, the sharer has no control over what happens to it. The recipient can alter the data, make copies, or share it with other parties. Third, the only entities interested in accessing such large datasets in this way are large companies that can afford the complex logistics of transferring and processing the data, as well as the high cost of such cumbersome data transfers. Therefore, smaller entities (e.g., small and medium-sized enterprises (SMBs), mom-and-pop shops, etc.), or even smaller, more agile cloud-focused startups, often cannot access this data due to its prohibitive cost, even though the data may be valuable to their businesses. This is likely because raw data assets are often too coarse and riddled with potentially sensitive information to be sold directly to other companies. Data owners must first perform data cleaning, de-identification, aggregation, concatenation, and other forms of data enrichment before sharing with another party. This is both time-consuming and expensive. Finally, for the reasons mentioned above, traditional data sharing methods do not allow for scalable sharing, making it difficult to share data assets with many entities. Traditional sharing methods also introduce latency and delays to all parties accessing the most recently updated data.

[0021] Private and public data exchanges enable data providers to more easily and securely share their data assets with other entities. Public data exchanges (also referred to herein as the "Snowflake datamarketplace" or "data marketplace") provide a centralized repository with open access, where data providers can publish and control live and read-only datasets to thousands of customers. Private data exchanges (also referred to herein as "data exchanges") can operate under the data provider's brand, and the data provider can control who has access to the data. Data exchanges can be for internal use only or open to customers, partners, vendors, or others. Data providers can control which data assets are listed and who has access to which datasets. This allows for a seamless way to discover and share data within the data provider's organization and with its business partners.

[0022] Data exchange offices can, for example, Cloud computing services facilitate and allow data providers to offer data assets directly from their own online domains (e.g., websites) on a private online marketplace with their own brand. Data exchange marketplaces provide entities with a centralized, managed hub to list internally or externally shared data assets, foster data collaboration, and maintain data governance and audit access. Through data exchange marketplaces, data providers are able to share data between companies without duplicating it. Data providers can invite other entities to view their data listings, control which data listings appear on their private online marketplace, control who can access the data listings, and how others can interact with the data assets connected to those listings. This can be viewed as a “walled garden” marketplace where visitors must be approved to enter, and access to certain listings can be restricted.

[0023] For example, Company A could be a consumer data company that has collected and analyzed the consumption habits of millions of individuals across several different categories. Their datasets could include data on categories such as online shopping, video streaming, electricity consumption, car use, internet use, clothing purchases, mobile app purchases, club memberships, and online subscription services. Company A might want to offer these datasets (or subsets or derivatives of these datasets) to other entities. For example, a new clothing brand might want access to datasets related to consumer clothing purchases and online shopping habits. Company A could support a page on its website that is, or functions essentially, a data exchange, where data consumers (e.g., the new clothing brand) can directly browse, explore, discover, access, and potentially purchase datasets from Company A. Furthermore, Company A can control: who can access the data exchange, which entities can view specific listings, the actions entities can take on the listings (e.g., view only), and any other appropriate actions. Additionally, data providers can combine their own data with other datasets from, for example, public data exchanges (also known as “snowflake data markets” or “data markets”) and use the combined data to create new listings.

[0024] Data exchanges can be suitable venues for discovering, aggregating, cleaning, and enriching data to make it more valuable. A large company on a data exchange can aggregate data from its various branches and departments that may be valuable to another company. Furthermore, participants in a private ecosystem data exchange can work together to connect their datasets and collaboratively create useful data products that none of them could produce alone. Once these related datasets are created, they can be listed on the data exchange or data marketplace.

[0025] Sharing data can be performed when a data provider creates a shared object (hereinafter referred to as a share) of the database in the data provider's account and grants shared access rights to specific objects of the database (e.g., tables, secure views, and secure user-defined functions (UDFs)). A read-only database can then be created using the information provided in the share. Access to the database can be controlled by the data provider. A "share" encapsulates all the information needed to share data in the database. A share may include at least three pieces of information: (1) permissions to access the database and the schema containing the objects to be shared, (2) permissions to access specific objects (e.g., tables, secure views, and secure UDFs), and (3) the consumer account with which the database and its objects are shared. When sharing data, data is not copied or transferred between users. This is achieved through methods such as... The cloud computing service provider uses cloud computing services to complete the sharing.

[0026] Data shared by providers (also known as "data providers") can be described through a manifest defined by the provider in a data exchange or data marketplace. Access control, management, and governance of the manifest are likely similar for both data marketplaces and data exchanges. The manifest may include metadata describing the shared data.

[0027] The shared data can then be used to process SQL queries, which may include joins, aggregations, or other analyses. In some cases, the data provider can define a share that allows "secure connections" to be performed on the shared data. Secure connections can be performed to enable analysis on the shared data, but the actual shared data cannot be accessed by the data consumer (e.g., the recipient of the share).

[0028] In public or private data exchanges, many requests for inventory may originate from remote deployments in regions different from the provider's on-premises deployment. While cross-regional functionality is possible within a data exchange, in some cases, the data exchange owner / administrator may wish to restrict where the data exchange is available (e.g., which regions or remote deployments). Furthermore, providers may want to control where their data inventory is visible. For example, companies and governments may have different and varying requirements / regulations regarding where certain data is available. Data providers themselves may have their own requirements / restrictions on who can see / access their data and from where. While control over inventory visibility can be implemented within a single instance of a data exchange, it is not feasible in cross-regional data exchanges across multiple remote deployments that do not share the same storage. Moreover, even if the inventory is visible across multiple remote deployments, the underlying data still resides in the on-premises deployment, thus requiring means for requesting and accessing the data.

[0029] The systems and methods described herein provide a secure and efficient way to manage the availability of data exchanges and the visibility of data lists within them. For example, the method may include having the data exchange administrator specify a set of areas where the data exchange is available, each of which includes one or more remote deployments. The method may also include having a data provider specify one or more areas within a set of areas where a data list owned by the data provider is visible. Upon receiving a request to access the data list from one or more remote deployments in one of the areas, the data provider can determine whether to deny or grant the request. In response to determining that the request should be granted, a processing device copies the data from the data list to the remote deployments.

[0030] Figure 1AThis is a block diagram of an example computing environment 100 in which the systems and methods disclosed herein can be implemented. Specifically, a cloud computing platform 110, such as Amazon Web Services, can be implemented. TM (AWS), MICROSOFT AZURE TM Google Cloud TM As is known in the art, cloud computing platform 110 provides computing and storage resources that can be acquired (purchased) or leased and configured to perform applications and store data.

[0031] Cloud computing platform 110 can host cloud computing service 112, which facilitates data storage (e.g., data management and access), analysis functions (e.g., SQL queries and analysis), and other computing capabilities (e.g., secure data sharing among users of cloud computing platform 110) on cloud computing platform 110. Cloud computing platform 110 may include a three-tier architecture: data storage 140, query processing 130, and cloud services 120.

[0032] Data storage 140 can facilitate storing data in one or more cloud databases 141 on cloud computing platform 110. Data storage 140 can use storage services such as Amazon S3 to store data and query results on cloud computing platform 110. In a particular embodiment, to load data into cloud computing platform 110, data tables can be horizontally partitioned into large, immutable files, which can resemble blocks or pages in a traditional database system. Within each file, the values ​​of each attribute or column are grouped together and compressed using a scheme sometimes referred to as a hybrid columnar approach. Each table has a header that contains offsets for each column within the file, in addition to other metadata.

[0033] In addition to storing table data, data storage 140 also helps store temporary data generated by query operations (such as joins) and data contained in the results of large queries. This allows the system to compute large queries without encountering out-of-memory or out-of-disk errors. Storing query results in this way simplifies query processing because it eliminates the need for server-side cursors found in traditional database systems.

[0034] Query processing 130 can handle query execution within an elastic cluster of virtual machines (referred to herein as a virtual warehouse or data warehouse). Therefore, query processing 130 may include one or more virtual warehouses 131, which may also be referred to herein as data warehouses. Virtual warehouse 131 may be one or more virtual machines running on cloud computing platform 110. Virtual warehouse 131 may be computing resources that can be created, destroyed, or resized at any time as needed. This functionality can create “elastic” virtual warehouses that can be expanded, shrunk, or shut down according to user needs. Expanding a virtual warehouse involves creating one or more compute nodes 132 to virtual warehouse 131. Shrinking a virtual warehouse involves removing one or more compute nodes 132 from virtual warehouse 131. More compute nodes 132 can result in faster computation times. For example, data loading that takes 15 hours on a system with four nodes may only take 2 hours on a system with 32 nodes.

[0035] Cloud service 120 may be a collection of services that coordinate activities on cloud computing service 110. These services bundle together all the different components of cloud computing service 110 to handle user requests from login to query distribution. Cloud service 120 can operate on computing instances provided by cloud computing service 110 from cloud computing platform 110. Cloud service 120 may include services that manage virtual repositories, queries, transactions, data exchange, and a collection of metadata (such as database schemas, access control information, encryption keys, and usage statistics) associated with these services. Cloud service 120 may include, but is not limited to, authentication engine 121, infrastructure manager 122, optimizer 123, exchange manager 124, security engine 125, and metadata store 126.

[0036] Figure 1BThis is a block diagram illustrating an example virtual warehouse 131. Exchange manager 124 can use, for example, a data exchange to facilitate data sharing between data providers and data consumers. For example, cloud computing service 112 can manage the storage and access to database 108. Database 108 may include various instances of user data 150 for different users (e.g., different enterprises or individuals). User data may include a user database 152 containing data stored and accessed by that user. User database 152 may be subject to access controls such that, after authentication with cloud computing service 112, only the data owner is allowed to change and access database 112. For example, data may be encrypted such that it can only be decrypted using decryption information possessed by the data owner. Using exchange manager 124, specific data from user database 152, which is subject to these access controls, can be shared with other users in a controlled manner according to the methods disclosed herein. In particular, users can specify share 154, which, as described above, can be shared uncontrolled in a public data exchange or within a data exchange, or shared in a controlled manner with specific other users. The term "share" encapsulates all the information required for sharing data in the database. Sharing may include at least three pieces of information: (1) granting permissions to access the database and the schema containing the objects to be shared; (2) granting permissions to access specific objects (e.g., tables, secure views, and secure UDFs); and (3) the consumer account with which the database and its objects are shared. Data is not copied or transferred between users when sharing data. Sharing is accomplished through cloud service 120 of cloud computing service 110.

[0037] Shared data can be executed when a data provider creates a database share in their account and grants access to specific objects (such as tables, secure views, and secure user-defined functions (UDFs)). A read-only database can then be created using the information provided in the share. Access to that database can be controlled by the data provider.

[0038] The shared data can then be used to process SQL queries, which may include joins, aggregations, or other analyses. In some cases, the data provider can define a share that allows "secure connections" to be performed on the shared data. Secure connections can be performed to allow analysis to be performed on the shared data, but the actual shared data cannot be accessed by the data consumer (e.g., the recipient of the share). Secure connections can be performed as described in U.S. Application Serial No. 16 / 368,339, filed March 18, 2019.

[0039] User devices 101-104, such as laptop computers, desktop computers, mobile phones, tablet computers, cloud-hosted computers, cloud-hosted serverless processes, or other computing processes or devices, can be used to access virtual warehouses 131 or cloud services 120 via networks 105 such as the Internet or private networks.

[0040] In the following description, actions are attributed to users, particularly consumers and providers. It should be understood that such actions pertain to devices 101-104 operated by such users. For example, a notification to a user can be understood as a notification sent to device 101-104, input or instructions from a user can be understood as being received through the user's device 101-104, and user interaction with an interface should be understood as interaction with an interface on the user's device 101-104. Furthermore, database operations (connections, aggregations, analyses, etc.) attributed to a user (consumer or provider) should be understood to include cloud computing service 110 performing such actions in response to instructions from that user.

[0041] Figure 2 This is a schematic block diagram illustrating data that can be used to implement public or data exchange data according to embodiments of the present invention. Exchange manager 124 can operate with respect to some or all of the exchange data 200 shown, which may be stored on the platform (e.g., cloud computing platform 110) executing exchange manager 124 or at some other location. Exchange data 200 may include multiple lists 202 describing data shared by a first user (“Provider”). Lists 202 may be lists within a data exchange or a data marketplace. Access control, management, and governance of the lists may be similar for both data marketplaces and data exchanges.

[0042] Listing 202 may include metadata 204 describing the shared data. Metadata 204 may include some or all of the following information: the identifier of the sharer of the shared data, the URL associated with the sharer, the name of the share, the name of the table, the category to which the shared data belongs, the update frequency of the shared data, the table's directory, the number of columns and rows in each table, and the column names. Metadata 204 may also include examples to help users use the data. Such examples may include a sample table (which includes a sample of the rows and columns of the sample table), sample queries that can be run against the table, sample views of the sample table, and sample visualizations based on the table data (e.g., charts, dashboards). Other information included in metadata 204 may be metadata for use by business intelligence tools, a textual description of the data contained in the table, keywords associated with the table to facilitate searches, links to documents related to the shared data (e.g., URLs), and refresh intervals indicating the update frequency of the shared data and the last update date of the data.

[0043] Listing 202 may include access control 206, which can be configured with any suitable access configuration. For example, access control 206 may indicate that shared data is available without restriction to any member of a private exchange (as used elsewhere in this document as "any share"). Access control 206 may specify user categories (members of a specific group or organization) who are allowed to access the data and / or view the list. Access control 206 may specify "peer-to-peer" sharing (see discussion in Figure 4), where a user can request access but is only allowed access with the provider's approval. Access control 206 may specify a set of user identifiers that are excluded from accessing the data referenced in Listing 202.

[0044] Note that some Listing 202 entries may be discovered by users without further authentication or access permission, and actual access is only granted after subsequent authentication steps (see Figure 4 and...). Figure 6 (Discussion). Access control 206 can specify that list 202 can only be discovered by a specific user or a specific category of users.

[0045] It should also be noted that the default functionality of Listing 202 is that the data referenced in the shared data cannot be exported by consumers. Alternatively, access control 206 can specify that this is not allowed. For example, access control 206 can specify that security operations (such as secure connections and security functions, discussed below) can be performed on the shared data, making it impossible to view and export the shared data.

[0046] In some embodiments, once a user is authenticated with respect to Listing 202, a reference to that user (e.g., the user identifier of the user's account in Virtual Repository 131) is added to Access Control 206, so that the user can subsequently access the data referenced in Listing 202 without further authentication.

[0047] Listing 202 may define one or more filters 208. For example, filter 208 may define a specific user identifier 214 for users who can view references to listing 202 when browsing directory 220. Filter 208 may define user categories (users in a specific industry, users associated with a specific company or organization, users within a specific geographic region or country) for users who can view references to listing 202 when browsing directory 220. In this way, a private exchange can be implemented by exchange manager 124 using the same components. In some embodiments, excluded users in the excluded access list 202 (i.e., adding list 202 to the consumption share 156 of excluded users) may still be allowed to view a representation of the list when browsing directory 220 and may be further allowed to request access to list 202, as described below. Requests for access to the list by such excluded users and other users may be listed in the interface presented to the provider of listing 202. The provider of List 202 can then view the access list requirements and select Extended Filter 208 to grant access rights to excluded users or categories of excluded users (e.g., users from excluded geographic regions or countries).

[0048] Filter 208 can further define which data a user can view. Specifically, filter 208 can instruct that a user who selects list 202 to add to a user's consumption share 156 is allowed access to only a filtered version of the data referenced by that list, which includes only data associated with the user's identifier 214, associated with the user's organization, or specific to a particular category of the user. In some embodiments, private exchanges are conducted by invitation: upon accepting an invitation received from the provider, a user invited by the provider to view list 202 of private exchanges is allowed to view it through exchange manager 124.

[0049] In some embodiments, Listing 202 can be addressed to a single user. Therefore, a reference to Listing 202 can be added to a set of "pending shares" that the user can view. Listing 202 can then be added to the user's set of shares after the user passes approval to the exchange manager 124.

[0050] Listing 202 may further include usage data 210. For example, cloud computing service 112 may implement a points system where points are purchased by users and consumed each time a user runs a query, stores data, or uses other services implemented by cloud computing service 112. Therefore, usage data 210 may record points consumed by accessing shared data. Usage data 210 may include other data, such as the number of queries, the number of aggregations of each of several types performed against the shared data, or other usage statistics. In some embodiments, usage data for listing 202 or more of listings 202 for a user is provided to the user in the form of a shared database (i.e., the exchange manager 124 adds a reference to the database including the usage data to the user's consumption share).

[0051] Listing 202 may also include a heat map 211, which may represent the geographic location clicked by a user on that particular listing. Cloud service 110 may use the heat map to make replication decisions or other decisions about the listing. For example, a data exchange may display a listing containing weather data for the state of Georgia. Heat map 211 may indicate that many users in California are selecting the listing to check the weather in Georgia. Based on this information, cloud service 110 may replicate the listing and make it available in a database (the server of which is physically located in the western United States), making the data accessible to consumers in California. In some embodiments, an entity may store its data on a server located in the western United States. A particular listing may be very popular with consumers. Cloud service 110 may replicate the data and store it on a server located in the eastern United States, making the data accessible to consumers in the Midwest and East Coast as well.

[0052] List 202 may also include one or more tags 213. Tags 213 can facilitate simpler sharing of data contained in one or more lists. For example, a large company may have a Human Resources (HR) list containing HR data for its internal employees in a data exchange. HR data may contain ten types of HR data (e.g., employee number, chosen health insurance, current retirement plan, job title, etc.). 100 people in the company (e.g., everyone in the HR department) can access the HR list. HR management may want to add an eleventh type of HR data (e.g., employee stock option plans). Instead of manually adding it to the HR list and granting access to the new data to each of the 100 people, management can simply apply an HR tag to the new dataset, and this HR tag can be used to categorize the data as HR data, list it along with the HR list, and grant the 100 people access to view the new dataset.

[0053] Listing 202 may also include version metadata 215. Version metadata 215 provides a way to track how the dataset changes. This helps ensure that data being viewed by an entity is not changed prematurely. For example, if a company owns the original dataset and then releases an updated version of that dataset, the update might interfere with another user's processing of the dataset because the update might have a different format, new columns, and other changes that may be incompatible with the recipient user's current processing mechanisms. To remedy this, cloud computing service 112 can use version metadata 215 to track version updates. Cloud computing service 112 can ensure that each data consumer accesses the same version of the data until they accept an updated version that does not interfere with the current processing of the dataset.

[0054] Exchange data 200 may further include user records 212. User records 212 may include data that identifies the user associated with user record 212, such as an identifier (e.g., a warehouse identifier) ​​of a user who has user data 150 in service database 128 and is managed by virtual warehouse 131.

[0055] User record 212 may list shares associated with a user, such as a reference list 202 created by the user. User record 212 may also list shares consumed by the user, such as a reference list 202 created by another user and associated with the user's account according to the methods described herein. For example, list 202 may have an identifier that will be used to reference it in the sharing or consumed sharing of user record 212.

[0056] Exchange data 200 may further include a directory 220. Directory 220 may include a list of all available listings 202 and may include an index of data from metadata 204 to facilitate browsing and searching according to the methods described herein. In some embodiments, listings 202 are stored in the directory as JavaScript Object Notation (JSON) objects.

[0057] Note that in the case of multiple instances of Virtual Repository 131 existing on different cloud computing platforms, the directory 220 of one instance of Virtual Repository 131 may store manifests or references to manifests from other instances on one or more other cloud computing platforms 110. Therefore, each manifest 202 can be globally unique (e.g., all instances across Virtual Repository 131 are assigned a globally unique identifier). For example, instances of Virtual Repository 131 may synchronize copies of their directory 220 such that each copy indicates a manifest 202 available from all instances of Virtual Repository 131. In some instances, the provider of manifest 202 may specify that it is only available on one or more specified computing platforms 110.

[0058] In some embodiments, directory 220 is available on the Internet, making it searchable by search engines such as Bing or Google. The directory may be subject to search engine optimization (SEO) algorithms to improve its visibility. Potential consumers can therefore browse directory 220 from any web browser. Exchange manager 124 may publicly link to a Uniform Resource Locator (URL) for each listing 202. This URL may be searchable and can be shared outside of any interface implemented by exchange manager 124. For example, a provider of listing 202 may publish the URL of their listing 202 to promote the use of their listing 202 and its brand.

[0059] Figure 3 The various components 300-310 that may be included in the exchange manager 124 are shown. The creation module 300 can provide an interface for creating the inventory 202. For example, a web interface to the virtual repository 131 allows users to select data for sharing (e.g., a specific table in the user's user data 150) on devices 101-104 and enter values ​​for some or all of the defining metadata 204, access control 206, and filters 208. In some embodiments, creation can be executed by the user via SQL commands in an SQL interpreter executed on the cloud computing platform 110 and accessed through a web interface on user devices 101-104.

[0060] Verification module 302 can verify the information provided by the provider when attempting to create inventory 202. Note that in some embodiments, the actions attributed to verification module 302 can be performed by a human reviewing the information provided by the provider. In other embodiments, these actions are performed automatically. Verification module 302 can perform or facilitate the performance of various functions by a human operator. These functions may include verifying that metadata 204 is consistent with the shared data it references, verifying that the shared data referenced by metadata 204 is not pirated data, personally identifiable information (PII), personal health information (PHI), or other data whose sharing would be undesirable or illegal. Verification module 302 can also facilitate verification of whether the data has been updated within a threshold time period (e.g., within the most recent 24 hours). Verification module 302 can also facilitate verification that the data is not static or cannot be obtained from other static public sources. Verification module 302 can also facilitate verification that the data is not merely a sample (e.g., the data is complete enough to be useful). For example, geographically restricted data may be undesirable, while an aggregation of data that is otherwise unrestricted may still be useful.

[0061] Exchange manager 124 may include search module 304. Search module 304 can be implemented via a web interface accessible to the user on user devices 101-104 to invoke a search string for metadata in directory 220, receive a response to the search, and select a reference to list 202 in the search results to add to the consumption share 156 of the user record 212 of the user performing the search. In some embodiments, the search may be performed by the user via SQL commands in an SQL interpreter executed on cloud computing platform 102 and accessible through a web interface on user devices 101-104. For example, a search share can be performed via an SQL query targeting directory 220 within SQL engine 310, discussed below.

[0062] The search module 304 can further implement recommendation algorithms. For example, the recommendation algorithm can recommend other lists 202 to the user based on the user's consumption shares 156 or other lists previously in the user's consumption shares. Recommendations can be based on logical similarity: one weather data source leads to a recommendation for a second weather data source. Recommendations can be based on differences: one list for data in one domain (geographic region, technical field, etc.) results in lists in different domains (different geographic regions, related technical fields, etc.) to facilitate complete coverage of the user's analysis.

[0063] Exchange manager 124 may include access management module 306. As described above, a user can add listing 202. This may require authentication of the provider of listing 202. Once listing 202 is added to the user's consumption share 156 of user record 212, the user can either (a) be required to authenticate each time they access the data referenced by listing 202, or (b) be automatically authenticated and allowed access to the data once listing 202 is added. Access management module 306 can manage automatic authentication for subsequent access to data in user's consumption share 156 to provide seamless access to the shared data as if it were part of that user's user data 150. To this end, access management module 306 can access the control 206 of listing 202, certificates, tokens, or other authentication materials to authenticate the user when performing access to the shared data.

[0064] The exchange manager 124 may include a connection module 308. Connection module 308 manages the integration of shared data (i.e., shared data from different providers) referenced by user consumption shares 156 with each other and with the user database 152 containing user-owned data. Specifically, connection module 308 can manage the execution of queries and other computational functions regarding these various data sources, making their access transparent to the user. Connection module 308 can further manage data access to impose restrictions on shared data, such as enabling the execution of analyses and the display of analysis results without exposing the underlying data to data consumers (wherein such restrictions are indicated by access control 206 in Listing 202).

[0065] The exchange manager 124 may further include a standard query language (SQL) engine 310, which is programmed to receive queries from users and execute queries on data referenced by the queries, including user consumption shares 156 and user-owned user data 112. The SQL engine 310 can perform any query processing functions known in the art. The SQL engine 310 may additionally or alternatively include any other database management or data analysis tools known in the art. The SQL engine 310 may define a web interface that executes on the cloud computing platform 102, through which SQL queries are entered and responses to the SQL queries are presented.

[0066] Figure 4A A cloud environment 400 is illustrated, comprising multiple remote cloud deployments 401, 402, and 403. Each of the remote deployments 401, 402, and 403 may include a cloud computing service 112 (in... Figure 1A The architecture is shown in the diagram. Remote deployments 401, 402, and 403 can all be physically located in separate remote geographic areas, but can all be deployments in a single data exchange or a single data marketplace. In cloud environment 400, requests for data such as data inventories, databases, or shared data on remote deployment 401 can originate from accounts on remote deployment 402 or remote deployment 403. Remote deployment 401 can be the original deployment of the data exchange or data marketplace, and appropriate data replication methods can be used to make such requested data available on remote deployments 402 and 403.

[0067] For example, if account A resides on remote deployment 401 in zone 1 and has a database DB1 on remote deployment 401, and wants to share database DB1 with account B residing in remote deployment 402 in zone 2, then account A can modify database DB1 to be a global type database (as opposed to zone-specific) and (e.g., by using the SQL command "alter database DB1 enable replication to accounts Reg_2.B") replicate the metadata of DB1 to remote deployment 402. Account B can obtain a list of databases they can access (e.g., using the SQL command "show replication databases"), which will return an identifier indicating DB1, "Reg_1.A.DB1(primary)". Account B can then (e.g., by using the SQL command "create database DB1R as replica of Reg_1.A.DB1") create a local copy of DB1 on remote deployment 402. Figure 4A The database is displayed as DB1R, which creates a global type database because it is created as a replica. Note that data replication has not yet started. At this point, the command "show replication databases" will return the identifiers "Reg_1.A.DB1(primary)" and "Reg_2.B.DB1(secondary)". Account B can initiate data replication using a command such as "alter database DB1 refresh", which is a synchronization operation whose duration can depend on the amount of data to be synchronized. Figure 4B As shown, each remote deployment includes certain local objects and objects that it accesses a global version of. Although databases are discussed, the methods described above can be used to replicate various types of data objects between remote deployments, including data exchanges, data inventories, and shares.

[0068] In some embodiments, remote deployments 401-403 can utilize a global messaging framework that leverages (as discussed further in detail herein) specific message types, each specifically enabling various functionalities. For each global message type, there is a corresponding processing function applied to messages of that type. Therefore, as discussed further in detail herein, a particular type of global message will include custom logic for determining what processing needs to be performed on that particular message type.

[0069] While the cross-regional functionality discussed above is achievable, in some scenarios, data exchange owners / administrators may wish to restrict where the data exchange is available (e.g., which regions or remote deployments). Furthermore, data providers may want to control where their data inventory is visible. For example, companies and governments may have different and varying requirements / regulations regarding where certain data is available. Data providers themselves may have their own requirements / restrictions on who can see / access their data and where it can be seen / accessed, and may also want to restrict where their inventory is visible. While control over inventory visibility can be implemented within a single instance of a data exchange, it is not feasible to achieve such control across regional data exchanges or on remote deployments that do not share the same storage. Moreover, even if the inventory is visible on multiple deployments (402 and 403), a means for requesting and accessing the data is still required because the data still resides on the local deployment (401).

[0070] Embodiments of this disclosure can utilize the data replication process and global messaging framework described herein to replicate data between remote deployments 401-403 based on custom logic, making the data exchange available in a specific region, which may be across clouds. Information regarding the visibility of each data inventory within the data exchange is also replicated to the specific region, enabling the implementation of such restrictions in each remote deployment even if the data inventory was not initially created there. Although discussed in the context of a data exchange, embodiments of this disclosure can also be implemented in a data marketplace. Figure 4B A cloud environment 400 according to some embodiments of the present disclosure is shown.

[0071] Figure 4B Remote deployment 401 is shown, which can be the original deployment of data exchange DX1, as well as remote deployments 402 and 403. Remote deployments 402 and 403 are remote deployments that make data exchange DX1 available, and as mentioned above, each remote deployment can reside in its own geographic region (hereinafter referred to as "region"), and in Figure 4B In the diagram shown as zones 1, 2, and 3), the data exchange DX1 can have a designated data exchange administrator account (hereinafter referred to as "exchange admin") and can provide the ability for the exchange administrator on the remote deployment 401 to specify the zones that the data exchange DX1 will have available (resolvable) zones and from which customers can be added as members of the data exchange DX1. It should be noted that the exchange administrator (like other Snowflake accounts) can include an account administrator role that can delegate the ability to specify the zones available to the data exchange DX1 to other roles within the exchange administrator group. The data exchange DX1 may also include features that allow data providers to restrict the list of allowed zones (e.g., ...). Figure 4BThe list shown (DXL1) illustrates the functionality of the visibility zone. Remote deployment 401 can provide exchange administrators with commands (e.g., SQL commands) to configure the availability zone. For example, an exchange administrator can use the command "Create data exchange..."<data_exchange_name> The command `regions=region1,...` creates a data exchange available in certain regions (e.g., region 1, etc.). Exchange administrators can use the command `Alter data exchange` to modify the available regions when they wish to change them.<data_exchange_name> The command `setregions = region1, region2...` modifies the regions available in a data exchange. For example, an exchange administrator can also use the command `Alter data exchange`.<data_exchange_name> The command "unset regions" removes all currently set availability regions. In some implementations, exchange administrators can modify availability regions, while data exchange account holders, administrators, and data providers can (e.g., using the command "Show regions in data exchange")<data_exchange_name> View the list of available regions. For Snowflake Data Marketplace (SDM), the available regions can be automatically set to the regions currently replicating the SDM.

[0072] When the exchange administrator sets up availability zones for the data exchange, this information can be stored as a list in a local database (not shown) on the remote deployment 401. The local database can be any suitable database, such as FoundationDB. The local database on the remote deployment 401 can include multiple Data Processing Objects (DPOs) where data related to the data exchange DX1 can be stored. For example, a basic dictionary DPO can include a set of database tables used to store information about database definitions, including information about database objects such as tables, indexes, columns, data types, and views.

[0073] Such a DPO can be an availability zone DPO that extends the base dictionary DPO, and the availability zone of data exchange DX1 can be stored within it. In other words, the specified availability zone can be an attribute of the base dictionary DPO. As can be seen from the example commands listed above, the exchange administrator can specify the areas available to data exchange DX1 on a region-by-region basis, rather than specifying the specific remote deployments available to DX1 on a deployment-by-deploy basis. Therefore, when the "Alter data exchange" command is executed, remote deployment 401 can store the deployment location ID of each area on which it wants to make data exchange DX1 available, rather than storing the deployment identifier (ID) of the remote deployment on which it wants to make data exchange DX1 available. The deployment location ID can be represented in any suitable alphanumeric form, such as 1001 or region1 (corresponding to region 1) and 1002 or region2 (corresponding to region 2). A list of available deployment location IDs can be stored as a string within the Availability Zone DPO (defined, for example, as the static final string AVAILABLE_DEPLOYMENT_LOCATION_IDS="availabledeploymentlocationIDs"), and this string can be parsed to determine the deployment location IDs of the regions available in Data Exchange DX1 when a member of DX1 wants to know the availability zones. It should be noted that any of Zones 1, 2, and 3 can contain multiple remote deployments, and each of these remote deployments can be referred to as a deployment shard. Each deployment shard in a specific zone will share the same deployment location IDs. Utilizing deployment location IDs is efficient because it eliminates the need to manually refresh the list of available deployment IDs (strings) in the Availability Zone DPO every time a new deployment is created. For example, if a new shard deployment is added to a zone, storing the deployment ID would require manually refreshing the list of available deployment IDs in the relevant DPO. By utilizing / storing the deployment location ID, if a new deployment / shard is created in any region, for example, a remote deployment 401 only needs to obtain the deployment region of the new deployment / shard, which is easy because it is included in the deployment metadata of the new deployment / shard.

[0074] Remote Deployment 401 can then use the database replication method discussed above to replicate Data Exchange DX1 to each remote deployment in each region available to the Data Exchange (as specified by the Exchange Administrator). For the global object corresponding to Data Exchange DX1, Remote Deployment 401 can determine which remote deployment the global object should be replicated to by parsing the string of the Deployment Location ID from the Available Region DPO. Figure 4BIn the example shown, the exchange administrator can set Zone 1 (where a zone already exists) and Zone 2 as available zones. When replicating data exchange DX1, remote deployment 401 needs to know which remote deployments are available in Zone 2 and can obtain all remote deployments in Zone 2 (e.g., deployment location ID 1002). Figure 4B In the example, this could include remote deployments 402, 402B, and 402C. More specifically, remote deployment 401 could include a mapping between the deployment location ID of region 2 and the deployment ID of each deployment shard in region 2. Therefore, the data exchange DX1 can easily look up all deployment shard IDs in region 2 (identified by its deployment location ID) and copy the information to all relevant deployment shards. Figure 4B As shown, the global object corresponding to data exchange DX1 is then copied to remote deployment 402. When a new deployment is created, the list can be populated by refreshing the list of remote deployments to be copied to the new remote deployment. Remote deployment 401 can then continue the data copying method described above to copy data exchange DX1 to each remote deployment (i.e., remote deployment 402) in region 2. In some embodiments, remote deployment 401 can perform a process of obtaining a list of available regions and copying data exchange DX1 to remote deployments in those regions at regular intervals. Figure 4B As can be seen, remote deployment 402 can now access a global copy of the data exchange DX1.

[0075] When setting up an availability zone for Data Exchange DX1, data providers for Data Exchange DX1 can set the zones in which their inventory will be visible (e.g., setting inventory visibility). An inventory can be a consumer-visible representation of the data that a data provider wishes to share. An inventory can describe what the underlying data is about, including examples of data use and other metadata discussed herein. Data providers create inventory, and at creation time, only data providers can see the inventory. Data providers can send the inventory to the exchange administrator for publication approval (as described further in detail herein, referred to as "inventory approval"). Once approved, the data provider can publish the globally available inventory in the zones where it is available in Data Exchange DX1.

[0076] Inventory visibility does not refer to physical limitations imposed due to the presence (or absence) of an inventory in remote deployments. This means that the inventory can still be replicated to those deployments while remaining invisible to consumers of those deployments. Once the exchange administrator determines which regions the data exchange DX1 is available in, the data provider can select a subset of those regions to make the inventory visible in that subset.

[0077] exist Figure 4BIn the example shown, in remote deployment 402, the data provider can (locally within remote deployment 402) generate a listing DXL1 to share specific data. A local copy of the data exchange DX1 (e.g., previously copied from remote deployment 401) can provide the data provider with a set of commands (e.g., SQL commands) to set the areas where listing DXL1 will be visible. For example, the data provider can use the command "Alter listing..."<listing_name> The command `set regions = region1, region2....` sets the regions visible to DXL1. Data providers can use the command "Alter listing" to specify the visible regions.<listing_name> The "unset regions" command removes all previously set regions (making the list invisible in any region), and the command "Show listings in data exchange" can be used to do so.<dx_name> Use the ;” option to view the current area visible to DXL1.

[0078] When a data provider sets the areas where listing DXL1 will be visible, this information can be stored as a list in the local database (not shown) of remote deployment 402. The local database of remote deployment 402 can be any suitable database, such as FoundationDB, and can include listing visible areas DPO (not shown), which extends the base dictionary DPO and can store one or more listing visible areas. As can be seen in the example commands listed above, a data provider can specify the areas where its listings are visible on a per-region basis, rather than specifying a specific deployment on which its listings are visible on a per-deployment basis. Therefore, when executing "Alter listing..."<listing_name> When using the "set regions" command, remote deployment 402 can store the deployment location ID of each region visible to list DXL1, instead of storing the deployment ID of the remote deployments visible to list DXL1. The list of deployment location IDs that make list DXL1 visible can be stored as a string (defined as, for example, the static final string VISIBLE_DEPLOYMENT_LOCATION_IDS="availabledeploymentlocationIDs") in the list-visible regions DPO, and this string can be parsed to determine the deployment location ID of the regions visible to list DXL1 when the data provider or exchange administrator wants to know which regions list DXL1 will be visible.

[0079] Using deployment location IDs is efficient because each time a new deployment is created, there's no need to manually refresh the list of deployment IDs visible in the DPO (Distributed Product Object) for the deployment on it. For example, if a new shard deployment is added to a region, storing deployment IDs would require manually refreshing the list of deployment IDs visible in the DPO. By leveraging / storing deployment location IDs, if a new deployment / shard is created, the data exchange only needs to retrieve the deployment location (region) of the new deployment / shard, which is straightforward because it resides in the deployment metadata of the new deployment / shard.

[0080] When a visible area for manifest DXL1 is set, remote deployment 402 can copy manifest DXL1 and the visibility list to each remote deployment in each area where manifest DXL1 is visible. As described above, remote deployment 402 can obtain a list of areas visible to manifest DXL1 by parsing a string from the deployment location ID of the manifest visible area DPO, and can package the list of areas along with other information about manifest DXL1, such as the type of manifest DXL1 and its metadata, into a single manifest information package. Remote deployment 402 can utilize the data replication method described herein, and when a global object corresponding to manifest DXL1 is created, it can include the manifest information package. In some embodiments, if the exchange administrator is located on a remote deployment different from the data provider (e.g., in...), Figure 4B In the example, the exchange administrator can obtain a list of regions where list DXL1 is visible from the global object corresponding to list DXL1 (which includes a copy of the list package). Remote deployment 402 can determine which remote deployment(s) the global object should be copied to based on the list of regions where list DXL1 is visible. Remote deployment 402 can then perform data copying to copy list DXL1 and the list package to each remote deployment in each region where list DXL1 is visible. Remote deployment 402 can perform a process of obtaining a list of regions where list DXL1 is visible and copying list DXL1 and the list package to the remote deployments in those regions at regular intervals. Figure 4B In the example shown, the data provider has set regions 1 and 2 as regions where manifest DXL1 is visible, and therefore DXL1 is copied to remote deployment 401.

[0081] In some embodiments, list DXL1 and the corresponding visibility list can be replicated to each region where data exchange DX1 is available, and list visibility restrictions can be logically enforced on remote deployments in regions specified by the data provider where the list does not imply visibility. For example, if the deployment location ID of region 3 is not included in the visibility list, list DXL1 and the visibility list can still be replicated to remote deployment 403 (if available in this data exchange), but when a consumer on remote deployment 403 wants to resolve the list available to them, remote deployment 403 can logically enforce the visibility restrictions set by the data provider, and the consumer on remote deployment 403 may not see list DXL1.

[0082] When a consumer in remote deployment 401 in Zone 1 (e.g., where the inventory is visible, as specified by the data provider) attempts to resolve the inventory available to them, they can see the data provider's inventory DXL1 and can request access to the data in inventory DXL1. If the inventory is pre-approved and data has been attached to inventory DXL1, the data in inventory DXL1 will be copied immediately / directly along with inventory DXL1 and the inventory information package. If data has not yet been attached to inventory DXL1, inventory DXL1 and the inventory information package will still be copied to remote deployment 401, but the consumer in Zone 1 will need to request the data.

[0083] If the data provider subsequently updates the list of visible areas of the DXL1 manifest so that the manifest is no longer visible in the areas where it was previously visible, then consumers on remote deployments in that area who are members of the data exchange DX1 will still be able to resolve the manifest when it is copied; however, consumers on remote deployments in that area who are new members of the data exchange DX1 may not be able to resolve the manifest.

[0084] When inventory DXL1 is replicated to each appropriate remote deployment, data exchange DX1 and inventory DXL1 are set to global, allowing requests from consumers in any appropriate remote deployment to consume the underlying data of inventory DXL1. However, although inventory DXL1 is visible across multiple remote deployments, the underlying data still resides in the local remote deployment 401. To request the underlying data and fulfill requests, an existing global messaging framework is used to manage consumer requests to the inventory and allows data providers to manage inventory approval requests.

[0085] Figure 5 A diagram of cloud environment 500 is shown, which can be similar to cloud environment 400 shown in Figure 4. Figure 5In the example, a consumer on remote deployment 503, where list DXL2 is visible, wants to request data for list DXL2 from a data provider that owns list DXL2 on remote deployment 502, which can communicate with the exchange administrator on remote deployment 501.

[0086] When a consumer in remote deployment 503 wishes to request list DXL2, they can utilize list metadata (included in a list information packet that contains a copy of the global object corresponding to list DXL2) that indicates who the data provider is and where they come from / originate from the remote deployment to determine where to send the request. Remote deployment 503 can utilize global messages with the global message type "DATA_EXCHANGE_LISTING_REQUEST_SYNC". As mentioned above, for each global message type, there is a corresponding processing function that applies to messages of that type. Therefore, a specific type of global message will include custom logic on what processing needs to be performed for that specific message type. The DATA_EXCHANGE_LISTING_REQUEST_SYNC message type can be used to manage consumer list requests to providers. This includes creating, canceling, rejecting, and fulfilling these requests, as well as cleaning up requests (expiring them) when removing members from the data exchange or deleting the list. These messages are sent between the data provider and the consumer. Remote deployment 503 can send a create message (of type DATA_EXCHANGE_LISTING_REQUEST_SYNC) to remote deployment 502. Remote deployment 502 may include a local database with an access request DPO (not shown), which can be used by a data provider to manage the approval / rejection of data list requests. As discussed in this paper regarding the global messaging framework, the create message may include dedicated logic to update the appropriate slice of the access request DPO with the requested information. Examples of the requested information may include the requester's contact information, the requester's Snowflake account and its Snowflake region, and the reasons / causes they may be interested in. As used herein, a slice of a multidimensional array (e.g., a DPO) is a column of data corresponding to a single value of one or more members of a particular dimension.

[0087] In a remote deployment 502, a data provider can request a list DXL2 by creating a share associated with the list and granting consumers access to that share. The ListingRequestFulfiller background service (BG) can synchronize list request implementation information and notify / replicate it to other regions / deployment shards that may be of interest. More specifically, the ListingRequestFulfiller BG can invoke an implementation (global) message (of type DATA_EXCHANGE_LISTING_REQUEST_SYNC) that will mark the list provider's request as implemented in the access request DPO, remove it from the access request DPO's "provider_pending" slice, and write it to the access request DPO's "provider_history" slice after setting its status to implemented. It should be noted that the share associated with this list DXL2 can be created (and granted access to) by a data provider or an implementer that is in the same remote deployment shard as the consumer (e.g., remote deployment 503) or in the same region as the consumer (e.g., region 3). If access is granted by an implementer in the same deployment shard as the consumer, this can trigger a write to the “listingShareUpdatedOn” slice in the shared state DPO on remote deployment 503, which the consumer uses to manage their list data requests. The “listingShareUpdatedOn” slice can be used to indicate the list of data for which the consumer has been granted access to the share. If access is granted by an implementer in a deployment shard not in the same deployment shard as the consumer but in the same region, the “RemoteShardAccountManager” BG, which synchronizes account and share information between deployment shards in the same region, can run on the consumer’s remote deployment 503, see the consumer added to the share, and update the “listingShareUpdatedOn” slice of the shared state DPO. The “ListingRequestFulfiller” BG will run in the consumer’s remote deployment 503, mark the request as local implementation in the shared state DPO, and send an implementation message (of type: DATA_EXCHANGE_LISTING_REQUEST_SYNC) to the provider on the remote deployment 502 to update the access request DPO by marking the request as implemented, removing it from the “provider_pending” slice, and writing it to the “provider_history” slice after setting its state to implemented.

[0088] If the provider denies the request, it can update the access request DPO and send a denial message (of type: DATA_EXCHANGE_LISTING_REQUEST_SYNC) to the remote deployment 503, which has the logic to update the appropriate slice of the shared state DPO.

[0089] In some embodiments, no request from a consumer is required, and the data provider can create a share (not shown) and attach it to the data inventory DXL2. The data provider can add consumers to the share, and consumers can consume data from the share. Note that in embodiments where no consumer requests the share, it can be created by the data provider or an implementer (the implementer being the data provider in the same remote deployment as the consumer).

[0090] Figure 6 A cloud environment 600 is shown, which can be similar to the cloud environment 400 shown in Figure 4. Figure 6In the example, a data provider on remote deployment 602 might want to send an approval request to the exchange administrator on remote deployment 601 for publishing its list DXL3. The data provider and exchange administrator can use a special global message type (e.g., global message type: DATA_EXCHANGE_LISTING_APPROVAL_REQUEST_SYNC) to manage the data provider's requests for approval to publish its list, which includes the creation, cancellation, rejection, and approval of publication requests. A publish request DPO on the local database of remote deployment 601 can be used by the exchange administrator to manage the approval / rejection of list publication requests. A publish request DPO can include multiple slices, where each slice is a data column containing a single value corresponding to each of one or more members of a specific dimension of the DPO. A publish request DPO can include a "exchange admin" slice for the exchange administrator, a "data provider" slice for the data provider, and an "updatedOn" slice for tracking the time when the request was last updated. Each slice can include one or more data categories, such as the local entity ID of the data exchange for the requested inventory, the deployment where the data exchange for the requested inventory resides, the deployment where the requested inventory resides, the local entity ID of the requested inventory, the account ID of the inventory owner (provider), the status of the request (e.g., pending, rejected, approved, etc.), a JSON string containing information for display in the user interface (UI), the reason why the request was rejected (if the request was rejected), the timestamp when the request was issued, and the timestamp when the request was last updated. The local database of a remote deployment 602 can include a separate inventory approval request DPO identical to the publish request DPO and used by the data provider to manage inventory publish requests. The inventory approval request DPO and the publish request DPO can share similar information because multiple accounts cannot modify the same object / DPO, and therefore two separate but similar DPOs (each owned by a single participant (e.g., exchange administrator and provider)) are utilized.

[0091] A data provider can generate an approval request instructing him / her to publish list DXL3 on remote deployment 601 of the exchange administrator, and update the relevant data categories of the "Provider" slice (of the list approval request DPO) with the information in the request. The data provider can then (e.g., via remote deployment 602) send a creation message to the exchange administrator on remote deployment 601 to request the publication of data list DXL3 on remote deployment 601. The creation message can write the approval request to the "exchange admin" slice and the "updatedOn" slice of the publication request DPO on remote deployment 601. More specifically, the creation message can update each of the relevant data categories listed above in each of the "exchange admin" slice and the "updatedOn" slice of the publication request DPO with the information in the approval request. The creation message can also remove any rejected or approved approval requests for the same list from the administrator slice.

[0092] If the exchange administrator decides to reject the approval request, they can update the "Request Status" and "Rejection Reason" fields in the "exchangeadmin" and "updatedOn" slices of the publishing request DPO, and use a rejection message to update the "Data Provider" slice of the inventory approval request DPO on the remote deployment. As part of updating the data provider slice, the rejection message can correspondingly update the "Request Status" and "Rejection Reason" fields in the "Data Provider" slice of the inventory approval request DPO.

[0093] If the exchange administrator decides to grant the approval request, they can update the "Request Status" and "Rejection Reason" fields in the "exchangeadmin" and "updatedOn" slices of the publishing request DPO, and use an implementation message to update the data provider slice of the inventory approval request DPO on the remote deployment 602. As part of updating the data provider slice, the implementation message can correspondingly update the "Request Status" and "Rejection Reason" fields in the "Data Provider" slice of the inventory approval request DPO.

[0094] Data providers can also utilize cancellation messages, which can remove any approved requests (with pending, approved, or rejected statuses) from the exchange administrator slice of the Release Request (DPO) on a remote deployment (401). When a data provider releases a list of approved requests, cleanup will "cancel" the request on their behalf using the same code path to remove the request from the exchange administrator's side.

[0095] Figure 7This is a flowchart of a method 700 for managing the availability of a data exchange and the visibility of a data inventory therein, according to some embodiments. Method 700 may be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, processor, processing device, central processing unit (CPU), system-on-a-chip (SoC), etc.), software (e.g., instructions that run / execute on the processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, method 700 may be executed by (in...) Figure 4B The corresponding processing devices for remote deployment of 401 and 402 are shown in the figure.

[0096] Also refer to Figure 4B At block 705, the exchange administrator can set the zones in which data exchange DX1 will be available. Data exchange DX1 can provide the following functionality: allowing the exchange administrator on remote deployment 401 to specify the zones in which data exchange DX1 will be available (resolvable) and from which zones customers can be added as members of data exchange DX1. Remote deployment 401 can provide the exchange administrator with commands (e.g., SQL commands) to set the available zones. When the exchange administrator sets the available zones for the data exchange, this information can be stored as a list in a local database (not shown) of remote deployment 401. The local database can be any suitable database, such as FoundationDB. The local database of remote deployment 401 can include multiple Data Processing Objects (DPOs) in which data related to data exchange DX1 can be stored. For example, a basic dictionary DPO can include a set of database tables used to store information about database definitions, including information about database objects such as tables, indexes, columns, data types, and views.

[0097] Such a DPO can be an Availability Region DPO that extends the base dictionary DPO, and the availability regions of data exchange DX1 can be stored within it. As can be seen from the example commands listed above, the exchange administrator can specify the regions in which data exchange DX1 is available on a region-by-region basis, rather than specifying a specific remote deployment where DX1 is available on a deployment-by-deployment basis. Remote deployment 401 can store the deployment location ID for each region in which the data exchange will be available. The deployment location ID can be represented in any suitable alphanumeric form, such as 1001 or region1 (corresponding to region 1), 1002 or region2 (corresponding to region 2). The list of available deployment location IDs can be stored as a string within the Availability Region DPO (defined, for example, as the static final string AVAILABLE_DEPLOYMENT_LOCATION_IDS=).

[0098] The string “availabledeploymentlocationIDs” can be parsed to determine the deployment location IDs of the regions available to data exchange DX1 when members of data exchange DX1 want to know the available regions.

[0099] At block 710, remote deployment 401 can then use the database replication method discussed above to replicate the data exchange DX1 to each remote deployment in each region where the data exchange DX1 is available (as specified by the exchange administrator). For the global object corresponding to the data exchange DX1, remote deployment 401 can determine which remote deployment the global object should be replicated to by parsing the string of the deployment location ID from the available region DPO to determine the list of regions where the data exchange DX1 is available.

[0100] Once an availability zone for the data exchange is set up, at block 715, the data provider for data exchange DX1 can set the zone in which its inventory (e.g., inventory DXL1) will be visible (e.g., setting inventory visibility). An inventory can be a client-visible representation of the data the data provider wishes to share. The inventory can describe what the underlying data is about, including examples of how the data will be used, and other metadata. The data provider creates the inventory, and at creation time, only the data provider can see the inventory. The data provider can send the inventory to the exchange administrator for publication approval (as described further in detail herein, referred to as "inventory approval"). Once approved, the data provider can publish the globally available inventory in the zone where data exchange DX1 is available.

[0101] When a data provider sets the regions where inventory DXL1 will be visible, this information can be stored as a list in the local database (not shown) of remote deployment 402. The local database of remote deployment 402 can be any suitable database, such as FoundationDB, and can include inventory visibility regions DPO (not shown), which extends the base dictionary DPO and can store one or more inventory visibility regions. As can be seen in the example commands listed above, a data provider can specify regions where its inventory is visible on a region-by-region basis, rather than specifying a specific deployment on which its inventory is visible on a deployment-by-deployment basis. The list of deployment location IDs where inventory DXL1 will be visible can be stored as a string in the inventory visibility regions DPO, and when the data provider or exchange administrator wants to know the regions where inventory DXL1 will be visible, this string can be parsed to determine the deployment location IDs of the regions where inventory DXL1 is visible.

[0102] When a visible region for manifest DXL1 is set, at block 720, remote deployment 402 can copy manifest DXL1 and the visibility list to each remote deployment in each region where manifest DXL1 is visible. As described above, remote deployment 402 can obtain a list of regions where the manifest is visible by parsing a string from the deployment location ID of the manifest visibility region DPO, and can package the list of regions along with other information about the manifest, such as the manifest type and manifest metadata, into a single manifest information package. Remote deployment 402 can utilize the copying method described above, and when a global object corresponding to manifest DXL1 is created, it can include the manifest information package.

[0103] Now also referencing Figure 5 When a consumer in remote deployment 503 wishes to request list DXL2, they can utilize list metadata (included in a list information packet that utilizes a copy of a global object corresponding to list DXL2) indicating who the data provider is and where they come from / originate from the remote deployment to determine where to send the request. Remote deployment 503 can utilize a global message with the global message type: DATA_EXCHANGE_LISTING_REQUEST_SYNC. This message type can be used to manage consumer list requests to providers. This includes creating, canceling, rejecting, and fulfilling these requests, as well as cleaning up requests (expiring them) when a member is removed from the data exchange or a list is deleted. At block 725, remote deployment 503 can send a creation message requesting access to list DXL2 to remote deployment 502, which may include a local database with an access request DPO, which can be used by the data provider to manage the approval / rejection of requests for the data list.

[0104] At block 730, a data provider in remote deployment 502 can request list DXL2 by creating a share associated with the list and granting the consumer access to the share associated with the list. It should be noted that the share associated with list DXL2 can be created (and access granted) by the data provider or an implementer of the data provider in the same remote deployment as the consumer (e.g., remote deployment 403).

[0105] Figure 8This is a flowchart of a method 800 for managing list approval requests according to some embodiments. Method 800 may be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, processor, processing device, central processing unit (CPU), system-on-a-chip (SoC), etc.), software (e.g., instructions that run / execute on the processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, method 800 may be executed by (in...) Figure 4B The corresponding processing devices for remote deployment of 401 and 402 are shown in the figure.

[0106] Also refer to Figure 6 A data provider on remote deployment 602 might want to send a request to the exchange administrator on remote deployment 601 for approval to publish its list DXL3. The data provider and exchange administrator can use a special global message type (e.g., global message type: DATA_EXCHANGE_LISTING_APPROVAL_REQUEST_SYNC) to manage the data provider's requests for approval to publish its list, which includes the creation, cancellation, rejection, and approval of publish requests. A publish request DPO on the local database of remote deployment 601 can be used by the exchange administrator to manage the approval / rejection of list publish requests. A publish request DPO can include multiple slices, where each slice is a data column corresponding to a single value of each of one or more members of a specific dimension of the DPO. A publish request DPO can include a "exchange admin" slice for the exchange administrator, a "data provider" slice for the data provider, and an "updatedOn" slice for tracking the time when the request was last updated. Each slice may include one or more data categories, such as the local entity ID of the data exchange for the requested manifest, the deployment where the data exchange for the requested manifest resides, the deployment where the requested manifest resides, the local entity ID of the requested manifest, the account ID of the manifest owner (provider), the status of the request (e.g., pending, rejected, approved, etc.), a JSON string containing information for the user interface (UI) display, the reason why the request was rejected (if the request was rejected), the timestamp when the request was issued, and the timestamp when the request was last updated. The local database of a remote deployment 602 may include a separate manifest approval request DPO identical to the publish request DPO and is used by the data provider to manage manifest publish requests.

[0107] At block 805, the data provider on remote deployment 602 can generate an approval request instructing him / her to publish list DXL3 on remote deployment 601 of the exchange administrator, and update the relevant data category of the "Provider" slice (of the list approval request DPO) with the information in the request. Subsequently, at block 810, the data provider (e.g., via remote deployment 602) can send a creation message to the exchange administrator on remote deployment 601 to request the publication of data list DXL3 on remote deployment 601. The creation message can write the approval request to the "exchangeadmin" slice and the "updatedOn" slice of the publication request DPO on remote deployment 601. More specifically, the creation message can update each of the above-listed relevant data categories in each of the "exchangeadmin" slice and the "updatedOn" slice of the publication request DPO with the relevant information in the approval request. The creation message can also remove any rejected or approved approval requests for the same list from the "admin" slice.

[0108] At block 815, if the exchange administrator decides to reject the approval request, it can update the "Request Status" and "Rejection Reason" fields in the "exchange admin" and "updatedOn" slices of the publishing request DPO, and at block 820, use a rejection message to update the data provider slice of the inventory approval request DPO on remote deployment 602. As part of updating the "Data Provider" slice, the rejection message can correspondingly update the "Request Status" and "Rejection Reason" fields in the "Data Provider" slice of the inventory approval request DPO.

[0109] If, at block 815, the exchange administrator decides to approve the request, it can update the "Request Status" and "Rejection Reason" fields in the "exchange admin" and "updatedOn" slices of the published request DPO, and at block 825, it uses an implementation message to update the data provider slice of the manifest-approved request DPO on remote deployment 602. As part of updating the data provider slice, the implementation message can correspondingly update the "Request Status" and "Rejection Reason" fields in the "Data Provider" slice of the manifest-approved request DPO.

[0110] Data providers can also utilize cancellation messages, which can remove any approved requests (with pending, approved, or rejected statuses) from the exchange management slice of the Release Request (DPO) on a remote deployment (401). When a data provider releases a list of approved requests, cleanup will "cancel" the request on their behalf using the same code path to remove the request from the exchange administrator's side.

[0111] Figure 9A schematic representation of a machine in an example form of a computer system 900 is shown, wherein a set of instructions causes the machine to perform any or more methods described herein for copying shared objects to a remote deployment. More specifically, the machine can modify a shared object of a first account into a global object, wherein the shared object includes authorization metadata indicating shared authorization for a set of objects in a database. The machine can create a local copy of the shared object on the remote deployment based on the global object in a second account located in the remote deployment, and copy a set of objects from the database to a local database copy on the remote deployment; and refresh the shared authorization for the local copy of the shared object.

[0112] In alternative embodiments, the machine may be connected (e.g., networked) to other machines on a local area network (LAN), intranet, extranet, or the Internet. The machine may operate as a server or client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web appliance, server, network router, switch or bridge, hub, access point, network access control device, or any machine capable of executing (sequentially or otherwise) a set of instructions specifying the action to be taken by the machine. Furthermore, although only a single machine is shown, the term "machine" should also be understood to include any set of machines that individually or in association execute a set (or more) of instructions to implement any one or more of the methods discussed herein. In one embodiment, computer system 900 may represent a server.

[0113] An exemplary computer system 900 includes a processing device 902, a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 918, which communicate with each other via a bus 930. Any signals provided on the various buses described herein may be time-multiplexed with other signals and provided via one or more common buses. Furthermore, interconnections between circuit components or blocks may be represented as buses or single signal lines. Each of the buses may optionally be one or more single signal lines, and each of the single signal lines may optionally be a bus.

[0114] The computing device 900 may further include a network interface device 908 capable of communicating with the network 920. The computing device 900 may also include a video display unit 910 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse), and a sound signal generation device 916 (e.g., a speaker). In one embodiment, the video display unit 910, the alphanumeric input device 912, and the cursor control device 914 may be combined into a single component or device (e.g., an LCD touchscreen).

[0115] Processing device 902 represents one or more general-purpose processing devices, such as microprocessors, central processing units, etc. More specifically, the processing device may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computer (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing combinations of instruction sets. Processing device 902 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processing device 902 is configured to execute data exchange and inventory visibility management instructions 925 to perform the operations and procedures discussed herein.

[0116] Data storage device 918 may include machine-readable storage medium 928 on which one or more sets of data exchange and inventory visibility management instructions 925 (e.g., software) embodying any or more of the functional methods described herein are stored. During execution of the data exchange and inventory visibility management instructions 925 by computer system 900, the data exchange and inventory visibility management instructions 925 may also reside wholly or at least partially in main memory 904 or in processing device 902; main memory 904 and processing device 902 also constitute machine-readable storage media. The data exchange and inventory visibility management instructions 925 may also be transmitted or received on network 920 via network interface device 908.

[0117] As described herein, the machine-readable storage medium 928 can also be used to store instructions for performing methods to determine the functionality to be compiled. Although the machine-readable storage medium 928 is shown as a single medium in exemplary embodiments, the term "machine-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) that store one or more sets of instructions. Machine-readable media includes any mechanism for storing information in a machine-readable (e.g., computer-readable) form (e.g., software, processing applications). Machine-readable media may include, but is not limited to, magnetic storage media (e.g., floppy disks); optical storage media (e.g., CD-ROMs); magneto-optical storage media; read-only memory (ROM); random access memory (RAM); erasable programmable memory (e.g., EPROM and EEPROM); flash memory; or another type of medium suitable for storing electronic instructions.

[0118] Unless otherwise expressly stated, terms such as “receive,” “route,” “grant,” “determine,” “publish,” “provide,” “designate,” “encode,” etc., refer to actions and processes performed or implemented by a computing device that manipulate and convert data represented as physical (electronic) quantities within the registers and memory of the computing device into other data similarly represented as physical quantities within the memory or registers of the computing device or other such information storage, transmission, or display devices. Furthermore, as used herein, the terms “first,” “second,” “third,” “fourth,” etc., are labels used to distinguish different elements and do not necessarily have ordinal meanings based on their numerical names.

[0119] The examples described herein also relate to apparatus for performing the operations described herein. This apparatus may be specifically constructed for a desired purpose, or it may comprise a general-purpose computing device selectively programmed by a computer program stored in a computing device. Such a computer program may be stored in a computer-readable, non-transitory storage medium.

[0120] The methods and illustrative examples described herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used based on the teachings described herein, or it can be demonstrated that it is convenient to construct more specialized devices to perform the desired method steps. As shown in the description above, the desired structures of various such systems will emerge.

[0121] The above description is intended to be illustrative and not restrictive. Although this disclosure has been described with reference to specific illustrative examples, it will be appreciated that this disclosure is not limited to the described examples. The scope of this disclosure should be determined by referring to the appended claims and the full scope of the equivalents given by the claims.

[0122] As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that, when used herein, the terms “comprises,” “comprising,” “includes,” and / or “including” specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Therefore, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0123] It should also be noted that in some alternative implementations, the functions / actions mentioned may not occur in the order shown in the diagram. For example, depending on the functions / actions involved, two diagrams shown consecutively may actually be executed substantially simultaneously, or sometimes in reverse order.

[0124] Although the method operations are described in a specific order, it should be understood that other operations may be performed between the described operations, the described operations may be adjusted so that they occur at slightly different times, or the described operations may be distributed in a system that allows processing operations to occur at various intervals associated with the processing.

[0125] Various units, circuits, or other components can be described or required to be "configured to" or "configurable to" perform one or more tasks. In such a context, the phrase "configured to" or "configurable to" is used to indicate a structure by indicating that the unit / circuit / component includes a structure (e.g., a circuit) that performs one or more tasks during operation. Thus, even when the specified unit / circuit / component is not currently in operation (e.g., not switched on), it can be said that the unit / circuit / component is configured to perform a task, or can be configured to perform a task. Units / circuit / components used with the language "configured to" or "configurable to" include hardware—e.g., circuits, memory storing program instructions that can be executed to perform operations, etc. Declaring that a unit / circuit / component is "configured to" perform one or more tasks, or is "configurable to" perform one or more tasks, is clearly not intended to invoke 35U.SC112, paragraph 6, for that unit / circuit / component. Additionally, "configured to" or "configurable to" can include general-purpose structures (e.g., general-purpose circuits) manipulated by software and / or firmware (e.g., FPGAs or general-purpose processors executing software) to operate in a manner capable of performing the tasks in discussion. "Configured to" can also include adapting a manufacturing process (e.g., a semiconductor manufacturing facility) to manufacture devices (e.g., integrated circuits) suitable for implementing or performing one or more tasks. It is explicitly stated that "configurable to" does not apply to blank media, unprogrammed processors or unprogrammed general-purpose computers, or unprogrammed programmable logic devices, programmable gate arrays, or other unprogrammed devices, unless accompanied by a programmed medium that endows the unprogrammed device with the ability to be configured to perform the disclosed functions.

[0126] Any combination of one or more computer-usable or computer-readable media may be used. For example, computer-readable media may include one or more of portable computer disks, hard disks, random access memory (RAM) devices, read-only memory (ROM) devices, erasable programmable read-only memory (EPROM or flash memory) devices, portable optical disc read-only memory (CDROM), optical storage devices, and magnetic storage devices. Computer program code for performing the operations of this disclosure may be written in any combination of one or more programming languages. Such code may be compiled from source code into computer-readable assembly language or machine code suitable for a device or computer on which the code will be executed.

[0127] The embodiments can also be implemented in a cloud computing environment. In this specification and the appended claims, “cloud computing” can be defined as a model that enables universal, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services), which can be rapidly configured (including via virtualization) and released with minimal management effort or service provider interaction, and then scaled accordingly. The cloud model can consist of various features (e.g., on-demand self-service, broad network access, resource pooling, rapid elasticity, and measurable services), service models (e.g., Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”)), and deployment models (e.g., private cloud, community cloud, public cloud, and hybrid cloud).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, including one or more executable instructions for implementing a specified logical function. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action. These computer program instructions may also be stored in a computer-readable medium that can instruct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of art including instruction means that implement the function / action specified in the flowchart and / or one or more block diagram blocks.

[0129] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments were chosen and described in order to best explain the principles of the embodiments and their practical application, thereby enabling others skilled in the art to best utilize the embodiments and various modifications as may be adapted to the particular uses contemplated. Therefore, these embodiments should be considered illustrative rather than restrictive, and the invention is not limited to the details given herein, but can be modified within the scope and equivalents of the appended claims.

Claims

1. A method comprising: The data exchange is specified as a set of available geographic regions by specifying a string within a first data processing object (DPO). The string includes a location identifier for each geographic region in the set of geographic regions, each geographic region comprising one or more remote deployments. The first DPO extends to store the base DPO of the data exchange, which defines the metadata of the data exchange. Each of the one or more remote deployments corresponds to a deployment shard. Each deployment shard in the geographic region has a similar location identifier for the geographic region. Information related to the data exchange in the geographic region is accessible within the deployment shard. Information related to the data exchange in the geographic region is copied to the new deployment shard in response to the creation of a new deployment shard in the geographic region. Specify that the data list is one or more of the set of geographic regions that are visible; In response to receiving a request to access the data inventory from a remote deployment in one or more geographic regions, a determination is made whether to deny or grant the request, wherein the request to access the data inventory from the remote deployment includes a global message for managing requests to access the data inventory made to a data provider, and wherein the global message updates the underlying DPO of the data exchange using information from the request, wherein the global message includes custom logic for performing the update of the underlying DPO using information from the request; and In response to determining that the request has been fulfilled, the processing device copies the data from the data list to the remote deployment.

2. The method according to claim 1, further comprising: Parse the string to determine if the set of geographic areas where the data exchange is available; Each remote deployment is obtained from each of the set of geographic regions where the data exchange is available; as well as The data exchange is copied to each remote deployment in each of the set of geographic regions where the data exchange is available.

3. The method of claim 1, wherein, Specifying one or more geographic regions in the set of geographic regions includes: The data processing object of the data exchange is modified using a string that includes the location identifier of each of the one or more geographic regions.

4. The method according to claim 3, further comprising: Parse the string to determine if the data list will be one or more visible geographic areas; Each remote deployment is obtained from each of the one or more geographic regions that are visible in the data list described therein; as well as The data list is copied to each remote deployment in each of the one or more geographic regions where the data list will be visible.

5. The method of claim 2, wherein, Copying the data exchange to a remote deployment includes: Generate a global representation of the data exchange; Copy the metadata of the data exchange to the remote deployment; and Based on the global representation, a local copy of the data exchange is created on the remote deployment.

6. The method of claim 1, wherein, The request to access the data list from the remote deployment includes a global message for managing requests to access the data list made to the data provider, wherein the global message updates the data processing object of the data exchange with information from the request.

7. The method of claim 1, further comprising receiving a request for approval from the exchange administrator to publish the data list.

8. The method of claim 7, wherein, The request for exchange administrators to approve the release of the data list includes a global message for managing requests from data providers to approve the release of the data list, wherein the global message uses the information in the request to update the remotely deployed data processing object hosted by the exchange administrator.

9. The method according to claim 3, further comprising: The data list is modified by adding one or more new location identifiers to the string or removing one or more location identifiers from the string, wherein the data list is a set of geographic regions that are visible.

10. A system comprising: Memory; as well as A processing device operatively coupled to the memory, the processing device being used for: The data exchange is specified as a set of available geographic regions by specifying a string within a first data processing object (DPO). The string includes a location identifier for each geographic region in the set of geographic regions, each geographic region comprising one or more remote deployments. The first DPO extends to store the base DPO of the data exchange, which defines the metadata of the data exchange. Each of the one or more remote deployments corresponds to a deployment shard. Each deployment shard in the geographic region has a similar location identifier for the geographic region. Information related to the data exchange in the geographic region is accessible within the deployment shard. Information related to the data exchange in the geographic region is copied to the new deployment shard in response to the creation of a new deployment shard in the geographic region. Specify that the data list is one or more of the set of geographic regions that are visible; In response to receiving a request to access the data inventory from a remote deployment in one or more geographic regions, a determination is made whether to deny or grant the request, wherein the request to access the data inventory from the remote deployment includes a global message for managing requests to access the data inventory made to a data provider, and wherein the global message updates the underlying DPO of the data exchange using information from the request, wherein the global message includes custom logic for performing the update of the underlying DPO using information from the request; and In response to determining that the request has been fulfilled, the data in the data inventory is copied to the remote deployment.

11. The system of claim 10, wherein, The processing equipment is also used for: Parse the string to determine if the set of geographic areas where the data exchange is available; Each remote deployment is obtained from each of the set of geographic regions where the data exchange is available; as well as The data exchange is copied to each remote deployment in each of the set of geographic regions where the data exchange is available.

12. The system of claim 10, wherein, In order to specify one or more geographic regions in the set of geographic regions, the processing device is configured to: The data processing object of the data exchange is modified using a string that includes the location identifier of each of the one or more geographic regions.

13. The system of claim 12, wherein, The processing equipment is also used for: Parse the string to determine if the data list will be one or more visible geographic areas; Each remote deployment is obtained from each of the one or more geographic regions that are visible in the data list described therein; as well as The data list is copied to each remote deployment in each of the one or more geographic regions where the data list will be visible.

14. The system of claim 11, wherein, In order to replicate the data exchange, the processing equipment is used for: Generate a global representation of the data exchange; Copy the metadata of the data exchange to the remote deployment; and Based on the global representation, a local copy of the data exchange is created on the remote deployment.

15. The system of claim 10, wherein, The request to access the data list from the remote deployment includes a global message for managing requests to access the data list made to the data provider, wherein the global message updates the data processing object of the data exchange with information from the request.

16. The system of claim 10, wherein, The processing equipment is also used for: Receive a request for the exchange administrator to approve the release of the data list.

17. The system of claim 16, wherein, The request for exchange administrators to approve the release of the data list includes a global message for managing requests from data providers to approve the release of the data list, wherein the global message uses the information in the request to update the remotely deployed data processing object hosted by the exchange administrator.

18. The system of claim 12, wherein, The processing equipment is also used for: The data list is modified by adding one or more new location identifiers to the string or removing one or more location identifiers from the string, wherein the data list is a set of geographic regions that are visible.

19. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the processing device, when executed by a processing device, to: The data exchange is specified as a set of available geographic regions by specifying a string within a first data processing object (DPO). The string includes a location identifier for each geographic region in the set of geographic regions, each geographic region comprising one or more remote deployments. The first DPO extends to store the base DPO of the data exchange, which defines the metadata of the data exchange. Each of the one or more remote deployments corresponds to a deployment shard. Each deployment shard in the geographic region has a similar location identifier for the geographic region. Information related to the data exchange in the geographic region is accessible within the deployment shard. Information related to the data exchange in the geographic region is copied to the new deployment shard in response to the creation of a new deployment shard in the geographic region. Specify that the data list is one or more of the set of geographic regions that are visible; In response to receiving a request to access the data inventory from a remote deployment in one or more geographic regions, a determination is made whether to deny or grant the request, wherein the request to access the data inventory from the remote deployment includes a global message for managing requests to access the data inventory made to a data provider, and wherein the global message updates the underlying DPO of the data exchange using information from the request, wherein the global message includes custom logic for performing the update of the underlying DPO using information from the request; and In response to determining that the request has been fulfilled, the processing device copies the data from the data list to the remote deployment.

20. The non-transitory computer-readable medium of claim 19, wherein, The processing equipment is also used for: Parse the string to determine if the set of geographic areas where the data exchange is available; Each remote deployment is obtained from each of the set of geographic regions where the data exchange is available; as well as The data exchange is copied to each remote deployment in each of the set of geographic regions where the data exchange is available.

21. The non-transitory computer-readable medium according to claim 19, wherein, In order to specify one or more geographic regions in the set of geographic regions, the processing device is configured to: The data processing object of the data exchange is modified using a string that includes the location identifier of each of the one or more geographic regions.

22. The non-transitory computer-readable medium of claim 21, wherein, The processing equipment is also used for: Parse the string to determine if the data list will be one or more visible geographic areas; Each remote deployment is obtained from each of the one or more geographic regions that are visible in the data list described therein; as well as The data list is copied to each remote deployment in each of the one or more geographic regions where the data list will be visible.

23. The non-transitory computer-readable medium of claim 20, wherein, In order to replicate the data exchange, the processing equipment is used for: Generate a global representation of the data exchange; Copy the metadata of the data exchange to the remote deployment; and Based on the global representation, a local copy of the data exchange is created on the remote deployment.

24. The non-transitory computer-readable medium of claim 19, wherein, The request to access the data list from the remote deployment includes a global message for managing requests to access the data list made to the data provider, wherein the global message updates the data processing object of the data exchange with information from the request.

25. The non-transitory computer-readable medium of claim 19, wherein, The processing equipment is also used for: Receive a request for the exchange administrator to approve the release of the data list.

26. The non-transitory computer-readable medium of claim 25, wherein, The request for exchange administrator approval to publish the data list includes a global message for managing requests from data providers to approve the publication of the data list, wherein the global message uses information from the request to update remotely deployed data processing objects hosted by the exchange administrator.

27. The non-transitory computer-readable medium of claim 25, wherein, The processing equipment is also used for: The data list is modified by adding one or more new location identifiers to the string or removing one or more location identifiers from the string, wherein the data list is a set of geographic regions that are visible.