A data product identity ID generation method based on a data grid

By constructing a communication architecture using BGP, IS-IS, and EBGP protocols in the data grid, and designing a distributed, multi-layered data product identity ID, the uniqueness and accessibility issues of data products in different domains and subdomains are resolved, enabling efficient cross-domain data transmission and sharing.

CN117435826BActive Publication Date: 2026-08-04江苏量界数据科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
江苏量界数据科技有限公司
Filing Date
2023-09-18
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies cannot design unique and cross-domain and subdomain-wide identity identifiers for data products in data grids, resulting in data products being unable to effectively communicate and access each other in different domains and subdomains, making it difficult to guarantee the global uniqueness and accessibility of data.

Method used

The system communication architecture is constructed using the BGP routing protocol. Combined with IS-IS and EBGP protocols, a distributed multi-level data product identity ID is designed, including identity IDs at the domain, subdomain, and sub-subdomain levels. A URL design style is adopted to ensure the global uniqueness and readability of the IDs, enabling cross-domain communication.

Benefits of technology

The generated identity ID is globally unique within the data grid, reducing the risk of conflicts, improving the accessibility and interoperability of data products, supporting cross-domain communication, ensuring the flow and sharing of data within the grid, and possessing reliability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117435826B_ABST
    Figure CN117435826B_ABST
Patent Text Reader

Abstract

This invention discloses a data product identity ID generation method based on a data grid, belonging to the field of data processing technology. It includes constructing a system communication architecture based on a data grid, employing a distributed, multi-layered data product identity ID design within the communication architecture. This solves the technical problem of achieving efficient ID generation algorithms and query mechanisms using data grid-based data product identity ID generation. This invention ensures that the generated identity ID is globally unique within the data grid, improving user understanding and interpretability of its meaning, facilitating the expression of relationships and dependencies between data products, possessing strong versatility, improving interoperability between the system and applications, allowing for flexible expansion and customization, ensuring efficient data transmission and low-latency communication, while also possessing reliability. The design employs the BGP routing protocol, enabling the selection of the optimal path, avoiding loops, and achieving data communication across autonomous systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology and relates to a method for generating data product identity IDs based on data grids. Background Technology

[0002] Data grids are a decentralized data architecture designed to address the bottlenecks and limitations of traditional centralized data management systems. Data grids treat data as a product and emphasize the autonomy and distributed nature of data products.

[0003] The existing technology has the following problems:

[0004] 1. Data product identity: It is impossible to design a unique identity for data products in the data grid that can span multiple domains and subdomains, and it is impossible to ensure the global uniqueness and accessibility of data products.

[0005] 2. It is impossible to achieve effective communication and access between data products in different fields and subdomains while also meeting the requirements of distributed multi-level identity IDs, making it difficult to guarantee the flow and sharing of data in the data grid. Summary of the Invention

[0006] The purpose of this invention is to provide a data product identity ID generation method based on data grids, which solves the technical problem of achieving an efficient ID generation algorithm and query mechanism using data product identity ID generation based on data grids.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A method for generating data product identity IDs based on data grids includes the following steps:

[0009] Step 1: Build a system communication architecture based on a data grid using the BGP routing protocol to realize cross-domain communication of distributed multi-level data product identity IDs, including determining domains and subdomains, using the BGP routing protocol to communicate between different domains or subdomains, and deploying communication equipment;

[0010] Step 2: Adopt a distributed, multi-tiered data product identity ID design in the communication design architecture, specifically including the following steps:

[0011] Step 2-1: Assign a unique identity ID to each data product. The identity ID includes <domain + subdomain + sub-subdomain + sub-sub-subdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>.

[0012] Step 2-2: Within each domain, use the IS-IS protocol to enable access communication for data products within the data grid;

[0013] Steps 2-3: Use the BGP protocol to achieve communication between data products across different domains;

[0014] Step 3: After completing the system communication architecture, test and deploy the system communication architecture to ensure that the data products can communicate normally and that the identity ID is valid.

[0015] Preferably, the domain and subdomain also represent different departments or fields.

[0016] Preferably, when performing step 1, the BGP routing protocol is used to construct a system communication architecture based on a data grid for distributed multi-level data product identity ID communication, specifically including the following steps:

[0017] Step 1-1: Identify the domains and subdomains. Each domain and subdomain represents a different organization, and each domain or subdomain has its own data products.

[0018] Steps 1-2: Data communication between different domains or subdomains is carried out using the BGP routing protocol;

[0019] Steps 1-3: Deploy communication devices within each domain or subdomain for data products to interact between domains or subdomains.

[0020] Preferably, when performing step 2-1, the specific meaning of the identity ID, including <domain + subdomain + sub-subdomain + sub-sub-subdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>, is as follows:

[0021] IP + port: Represents the service address registered and published in the subnet where the data grid is located. Each data product is assigned a specific IP address and port for external communication and access.

[0022] Domain Name: This indicates that each data product is domain-autonomous;

[0023] Data Product Name: Each data product has a unique data product name. When registering a data product on the self-service platform, there is no data product with the same name in the current subnet, and there is no data product with the same name in the entire data grid.

[0024] Processing time / type / other identifying parameters: In the data grid architecture design, the same data product has different versions. The distinction between different versions is based on the processing time. That is, data products generated at different processing times are classified as a new version, and the version increases with the processing time.

[0025] Domain: A shared domain name is set for data products within the same local area network;

[0026] Subdomains, sub-subdomains, and sub-subdomains: Subdomains can have N levels, where N is a positive integer. Domains are used to define the first level of data products, providing the highest level of classification and differentiation. Data products are further subdivided and differentiated using the domain level of subdomains, sub-subdomains, and sub-subdomains.

[0027] Preferably, during step 2-2, communication access between data products within the same subdomain is achieved through IS-IS three-level access.

[0028] Preferably, during steps 2-3, data product communication across regions, between different regions, and between different fields is implemented using the EBGP protocol.

[0029] This invention discloses a data product identity ID generation method based on a data grid. It solves the technical problem of achieving efficient ID generation algorithms and query mechanisms using a data grid-based data product identity ID generation method. This invention ensures that the generated identity ID is globally unique within the data grid, thereby reducing the risk of conflicts. The URL design style makes the data product identity ID more intuitive and easier to understand, improving user comprehension and interpretability. The URL design style reflects the hierarchical structure between resources in the identity ID, facilitating referencing and querying, and helping to express the relationships and dependencies between data products. It has strong versatility, improves interoperability between systems and applications, and is flexibly expandable and customizable to meet the scalability needs of different business scenarios. The BGP routing protocol-based communication architecture achieves the global and distributed nature of data products, enabling effective communication and access between data products in different domains and subdomains, promoting data flow and sharing within the data grid. The communication architecture design considers the cross-domain communication needs of data products to ensure efficient data transmission and low-latency communication, while also ensuring reliability. The BGP routing protocol design can select the optimal path, avoid loops, and achieve data communication across autonomous systems. Attached Figure Description

[0030] Figure 1 This is the main flowchart of the present invention;

[0031] Figure 2 This is a schematic diagram of the data grid of the present invention;

[0032] Figure 3 This is a schematic diagram of the communication topology of the data products within the grid according to the present invention;

[0033] Figure 4 This is a schematic diagram of the inter-mesh data product communication topology of the present invention;

[0034] Figure 5 This invention relates to a cross-regional communication architecture for data grid products based on inter-grid and intra-grid communication. Detailed Implementation

[0035] like Figures 1-5 The method for generating data product identity IDs based on data grids includes the following steps:

[0036] Step 1: Construct a system communication architecture based on a data grid using the BGP routing protocol to enable cross-domain communication of distributed multi-level data product identity IDs. This includes determining domains and subdomains, using the BGP routing protocol to communicate between different domains or subdomains, and deploying communication equipment. Domains and subdomains also represent different departments or fields.

[0037] Specifically, the steps include the following:

[0038] Step 1-1: Identify the domains and subdomains. Each domain and subdomain represents a different organization, and each domain or subdomain has its own data products.

[0039] Steps 1-2: Data communication between different domains or subdomains is carried out using the BGP routing protocol;

[0040] Steps 1-3: Deploy communication devices within each domain or subdomain for data products to interact between domains or subdomains.

[0041] Step 2: Adopt a distributed, multi-tiered data product identity ID design in the communication design architecture, specifically including the following steps:

[0042] Step 2-1: Assign a unique identity ID to each data product. The identity ID includes <domain + subdomain + sub-subdomain + sub-sub-subdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>.

[0043] The specific meanings of the identity ID, including <domain + subdomain + sub-subdomain + sub-sub-subdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>, are as follows:

[0044] IP + port: Represents the service address registered and published in the subnet where the data grid is located. Each data product is assigned a specific IP address and port for external communication and access.

[0045] Domain Name: This indicates that each data product is domain-autonomous;

[0046] Data Product Name: Each data product has a unique data product name. When registering a data product on the self-service platform, there is no data product with the same name in the current subnet, and there is no data product with the same name in the entire data grid.

[0047] Processing time / type / other identifying parameters: In the data grid architecture design, the same data product has different versions. The distinction between different versions is based on the processing time. That is, data products generated at different processing times are classified as a new version, and the version increases with the processing time.

[0048] Domain: A shared domain name is set for data products within the same local area network;

[0049] Subdomains, sub-subdomains, and sub-subdomains: Subdomains can have N levels, where N is a positive integer. Domains are used to define the first level of data products, providing the highest level of classification and differentiation. Data products are further subdivided and differentiated using the domain level of subdomains, sub-subdomains, and sub-subdomains.

[0050] In this embodiment, the identity ID is designed using a URL design style. A URL stands for Uniform Resource Locator, also known as a web address, and is the standard address of a resource on the Internet. Every file on the Internet has a unique URL, which contains information indicating the file's location and how the browser should handle it.

[0051] A URL is a core concept on the Web. It is the mechanism by which browsers retrieve any resources published on the Web.

[0052] The URL consists of: Scheme\Domain Name\Port\Path to the file\Parameters\Anchor.

[0053] Data grids are domain-oriented data management architectures, where each domain contains its own published data products, as follows: Figure 2 As shown.

[0054] The modules in each domain of the data grid form a dynamic topological relationship.

[0055] Each domain is relatively autonomous, and globally accessible (i.e., access to data products).

[0056] Data products are required to have a globally unique identifier in the grid, meaning that each data product has a unique address in the dynamic grid.

[0057] Data products can be addressed in the data grid through a self-service platform, meaning that data products can access each other.

[0058] This invention uses a URL-based design style to uniquely identify each data product, offering the following advantages for data grid implementation:

[0059] Uniqueness Guarantee: URLs are globally unique; each URL has a unique identifier. By adopting a URL design style, a unique identity ID can be generated for each data product, reducing the risk of conflicts.

[0060] Readability and understandability: URLs typically represent resources in a readable and understandable format. For generated data product identity IDs, adopting a URL design style can make the ID more intuitive and easier to understand, improving user comprehension and interpretability.

[0061] Relationships and Dependencies: URL design styles typically use paths, query parameters, etc., to represent the relationships and dependencies between resources. For data products, using URL design styles can reflect the hierarchical structure between resources in the ID, facilitating referencing and querying.

[0062] Standardization and Interoperability: URLs are a universal standard on the internet, and almost all systems and applications support parsing and processing them. Adopting a URL-based design style can improve interoperability between systems and applications, making the identity ID of data products more convenient and flexible to use in different environments.

[0063] Scalability and flexibility: The URL design style can be flexibly extended and customized. By adding additional information to the URL path or query parameters, more metadata can be added to the identity ID of the data product, meeting the scalability needs of different business scenarios.

[0064] In this embodiment, in addition to adopting the URL design style, the NAT-DDNS design principle is also added. That is, a common domain name is set for data products in the same local area network. The identity ID of the data product is upgraded to: domain + IP + port + [domain name] + data product name + processing time / type / other identification parameters. At the same time, each domain is added to the DDNS so that external data products can be accessed through the domain name.

[0065] In this embodiment, a higher-level identity definition for data products is achieved, that is, the data products are published to the entire network and truly have a globally unique identity. The mutual discovery and accessibility of data products in the data grid are guaranteed, and a real and specific distributed data grid implementation path can be obtained.

[0066] Step 2-2: Within each domain, the IS-IS protocol is used to enable access and communication of data products within the data grid; communication and access between data products within the same subdomain are achieved through IS-IS three-level access.

[0067] IS-IS (Intermediate System to Intermediate System) is a protocol for routing and forwarding that ensures that data products within the same domain can discover and access each other.

[0068] like Figure 3 As shown, this embodiment sets up three different levels of communication devices: Level-1 routing devices are deployed within the area, Level-2 routing devices are deployed between areas, and Level-1-2 routing devices are deployed between the Level-1 and Level-2 routing devices. A dynamic mesh is used, and the inter-area (backbone) includes not only all Level-2 routing devices in Area 1 but also Level-1-2 routing devices in other areas.

[0069] Communication between data products within the same subdomain is achieved through IS-IS three-level access.

[0070] L2 represents different organizations, fields, or departments within the same region.

[0071] L1 represents the specific data product under the corresponding organization, field, or department.

[0072] Intercommunication between data products in L1 is routed and forwarded through L1-2.

[0073] Steps 2-3: Use the BGP protocol to realize communication of data products across different domains; communication of data products across regions, between different regions, and between different fields is realized by nesting the EBGP protocol.

[0074] BGP (Border Gateway Protocol) is a routing protocol used for route reachability and optimal path selection between different domains. Through BGP, data products within different domains can achieve cross-domain communication.

[0075] like Figure 4 The diagram shows the inter-mesh routing protocol relationship based on BGP mode. Border Gateway Protocol (BGP) is a distance-vector routing protocol that enables route reachability between Autonomous Systems (AS) and selects the best route. It includes IBGP (Interior Gateway Protocol) and EBGP (Exterior Gateway Protocol).

[0076] To facilitate the management of the ever-expanding grid, the grid was divided into different autonomous systems.

[0077] An AS (Application Server) refers to an IP network that shares the same routing policy within the same grid. Each AS in the grid is assigned a unique AS number to distinguish it from other ASes.

[0078] Data products within the grid have fast addressing speed and clear hierarchy, but only three levels of access can be designed, meaning that access can only be made to products in different domains within the grid.

[0079] A multi-level gateway access is achieved by adopting an inter-mesh data product communication design.

[0080] Similar to the region, the grid hierarchy consists of domains, departments, and data products. Routing access between departments and data products can be achieved using IBGP (IS-IS protocol); communication between data products in different domains is achieved using EBGP.

[0081] In cross-regional scenarios, communication between data products in different regions can still be achieved through EBGP nesting.

[0082] like Figure 5 The diagram illustrates a BGP-based cross-regional communication architecture design for data products across data meshes. This architecture includes two autonomous regions, AS100 and AS200 (which can be within the same region or different regions). Communication between data products in different domains within AS100 is achieved through IBGP (specifically through IS-IS). Communication between data products in different autonomous regions between AS100 and AS200 is achieved through EBGP. BGP can be nested, and its specific implementation needs to be tailored to the size and architecture of the organization.

[0083] This embodiment ultimately enhances the identity ID, namely the ID design pattern of <domain + IP + port + [domain name] + data product name + processing time / type / other identifier parameters>, to <domain + subdomain + subsubdomain + subsubsubdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifier parameters>.

[0084] Domain names define the first level of a data product, providing the highest level of classification and differentiation, and may represent a specific organization or company. By using subdomains, sub-subdomains, and other levels in the domain name hierarchy, data products can be further subdivided and differentiated. These subdomains can represent different departments, teams, business lines, etc., to more accurately locate and identify data products. Combining IP addresses and ports with a unique ID for the data product gives each specific data product a unique identifier within the network. Each data product is assigned a specific IP address and port for external communication and access.

[0085] Step 3: After completing the system communication architecture, test and deploy the system communication architecture to ensure that the data products can communicate normally and that the identity ID is valid.

[0086] In this embodiment, step 3 may specifically include testing communication between different domains to verify that the data product can correctly address and transmit data, deploying the communication design architecture to the actual data grid environment to ensure that all organizations or departments can participate in data sharing and communication, establishing a monitoring mechanism to monitor the performance of communication and data transmission at any time, and performing necessary maintenance and upgrades, etc.

[0087] This invention discloses a data product identity ID generation method based on a data grid. It solves the technical problem of achieving efficient ID generation algorithms and query mechanisms using a data grid-based data product identity ID generation method. This invention ensures that the generated identity ID is globally unique within the data grid, thereby reducing the risk of conflicts. The URL design style makes the data product identity ID more intuitive and easier to understand, improving user comprehension and interpretability. The URL design style reflects the hierarchical structure between resources in the identity ID, facilitating referencing and querying, and helping to express the relationships and dependencies between data products. It has strong versatility, improves interoperability between systems and applications, and is flexibly expandable and customizable to meet the scalability needs of different business scenarios. The BGP routing protocol-based communication architecture achieves the global and distributed nature of data products, enabling effective communication and access between data products in different domains and subdomains, promoting data flow and sharing within the data grid. The communication architecture design considers the cross-domain communication needs of data products to ensure efficient data transmission and low-latency communication, while also ensuring reliability. The BGP routing protocol design can select the optimal path, avoid loops, and achieve data communication across autonomous systems.

Claims

1. A method for generating data product identity IDs based on data grids, characterized in that: Includes the following steps: Step 1: Build a system communication architecture based on a data grid using the BGP routing protocol to realize cross-domain communication of distributed multi-level data product identity IDs, including determining domains and subdomains, using the BGP routing protocol to communicate between different domains or subdomains, and deploying communication equipment; Step 2: Adopt a distributed, multi-tiered data product identity ID design in the communication design architecture, specifically including the following steps: Step 2-1: Assign a unique identity ID to each data product. The identity ID includes <domain + subdomain + sub-subdomain + sub-sub-subdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>. Step 2-2: Within each domain, use the IS-IS protocol to enable access communication for data products within the data grid; Steps 2-3: Use the BGP protocol to enable communication between data products across different domains; Step 3: After completing the system communication architecture, test and deploy the system communication architecture to ensure that the data products can communicate normally and that the identity ID is valid.

2. The data product identity ID generation method based on data grid as described in claim 1, characterized in that: Domains and subdomains also represent different departments or fields.

3. The data product identity ID generation method based on data grid as described in claim 1, characterized in that: When performing step 1, the BGP routing protocol is used to construct a system communication architecture based on a data grid for distributed multi-level data product identity ID communication. The specific steps include the following: Step 1-1: Identify the domains and subdomains. Each domain and subdomain represents a different organization, and each domain or subdomain has its own data products. Steps 1-2: Data communication between different domains or subdomains is carried out using the BGP routing protocol; Steps 1-3: Deploy communication devices within each domain or subdomain for data products to interact between domains or subdomains.

4. The data product identity ID generation method based on data grid as described in claim 1, characterized in that: When performing step 2-1, the specific meaning of the identity ID, which includes <domain + subdomain + subsubdomain + subsubsubdomain + ... + IP + port + [domain name] + data product name + processing time / type / other identifying parameters>, is as follows: IP + port: Represents the service address registered and published in the subnet where the data grid is located. Each data product is assigned a specific IP address and port for external communication and access. Domain Name: This indicates that each data product is domain-autonomous; Data Product Name: Each data product has a unique data product name. When registering a data product on the self-service platform, there is no data product with the same name in the current subnet, and there is no data product with the same name in the entire data grid. Processing time / type / other identifying parameters: In the data grid architecture design, the same data product has different versions. The distinction between different versions is based on the processing time. That is, data products generated at different processing times are classified as a new version, and the version increases with the processing time. Domain: A shared domain name is set for data products within the same local area network; Subdomains, sub-subdomains, and sub-subdomains: Subdomains can have N levels, where N is a positive integer. Domains are used to define the first level of data products, providing the highest level of classification and differentiation. Data products are further subdivided and differentiated using the domain level of subdomains, sub-subdomains, and sub-subdomains.

5. The data product identity ID generation method based on data grid as described in claim 1, characterized in that: When performing step 2-2, communication access between data products within the same subdomain is achieved through IS-IS three-level access.

6. The data product identity ID generation method based on data grid as described in claim 1, characterized in that: During steps 2-3, data product communication across regions, between different regions, and between different fields is achieved through nested EBGP protocols.