Multi-cloud multi-region data platform construction method and system based on distributed architecture

By employing a distributed architecture approach, regional adaptability assessment and role allocation are conducted. Combined with pattern matching and link planning, the problems of regional selection and component deployment in the construction of a multi-cloud, multi-region data platform are solved, enabling efficient operation and reasonable layout of the platform.

CN121967412APending Publication Date: 2026-05-01ZHEJIANG SHUXIN NETWORK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG SHUXIN NETWORK CO LTD
Filing Date
2026-04-03
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing multi-cloud, multi-region data platform construction methods have shortcomings in terms of region selection and adaptability assessment, component deployment, and communication configuration, resulting in an unreasonable platform architecture that affects service effectiveness and performance.

Method used

By employing a distributed architecture approach, regional adaptability assessment, role allocation, pattern matching, and link planning are performed to generate component deployment configurations and cross-regional interaction channels, ensuring the platform's efficient operation.

Benefits of technology

It enables precise layout and efficient operation of multi-cloud and multi-region data platforms, solves the shortcomings of traditional technologies in regional planning, component deployment and service instantiation, and provides technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967412A_ABST
    Figure CN121967412A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-cloud multi-region data platform construction method and system based on a distributed architecture, and the method and system achieve the precise planning of layout through adaptability analysis and role distribution. And a deployment mechanism is constructed, and a reliable platform architecture strategy is established in combination with mode matching and link planning. Service optimization is introduced, and efficient operation of the platform is ensured through instantiation deployment and channel configuration. According to the method, the defects of the traditional technology in the aspects of regional planning, component deployment, service instantiation and the like are effectively overcome, and a technical guarantee is provided for construction of a multi-cloud multi-regional data platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, specifically to a method and system for constructing a multi-cloud, multi-region data platform based on a distributed architecture. Background Technology

[0002] Existing methods for building multi-cloud, multi-region data platforms have significant shortcomings. Traditional systems perform poorly in terms of region selection and adaptability assessment, failing to effectively achieve a reasonable platform layout and impacting service performance.

[0003] Furthermore, existing technologies suffer from bottlenecks in component deployment and communication configuration. Most systems lack robust pattern matching mechanisms and link planning strategies, resulting in inadequate platform architecture.

[0004] The existing system has technical shortcomings in service instantiation. A lack of in-depth analysis of regional characteristics makes it difficult to achieve efficient resource allocation through role assignment, impacting platform performance. Solving these problems is crucial for improving the data platform's construction capabilities. Summary of the Invention

[0005] To address the problems in existing technologies, this application provides a method and system for constructing a multi-cloud, multi-region data platform based on a distributed architecture. This method effectively solves the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, and provides technical support for the construction of multi-cloud, multi-region data platforms.

[0006] To solve at least one of the above problems, this application provides the following technical solution:

[0007] Firstly, this application provides a method for constructing a multi-cloud, multi-region data platform based on a distributed architecture, including:

[0008] Based on the regional distribution information of multiple cloud service providers, a candidate region set is generated by performing a regional adaptability assessment according to preset regional filtering rules. The candidate region set is then filtered according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table. Each region in the target region configuration table is weighted and scored according to preset center selection rules to generate a regional role allocation scheme.

[0009] Based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme, a platform deployment mode identifier is generated by pattern matching according to a preset constraint model. The platform deployment mode identifier is associated and mapped with a preset mode configuration template to obtain a component deployment configuration set. The component deployment configuration set is instantiated and expanded according to the regional role allocation scheme to generate a platform component deployment list. The platform component deployment list is used to plan communication links according to a preset communication rule base to generate cross-regional interaction channel configuration.

[0010] Based on the platform component deployment list, services are deployed in the central management area according to a preset mode to generate central service instances. The central service instances are then connected to the central nodes according to the cross-regional interaction channel configuration. Based on the platform component deployment list, services are deployed in each subordinate region according to a preset mode to generate regional service instances. The regional service instances and the central nodes are then connected to establish a communication link to generate a multi-cloud, multi-regional data platform.

[0011] Furthermore, it also includes: collecting regional distribution information of multiple cloud service providers and extracting the geographic location identifier, available service type and resource quota parameters of each region according to preset information parsing rules to generate a set of regional attribute information; and performing field mapping and format standardization processing on the set of regional attribute information according to a preset structured template to generate a standardized regional attribute table.

[0012] For each region in the standardized regional attribute table, service type matching degree is calculated and resource quota sufficiency is determined according to preset regional filtering rules to generate a regional adaptability score sequence. The regional adaptability score sequence is then filtered according to preset scoring threshold conditions to generate a candidate region set.

[0013] Furthermore, it also includes: performing compliance attribute matching verification on each region in the candidate region set according to data compliance requirements to generate a compliance verification result set; and performing bandwidth capacity and transmission delay determination on the regions that pass the verification in the compliance verification result set according to network bandwidth constraints to generate a target region configuration table.

[0014] The network connectivity index and resource carrying capacity index of each region in the target region configuration table are weighted and scored according to the preset center selection rules to generate a regional comprehensive score sequence. The regional comprehensive score sequence is then used to assign roles to the central management region and subordinate regions according to the preset role division threshold conditions to generate a regional role allocation scheme.

[0015] Furthermore, it also includes: extracting global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters from the regional role allocation scheme, and vectorizing them according to preset parameter encoding rules to generate constraint parameter vectors; calculating the similarity between the constraint parameter vectors and the pattern feature vectors in the preset constraint model, and selecting the optimal pattern according to preset matching threshold conditions to generate a platform deployment mode identifier;

[0016] The platform deployment mode identifier is indexed and queried against the preset mode configuration template library to obtain the corresponding mode configuration template. The data development platform deployment location parameters and data asset management platform deployment location parameters in the mode configuration template are bound to the region according to the regional role allocation scheme to generate a component deployment configuration set.

[0017] Furthermore, it also includes: instantiating and expanding the deployment location parameters of each component in the component deployment configuration set according to the regional identifier in the regional role allocation scheme to generate a component instance configuration sequence, and assembling the component instance configuration sequence in a structured manner according to a preset list encapsulation rule to generate a platform component deployment list;

[0018] The regional affiliation and functional type of each component instance in the platform component deployment list are extracted, and a communication requirement matrix is ​​generated by analyzing the communication requirements between components according to a preset communication rule library. The communication requirement matrix is ​​then configured with a dual link of synchronous call channel and asynchronous message channel according to a preset link planning rule to generate a cross-regional interaction channel configuration.

[0019] Furthermore, it also includes: extracting the component configuration parameters corresponding to the central management area from the platform component deployment list and deploying services according to a preset mode to generate a central service instance; and registering the service endpoint and initializing the service status of the central service instance according to a preset service registration rule to generate a registered central service instance.

[0020] Extract the synchronous call channel configuration and asynchronous message channel configuration related to the registered center service instance from the cross-regional interaction channel configuration, bind the channel port and adapt the communication protocol to generate the channel access configuration, and apply the channel access configuration to the registered center service instance to load cross-regional communication capabilities and generate the center node.

[0021] Furthermore, it also includes: extracting the component configuration parameters corresponding to each subordinate region from the platform component deployment list and instantiating and deploying the data storage computing cluster and unified engine gateway service according to the preset cluster deployment rules to generate regional service instances; and registering the service endpoints and initializing the service status of the regional service instances according to the preset service registration rules to generate registered regional service instances.

[0022] Extract the channel configuration parameters related to the registered regional service instances from the cross-regional interaction channel configuration, and perform bidirectional channel handshake and link connectivity verification with the central node to generate an established communication link. Then, assemble the established communication link with the central node and each of the registered regional service instances according to the preset platform encapsulation rules to generate a multi-cloud, multi-regional data platform.

[0023] Secondly, this application provides a multi-cloud, multi-region data platform construction system based on a distributed architecture, including:

[0024] The region allocation module is used to generate a candidate region set by performing a region adaptability assessment based on the regional distribution information of multiple cloud service providers and according to preset region filtering rules. The candidate region set is then filtered according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table. Each region in the target region configuration table is weighted and scored according to preset center selection rules to generate a region role allocation scheme.

[0025] The cross-regional interaction module is used to generate a platform deployment mode identifier by performing pattern matching according to a preset constraint model based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme. The platform deployment mode identifier is associated and mapped with a preset mode configuration template to obtain a component deployment configuration set. The component deployment configuration set is instantiated and expanded according to the regional role allocation scheme to generate a platform component deployment list. The platform component deployment list is used to generate a cross-regional interaction channel configuration by planning communication links according to a preset communication rule base.

[0026] The unified deployment module is used to deploy services in the central management area according to a preset mode based on the platform component deployment list to generate central service instances, to access the central service instances through the cross-regional interaction channel configuration to generate central nodes, to deploy services in each subordinate area according to a preset mode based on the platform component deployment list to generate regional service instances, and to establish communication links between the regional service instances and the central nodes to generate a multi-cloud, multi-regional data platform.

[0027] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture.

[0028] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture.

[0029] Fifthly, this application provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture.

[0030] As can be seen from the above technical solutions, this application provides a method and system for constructing a multi-cloud, multi-region data platform based on a distributed architecture. Through adaptability analysis and role allocation, it achieves precise layout planning. A deployment mechanism is constructed, combining pattern matching and link planning to establish a reliable platform architecture strategy. Service optimization is introduced, ensuring efficient platform operation through instantiation deployment and channel configuration. This method effectively solves the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, providing technical support for the construction of multi-cloud, multi-region data platforms. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture in this application embodiment;

[0033] Figure 2 This is a structural diagram of the multi-cloud, multi-region data platform construction system based on a distributed architecture in the embodiments of this application. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.

[0036] To address the shortcomings of existing technologies, this application provides a method and system for constructing a multi-cloud, multi-region data platform based on a distributed architecture. Through adaptability analysis and role allocation, it achieves precise layout planning. A deployment mechanism is built, combining pattern matching and link planning to establish a reliable platform architecture strategy. Service optimization is introduced, ensuring efficient platform operation through instantiation deployment and channel configuration. This method effectively solves the deficiencies of traditional technologies in areas such as region planning, component deployment, and service instantiation, providing technical support for the construction of multi-cloud, multi-region data platforms.

[0037] To effectively address the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, and to provide technical support for the construction of multi-cloud, multi-region data platforms, this application provides an embodiment of a method for constructing a multi-cloud, multi-region data platform based on a distributed architecture. See [link to embodiment]. Figure 1 The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture specifically includes the following:

[0038] Step S101: Based on the regional distribution information of multiple cloud service providers, perform regional adaptability assessment according to preset regional filtering rules to generate a candidate region set. Filter the candidate region set according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table. Perform weighted scoring on each region in the target region configuration table according to preset center selection rules to generate a regional role allocation scheme.

[0039] This embodiment first collects regional distribution information from multiple cloud service providers, summarizing and reading the publicly released regional lists and service capability lists of each provider. For each region, three attribute fields are extracted: geographic location identifier, available service type, and resource quota parameters. These fields are then structured and organized according to preset information parsing rules to generate a set of regional attribute information.

[0040] After the set of regional attribute information is constructed, it undergoes format standardization. Since different cloud service providers use different naming and encoding methods for regional attributes, the geographic location identifiers of each region are uniformly mapped to standard geographic codes according to a preset structured template. Available service types are uniformly mapped to preset service type enumeration values, and resource quota parameters are uniformly converted to standard units of measurement. After the above mapping and conversion are completed, a standardized regional attribute table is generated. Each record in the standardized regional attribute table corresponds to one region and contains three standardized attribute fields.

[0041] Accordingly, this embodiment performs a suitability assessment on each region in the standardized regional attribute table. Following preset regional filtering rules, the matching degree between the available service types of each region and the service type list required for the enterprise data platform construction is calculated, and the sufficiency of each region's resource quota parameters is determined by comparing them with the enterprise's estimated resource needs. The service type matching degree and resource quota sufficiency indicators are weighted and combined to generate a regional suitability score. This process is repeated across all regions to generate a regional suitability score sequence. The regional suitability score sequence is then filtered according to preset scoring threshold conditions, and regions that meet the score criteria are retained to generate a candidate region set.

[0042] After the candidate region set is generated, this embodiment performs data compliance verification. For each candidate region, its geographic location identifier is matched against the list of permitted deployment regions in the enterprise's data compliance requirements to determine whether the region meets the compliance constraints for data storage and cross-border transmission. Regions that pass the compliance verification proceed to the next round of network condition assessment, while regions that fail the verification are excluded. For regions that pass the compliance verification, bandwidth capacity and transmission latency are determined according to network bandwidth constraints. Regions that simultaneously meet the minimum bandwidth requirement and the maximum latency requirement are retained to generate a target region configuration table.

[0043] Based on the target area configuration table, this embodiment assigns roles to each area. The network connectivity and resource carrying capacity indicators for each area are extracted, and a weighted score is calculated according to a preset center selection rule, denoted as the formula:

[0044] Q = w1·N + w2·R.

[0045] In the formula, Q represents the regional comprehensive score, N represents the network connectivity index between this region and other regions, R represents the resource carrying capacity index of this region, and w1 and w2 are weighting coefficients, both of which are positive numbers. A regional comprehensive score sequence is generated after traversing all regions in the target region configuration table.

[0046] After the comprehensive score sequence of the regions is generated, the central management region and subordinate regions are labeled according to the preset role division threshold conditions. The region with the highest comprehensive score and meeting the central role threshold condition is labeled as the central management region, and the remaining regions are labeled as subordinate regions. The role labeling results of each region are associated and encapsulated with the corresponding region identifier to generate a region role allocation scheme. The region role allocation scheme will be read in the subsequent step S201 for matching the platform deployment mode and generating component configurations.

[0047] Step S102: Based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme, perform pattern matching according to the preset constraint model to generate a platform deployment mode identifier. Associate and map the platform deployment mode identifier with the preset mode configuration template to obtain a component deployment configuration set. Instantiate and expand the component deployment configuration set according to the regional role allocation scheme to generate a platform component deployment list. Perform communication link planning on the platform component deployment list according to the preset communication rule base to generate a cross-regional interaction channel configuration.

[0048] Based on the regional role allocation scheme generated in step S101 above, this embodiment extracts global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters. The global data management requirement parameters characterize the enterprise's need for centralized management of global data, while the regional fault tolerance parameters characterize the enterprise's need for independent operation capabilities in each region. Following preset parameter coding rules, the above two types of parameters are quantized numerically and then horizontally concatenated to generate a constraint parameter vector.

[0049] After the constraint parameter vector is constructed, this embodiment performs pattern matching with a preset constraint model. The preset constraint model pre-stores pattern feature vectors corresponding to three modes: distributed data platform, centralized data platform, and distributed-data-integrated platform. The similarity between the constraint parameter vector and each mode feature vector is calculated one by one. The mode with the highest similarity is selected as the target mode according to a preset matching threshold, generating a platform deployment mode identifier. For example, when the global data management requirement parameter is high and the regional fault tolerance parameter is low, the similarity calculation result tends to match the centralized data platform mode.

[0050] Accordingly, this embodiment indexes and queries the platform deployment mode identifier against the preset mode configuration template library. The preset mode configuration template library stores corresponding mode configuration templates according to the mode identifier. Each template contains two types of configuration items: deployment location parameters for the data development platform and deployment location parameters for the data asset management platform. After obtaining the corresponding mode configuration template, the deployment location parameters in the template are bound to regions according to the region role allocation scheme. When the deployment location parameter points to the center, it is bound to the region identifier corresponding to the center management region; when the deployment location parameter points to a region, it is bound to the region identifier corresponding to each subordinate region. After the above binding is completed, a component deployment configuration set is generated.

[0051] After the component deployment configuration set is generated, this embodiment instantiates and expands it region by region. For the deployment location parameters of each component in the component deployment configuration set, instantiation is performed according to the region identifier in the region role allocation scheme, generating an instance configuration record for each component in each region. After traversing all components and all regions, a component instance configuration sequence is generated. This sequence is then structured and assembled according to a preset list encapsulation rule to generate a platform component deployment list. Each record in the platform component deployment list contains three fields: component type, target region identifier, and instance configuration parameters.

[0052] Based on the platform component deployment list, this embodiment plans communication links. The region affiliation and function type of each component instance in the platform component deployment list are extracted, and the communication requirements between components are analyzed according to a preset communication rule base. Component pairs requiring real-time interaction are marked as synchronous call requirements, and component pairs requiring asynchronous message passing are marked as asynchronous message requirements. These are then aggregated to generate a communication requirement matrix. The row and column indices of the communication requirement matrix are both component instance identifiers, and the matrix element values ​​represent the communication type between corresponding component pairs.

[0053] After the communication demand matrix is ​​constructed, dual-link configuration is performed according to preset link planning rules. For component pairs marked as synchronous call requirements in the communication demand matrix, synchronous call channels are configured and communication protocols and port parameters are specified; for component pairs marked as asynchronous message requirements, asynchronous message channels are configured and message queues and subscription rules are specified. All channel configurations are summarized to generate a cross-regional interaction channel configuration. This cross-regional interaction channel configuration will be read in subsequent step S103 and used to establish communication links between the central node and regional service instances.

[0054] Step S103: Based on the platform component deployment list, deploy services in the central management area according to the preset mode to generate central service instances. Configure the central service instances to access the channel according to the cross-regional interaction channel configuration to generate central nodes. Based on the platform component deployment list, deploy services in each subordinate area according to the preset mode to generate regional service instances. Establish communication links between the regional service instances and the central nodes to generate a multi-cloud, multi-regional data platform.

[0055] Based on the platform component deployment list generated in step S102 above, this embodiment extracts the component configuration parameters corresponding to the central management area. According to preset service deployment rules, the management control service and data asset management service are instantiated and deployed sequentially in the cloud resource environment of the central management area. The management control service is responsible for user management, tenant management, permission management, and resource management functions, while the data asset management service is responsible for centralized governance of cross-regional data assets and metadata aggregation functions. After the above two services are deployed, a central service instance is generated.

[0056] After the central service instance is deployed, this embodiment performs service registration and status initialization. According to preset service registration rules, the service endpoint address and service type identifier of the central service instance are written to the service registry, and the service running status is initialized to the ready state, generating a registered central service instance. The registered central service instance has the ability to be discovered and invoked by other components.

[0057] Accordingly, this embodiment loads cross-regional communication capabilities onto the registered center service instance. Synchronous call channel configuration and asynchronous message channel configuration related to the registered center service instance are extracted from the cross-regional interaction channel configuration generated in step S102. For the synchronous call channel configuration, channel ports are bound and the corresponding communication protocol is adapted; for the asynchronous message channel configuration, message queue connections are established and message publishing and subscription rules are configured. After the above configuration is applied to the registered center service instance, a central node is generated, which has the ability to communicate bidirectionally with each subordinate region.

[0058] After the central node is generated, this embodiment deploys services to each subordinate region. The component configuration parameters corresponding to each subordinate region are extracted one by one from the platform component deployment list. According to preset cluster deployment rules, the data storage and computing cluster and the unified engine gateway service are instantiated and deployed in the cloud resource environment of each subordinate region. The data storage and computing cluster provides data storage and computing capabilities for that region, and the unified engine gateway service is responsible for receiving and executing task scheduling instructions from the central node. After the above deployment is completed, a regional service instance is generated.

[0059] Based on the aforementioned regional service instances, this embodiment performs service registration and link establishment. According to preset service registration rules, the service endpoint address and regional identifier of each regional service instance are written into the service registration center, and the service running status is initialized to a ready state, generating registered regional service instances. Channel configuration parameters related to each registered regional service instance are extracted from the cross-regional interaction channel configuration, and a bidirectional channel handshake is performed with the central node. During the handshake process, channel port connectivity and communication protocol compatibility are verified; upon successful verification, an established communication link is generated.

[0060] After the established communication links are generated, this embodiment performs overall platform assembly. The central node and each registered regional service instance are associated and bound according to preset platform encapsulation rules, and the established communication links serve as data transmission and command interaction channels between nodes. After the above assembly is completed, a multi-cloud, multi-region data platform is generated. This multi-cloud, multi-region data platform supports centralized management and unified governance of data assets by the central node over its subordinate regions.

[0061] As described above, the multi-cloud, multi-region data platform construction method based on a distributed architecture provided in this application can achieve precise layout planning through adaptability analysis and role allocation. It establishes a deployment mechanism, combining pattern matching and link planning, to build a reliable platform architecture strategy. Service optimization is introduced, ensuring efficient platform operation through instantiation deployment and channel configuration. This method effectively addresses the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, providing technical support for the construction of multi-cloud, multi-region data platforms.

[0062] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0063] Step S201: Collect the regional distribution information of multiple cloud service providers and extract the geographic location identifier, available service type and resource quota parameters of each region according to the preset information parsing rules to generate a set of regional attribute information. Then, perform field mapping and format standardization processing on the set of regional attribute information according to the preset structured template to generate a standardized table of regional attributes.

[0064] Step S202: Calculate the service type matching degree and determine the resource quota sufficiency of each region in the regional attribute standardization table according to the preset regional filtering rules to generate a regional adaptability score sequence. Filter the regional adaptability score sequence according to the preset scoring threshold conditions to generate a candidate region set.

[0065] This embodiment first collects regional distribution information from multiple cloud service providers. By calling the publicly available regional query interfaces of each cloud service provider, it obtains a global list of regions and detailed regional data. For each region, three attribute fields are extracted from the returned data according to preset information parsing rules: geographic location identifier, available service type, and resource quota parameter. The geographic location identifier represents the geographic location code of the region, the available service type represents the list of cloud service types supported by the region, and the resource quota parameter represents the upper limit of computing and storage resources that can be applied for in the region. The above three attribute fields for each region are summarized and encapsulated to generate a regional attribute information set.

[0066] After the set of regional attribute information is constructed, this embodiment performs field mapping processing. Since different cloud service providers use different encoding systems for geographic location identifiers, the geographic location codes of each provider are uniformly mapped to standard geographic region codes according to a preset structured template. Simultaneously, the available service type names of each provider are mapped to preset service type enumeration values, eliminating service naming differences between different providers.

[0067] Accordingly, this embodiment performs format standardization processing on the set of regional attribute information. For resource quota parameters, the units of measurement used by different providers are uniformly converted to standard units of measurement; computing resource quotas are converted to standard computing power units; and storage resource quotas are converted to standard capacity units. After the above field mapping and format conversion are completed, the processing results are reorganized according to a preset structured template to generate a standardized regional attribute table. Each record in the standardized regional attribute table corresponds to a region and includes three fields: standardized geographic location identifier, available service type, and resource quota parameters.

[0068] Based on the standardized regional attribute table generated in step S201 above, this embodiment calculates the service type matching degree for each region. A list of required service types is extracted from the enterprise data platform construction requirements, and the available service types for each region in the standardized regional attribute table are compared item by item with the list of required service types. The number of successfully matched service types is counted and divided by the total number of required service types to obtain the service type matching degree value for that region.

[0069] After the service type matching degree is calculated, this embodiment determines the sufficiency of resource quotas for each region. The estimated resource requirements are extracted from the enterprise data platform construction needs, and the resource quota parameters for each region in the standardized regional attribute table are compared with the estimated resource requirements. When a region's computing resource quota is not lower than the estimated computing resource requirements and its storage resource quota is not lower than the estimated storage resource requirements, the region's resource quota sufficiency is marked as satisfactory; otherwise, it is marked as unsatisfactory.

[0070] Accordingly, this embodiment comprehensively scores service type matching degree and resource quota sufficiency. For regions that meet the resource quota sufficiency standard, their service type matching degree value is used as the region's suitability score; for regions that do not meet the resource quota sufficiency standard, their suitability score is set to zero. After traversing all regions in the region attribute standardization table, a region suitability score sequence is generated. The region suitability score sequence is filtered according to a preset scoring threshold condition, and regions whose suitability scores reach the threshold are retained to generate a candidate region set. The candidate region set will be read in subsequent step S301 for data compliance verification and network condition determination.

[0071] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0072] Step S301: Perform compliance attribute matching verification on each region in the candidate region set according to data compliance requirements to generate a compliance verification result set. For the regions in the compliance verification result set that pass the verification, determine the bandwidth capacity and transmission delay according to network bandwidth constraints to generate a target region configuration table.

[0073] Step S302: Calculate a comprehensive regional score sequence by weighting the network connectivity index and resource carrying capacity index of each region in the target region configuration table according to the preset center selection rules. Then, assign roles to the central management region and subordinate regions according to the preset role division threshold conditions to generate a regional role allocation scheme.

[0074] Based on the candidate region set generated in step S202 above, this embodiment performs data compliance verification on each region. A list of permitted deployment regions and data cross-border transmission restriction rules are extracted from the enterprise's data compliance requirements. The geographic location identifier of each region in the candidate region set is matched against the list of permitted deployment regions. When the geographic location identifier of a region exists in the list of permitted deployment regions, the compliance attribute matching result of that region is marked as passed. When the geographic location identifier of a region does not exist in the list of permitted deployment regions, it is further determined whether the region meets specific exemption conditions according to the data cross-border transmission restriction rules. If it meets the conditions, it is marked as passed; otherwise, it is marked as failed. After traversing all candidate regions, a compliance verification result set is generated.

[0075] After the compliance verification result set is generated, this embodiment performs network condition determination on the regions that pass the verification. Regions with compliance attribute matching results of "pass" are selected from the compliance verification result set, and bandwidth capacity is determined for each region according to network bandwidth constraints. The available bandwidth value between the region and the enterprise's core business area is obtained by calling the network detection interface, and compared with a preset bandwidth lower limit. Regions with available bandwidth not lower than the preset bandwidth lower limit pass the bandwidth capacity determination.

[0076] Accordingly, this embodiment performs transmission delay determination on regions that pass the bandwidth capacity determination. The round-trip transmission delay value between this region and the enterprise's core business area is obtained through a network probing interface and compared with a preset delay upper limit. Regions whose transmission delay does not exceed the preset delay upper limit pass the transmission delay determination. Regions that pass both the bandwidth capacity and transmission delay determinations are aggregated to generate a target region configuration table. Each record in the target region configuration table contains four fields: region identifier, geographic location identifier, available bandwidth value, and transmission delay value.

[0077] Based on the target area configuration table generated in step S301 above, this embodiment extracts network connectivity indicators for each area. For each area, the network connectivity status between it and other areas in the target area configuration table is statistically analyzed, and the number of reachable areas is divided by the total number of other areas to obtain the network connectivity indicator for that area. Simultaneously, the resource quota parameters for that area are read from the area attribute standardization table, and the calculated resource quota and storage resource quota are standardized according to a preset normalization rule, then averaged to obtain the resource carrying capacity indicator for that area.

[0078] After the network connectivity and resource carrying capacity indicators are extracted, this embodiment performs a weighted score calculation according to a preset center selection rule. For each region, its network connectivity indicator is multiplied by a first weighting coefficient, and its resource carrying capacity indicator is multiplied by a second weighting coefficient. The two are then added together to obtain the comprehensive score for that region. After traversing all regions in the target region configuration table, a comprehensive regional score sequence is generated.

[0079] Accordingly, this embodiment assigns roles to the comprehensive regional score sequence. Based on preset role classification thresholds, the region with the highest comprehensive score and reaching the central role threshold is designated as the central management region, which will assume responsibility for cross-regional control and centralized governance of data assets. The remaining regions are designated as subordinate regions, each responsible for data storage, computation, and task execution within its own region. The role assignment results for each region are associated and encapsulated with the corresponding region identifier and comprehensive score to generate a regional role allocation scheme. This regional role allocation scheme will be read in subsequent step S401 for matching platform deployment modes and vectorizing constraint parameters.

[0080] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0081] Step S401: Extract global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters from the regional role allocation scheme, and vectorize them according to the preset parameter encoding rules to generate constraint parameter vectors. Calculate the similarity between the constraint parameter vectors and the pattern feature vectors in the preset constraint model, and select the optimal mode according to the preset matching threshold conditions to generate a platform deployment mode identifier.

[0082] Step S402: Index the platform deployment mode identifier and the preset mode configuration template library to obtain the corresponding mode configuration template, and bind the data development platform deployment location parameters and data asset management platform deployment location parameters in the mode configuration template to the region according to the regional role allocation scheme to generate a component deployment configuration set.

[0083] Based on the regional role allocation scheme generated in step S302 above, this embodiment extracts the global data management requirement parameter, regional fault tolerance parameter, and total cost of ownership parameter. The global data management requirement parameter is configured and input by the enterprise, representing the enterprise's demand for centralized management capabilities across regions. Its value ranges from zero to one, with higher values ​​indicating a stronger demand for centralized management. The regional fault tolerance parameter is also configured and input by the enterprise, representing the enterprise's demand for independent operation capabilities in each region. Its value range is consistent with that of the global data management requirement parameter.

[0084] After the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters are extracted, this embodiment performs vectorization encoding according to preset parameter encoding rules. The global data management requirement parameters are used as the first component of the vector, the regional fault tolerance parameters as the second component, and the total cost of ownership parameters as the third component. These three components are then horizontally concatenated to generate a constraint parameter vector. This constraint parameter vector is a three-dimensional vector structure used to characterize the enterprise's comprehensive demand for the capabilities and characteristics of a multi-regional data platform.

[0085] Accordingly, this embodiment calculates the similarity between the constraint parameter vector and the preset constraint model. The preset constraint model stores feature vectors corresponding to three modes: distributed data platform mode feature vector, centralized data platform mode feature vector, and distributed-centralized data platform mode feature vector. The distributed data platform mode feature vector corresponds to a feature combination with low global data management requirements, high regional fault tolerance, and low overall ownership cost; the centralized data platform mode feature vector corresponds to a feature combination with low global data management requirements, high regional fault tolerance, and low overall ownership cost; and the distributed-centralized data platform mode feature vector corresponds to a feature combination with high global data management requirements, high regional fault tolerance, and high overall ownership cost. The constraint parameter vector and each mode feature vector are then used to calculate cosine similarity, generating a similarity numerical sequence.

[0086] After the similarity value sequence is generated, this embodiment selects the optimal mode according to a preset matching threshold. The mode with the highest similarity value is selected from the similarity value sequence. When this similarity value reaches the preset matching threshold, the corresponding mode is determined as the target deployment mode, and a platform deployment mode identifier is generated. The platform deployment mode identifier is an enumeration type, taking the value of one of three types: distributed, centralized, or aggregated / distributed.

[0087] Based on the platform deployment mode identifier generated in step S401, this embodiment performs an index query against a preset mode configuration template library. The preset mode configuration template library stores corresponding mode configuration templates according to the mode identifier. Each template defines the deployment location parameters for the data development platform and the data asset management platform. When the platform deployment mode identifier is distributed, the queried template specifies that both the data development platform and the data asset management platform are deployed in various regions; when the platform deployment mode identifier is centralized, the queried template specifies that both are deployed in the central management region; when the platform deployment mode identifier is distributed, the queried template specifies that the data development platform is deployed in various regions and the data asset management platform is deployed in the central management region.

[0088] After the pattern configuration template is queried, this embodiment performs region binding processing. The central management region identifier and each subordinate region identifier are read from the region role allocation scheme generated in step S302. Components whose deployment location parameters point to the center in the template are bound to the central management region identifier, and components whose deployment location parameters point to regions are bound to each subordinate region identifier. After the above binding is completed, a component deployment configuration set is generated. Each record in the component deployment configuration set contains two fields: component type and target region identifier. The component deployment configuration set will be read in subsequent step S501 for the region-by-region expansion of component instances and the generation of the platform component deployment list.

[0089] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0090] Step S501: The deployment location parameters of each component in the component deployment configuration set are instantiated and expanded according to the regional identifier in the regional role allocation scheme to generate a component instance configuration sequence. The component instance configuration sequence is then structurally assembled according to the preset list encapsulation rules to generate a platform component deployment list.

[0091] Step S502: Extract the regional affiliation and functional type of each component instance in the platform component deployment list, and perform inter-component communication requirement analysis according to the preset communication rule library to generate a communication requirement matrix. Then, perform dual-link configuration of synchronous call channel and asynchronous message channel on the communication requirement matrix according to the preset link planning rules to generate cross-regional interaction channel configuration.

[0092] Based on the component deployment configuration set generated in step S402, this embodiment instantiates and expands each component region by region. Component configuration records are read one by one from the component deployment configuration set, and for each record, two fields are extracted: component type and target region identifier. When the target region identifier points to the central management region, an instance configuration record is generated for the component, and the instance's belonging region is set to the central management region identifier. When the target region identifier points to a subordinate region, all subordinate region identifiers are read from the region role allocation scheme generated in step S302, and an instance configuration record is generated for the component in each subordinate region.

[0093] During the generation of the instance configuration record, this embodiment assigns a unique instance identifier to each record. The instance identifier is generated by concatenating the component type code and the region identifier code, and is used to uniquely identify each component instance in subsequent communication link planning. After aggregating all the instance configuration records of all components, a component instance configuration sequence is generated. Each record in the component instance configuration sequence contains three fields: instance identifier, component type, and region identifier.

[0094] Accordingly, this embodiment performs structured assembly of the component instance configuration sequence. Following preset list encapsulation rules, each record in the component instance configuration sequence is grouped and aggregated according to its region identifier, with instance configuration records from the same region grouped into the same group. After adding region-level metadata information to each group, hierarchical encapsulation is performed to generate a platform component deployment list. The platform component deployment list adopts a region-grouped structure, with the top level divided by region, and each region group containing all component instance configuration records to be deployed in that region.

[0095] Based on the platform component deployment list generated in step S501 above, this embodiment performs communication requirement analysis on each component instance. It iterates through all component instances in the platform component deployment list, extracting the region affiliation and function type attributes for each instance. According to the inter-component communication rules defined in the preset communication rule base, it is determined whether there is a communication requirement between any two component instances. The preset communication rule base defines communication requirement types based on function type combinations, including three types: no communication requirement, synchronous call requirement, and asynchronous message requirement.

[0096] In the process of determining communication requirements, this embodiment focuses on analyzing the communication requirements of cross-regional component pairs. When two component instances belong to different regions, their communication requirement type is determined based on a preset communication rule base. The communication between the management control service and the unified engine gateway service is marked as a synchronous call requirement, used for the real-time issuance of task scheduling instructions and the immediate feedback of execution status; the communication between the data asset management service and the data development platforms of each region is marked as an asynchronous message requirement, used for metadata reporting and logical model distribution. A communication requirement matrix is ​​constructed by aggregating the communication requirement types of all component instance pairs. The row and column indices of the communication requirement matrix are instance identifiers, and the matrix element values ​​represent the communication requirement type encoding between the corresponding instance pairs.

[0097] Accordingly, this embodiment performs dual-link configuration on the communication demand matrix. It iterates through instance pairs in the communication demand matrix whose communication demand type is synchronous call demand, configuring a synchronous call channel for each pair, specifying the communication protocol type and channel port parameters. It also iterates through instance pairs whose communication demand type is asynchronous message demand, configuring an asynchronous message channel for each pair, specifying the message queue name and message subscription rules. All synchronous call channel configurations and asynchronous message channel configurations are then aggregated to generate a cross-regional interaction channel configuration. This cross-regional interaction channel configuration will be read in subsequent step S601 and used for channel access and cross-regional communication capability loading for the central service instance.

[0098] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0099] Step S601: Extract the component configuration parameters corresponding to the central management area from the platform component deployment list and deploy services according to the preset mode to generate a central service instance. Register the service endpoint and initialize the service status of the central service instance according to the preset service registration rules to generate a registered central service instance.

[0100] Step S602: Extract the synchronous call channel configuration and asynchronous message channel configuration related to the registered center service instance from the cross-regional interaction channel configuration, and perform channel port binding and communication protocol adaptation to generate channel access configuration. Apply the channel access configuration to the registered center service instance to load cross-regional communication capabilities and generate a center node.

[0101] Based on the platform component deployment list generated in step S501 above, this embodiment extracts the component configuration parameters corresponding to the central management area. The group in the platform component deployment list with the area identifier "central management area" is located, and instance configuration records for the two types of components—management control service and data asset management service—are read from this group. Each record contains an instance identifier, component type, and resource specification parameters required for deployment.

[0102] After the component configuration parameters are extracted, this embodiment instantiates and deploys the service according to the preset service deployment rules.

[0103] Feasible, this embodiment includes three preset modes: centralized mode, decentralized mode, and centralized-distributed mode. In centralized mode, this embodiment can instantiate management control service, data asset management service, data development service, and unified metadata service; in decentralized mode, this embodiment can instantiate management control service; in centralized-distributed mode, this embodiment can instantiate management control service and data asset management service.

[0104] For the management and control service, computing and storage resources are requested in the cloud resource environment of the central management area, and the service container is created and started according to the resource specifications in the configuration parameters. After startup, the management and control service has four core functional modules: user management, tenant management, permission management, and resource management. For the data asset management service, the same method is used to create and start the service container. After startup, the data asset management service has three core functional modules: metadata aggregation, data asset catalog construction, and data governance rule execution. After the above two services are deployed, a central service instance is generated.

[0105] Accordingly, this embodiment performs service registration processing on the central service instance. According to preset service registration rules, the service endpoint address and service type identifier of the management and control service are written into the service registry, and the service endpoint address and service type identifier of the data asset management service are also written into the service registry. The service endpoint address includes the network address and service listening port of the host where the service resides, used for addressing and locating when other components initiate service calls.

[0106] After the service endpoint registration is completed, this embodiment performs state initialization on the central service instance. The running states of both the management control service and the data asset management service are initialized to the ready state, indicating that the services are capable of receiving external requests. The service health check probe is activated and a heartbeat reporting cycle is configured for continuous monitoring of service availability by the service registry. After the above registration and initialization are completed, a registered central service instance is generated.

[0107] Based on the cross-regional interaction channel configuration generated in step S502 above, this embodiment extracts the channel configuration related to the registered center service instance. All channel records in the cross-regional interaction channel configuration are traversed, and records where the source instance or target instance is the registered center service instance are selected. From the selection results, two categories are separated: synchronous call channel configuration and asynchronous message channel configuration. The synchronous call channel configuration includes the channel port and synchronous communication protocol parameters, while the asynchronous message channel configuration includes the message queue identifier and message subscription rule parameters.

[0108] After the channel configuration is extracted, this embodiment performs channel port binding and communication protocol adaptation. For synchronous call channel configuration, the channel port specified in the configuration is bound to the service process of the registered center service instance, and the corresponding synchronous communication protocol stack is loaded to complete protocol adaptation. For asynchronous message channel configuration, the registered center service instance is connected to the message queue specified in the configuration, and message publishing and receiving capabilities are configured according to the message subscription rule parameters. After the above configuration is completed, the channel access configuration is generated.

[0109] Accordingly, this embodiment applies the channel access configuration to the registered center service instance. The port binding result and protocol adaptation result of the synchronous call channel are loaded into the management control service, enabling it to issue synchronous call instructions to each subordinate region. The queue connection result and subscription configuration result of the asynchronous message channel are loaded into the data asset management service, enabling it to receive metadata reporting messages from each region and messages issued by the publishing logic model. After the above loading is completed, a center node is generated. This center node will be read in subsequent step S701 and used to establish communication links with each regional service instance.

[0110] In one embodiment of the multi-cloud, multi-region data platform construction method based on a distributed architecture in this application, the following may also be included:

[0111] Step S701: Extract the component configuration parameters corresponding to each subordinate region from the platform component deployment list, and instantiate and deploy the data storage computing cluster and unified engine gateway service according to the preset cluster deployment rules to generate regional service instances. Register the service endpoints and initialize the service status of the regional service instances according to the preset service registration rules to generate registered regional service instances.

[0112] Step S702: Extract the channel configuration parameters related to the registered regional service instances from the cross-regional interaction channel configuration, and perform a two-way channel handshake and link connectivity verification with the central node to generate an established communication link. Assemble the established communication link, the central node, and each of the registered regional service instances according to the preset platform encapsulation rules to generate a multi-cloud multi-regional data platform.

[0113] Based on the platform component deployment list generated in step S501 above, this embodiment extracts the component configuration parameters corresponding to each subordinate region. It iterates through all groups in the platform component deployment list whose regions are identified as subordinate regions, and reads the instance configuration records of the data storage computing cluster and unified engine gateway service components from each group. Each record contains the instance identifier, component type, belonging region identifier, and resource specification parameters required for deployment.

[0114] After the component configuration parameters are extracted, this embodiment performs region-by-region instantiation and deployment according to the preset cluster deployment rules.

[0115] Feasible, this embodiment includes three preset modes: centralized mode, decentralized mode, and centralized-distributed mode. In centralized mode, this embodiment can instantiate the unified engine gateway service and the data storage computing cluster service; in decentralized mode, this embodiment can instantiate the data asset management service, data development service, unified metadata service, unified engine gateway service, and data storage computing cluster service; in centralized-distributed mode, this embodiment can instantiate the data development service, unified metadata service, unified engine gateway service, and data storage computing cluster service.

[0116] For each subordinate region, computing and storage resources are first requested in the cloud resource environment of that region. The data storage and computing cluster is then created and initialized according to the resource specifications in the configuration parameters. This data storage and computing cluster includes distributed storage nodes and distributed computing nodes, and after startup, it possesses the capability for persistent data storage and batch computing task execution within that region.

[0117] Accordingly, this embodiment continues to deploy the unified engine gateway service within the same subordinate region. Independent computing resources are requested within the cloud resource environment of this region, and the service container is created and started according to the configuration parameters. After startup, the unified engine gateway service possesses three core functions: protocol conversion, routing selection, and task distribution. It receives task scheduling instructions from the central node and forwards them to the data storage and computing cluster in this region for execution. After traversing all subordinate regions and completing the above deployment process, a regional service instance is generated.

[0118] After the regional service instances are deployed, this embodiment performs service registration. According to preset service registration rules, the data storage computing cluster endpoint address and the unified engine gateway service endpoint address of each subordinate region are written to the service registration center, and the region identifier and service type identifier of each service instance are recorded. The running status of each service instance is initialized to the ready state, the health check probe is activated, and the heartbeat reporting cycle is configured. After the above registration and initialization are completed, a registered regional service instance is generated.

[0119] Based on the cross-regional interaction channel configuration generated in step S502 above, this embodiment extracts the channel configuration parameters related to the registered regional service instance. All channel records in the cross-regional interaction channel configuration are traversed, and records where the source instance or target instance is the registered regional service instance are selected. The channel port parameters and communication protocol parameters are then extracted from these records.

[0120] After the channel configuration parameters are extracted, this embodiment performs a bidirectional channel handshake with the central node generated in step S602. For synchronous call channels, the registered regional service instance initiates a handshake request to the central node, and the central node establishes a bidirectional synchronous call link after responding with a handshake confirmation. For asynchronous message channels, the registered regional service instance connects to the message queue subscribed to by the central node, completing the connection between the message publisher and the message receiver. After the handshake is completed, link connectivity is verified by sending probe messages to verify the reachability and latency of bidirectional data transmission. After the verification is successful, an established communication link is generated.

[0121] Accordingly, this embodiment assembles the established communication links and each node as a whole. Following preset platform encapsulation rules, the central node serves as the platform's management core, each registered regional service instance serves as a distributed execution unit, and the established communication links serve as the data transmission and instruction interaction channel between the core and the execution units. These components are associated and bound in a hierarchical structure, with the central node at the top level responsible for global management, and each regional service instance at the bottom level responsible for local execution. The communication links run through all layers to achieve cross-regional collaboration. After assembly, a multi-cloud, multi-region data platform is generated, which supports cross-cloud, cross-regional data storage and computing, and unified asset governance based on a distributed architecture.

[0122] To effectively address the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, and to provide technical support for the construction of multi-cloud, multi-region data platforms, this application provides an embodiment of a distributed architecture-based multi-cloud, multi-region data platform construction system for implementing all or part of the aforementioned distributed architecture-based multi-cloud, multi-region data platform construction method. See [link to embodiment]. Figure 2 The multi-cloud, multi-region data platform construction system based on a distributed architecture specifically includes the following components:

[0123] The region allocation module 10 is used to perform a region adaptability assessment based on the region distribution information of multiple cloud service providers according to a preset region filtering rule to generate a candidate region set, filter the candidate region set according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table, and perform a weighted score on each region in the target region configuration table according to a preset center selection rule to generate a region role allocation scheme.

[0124] The cross-regional interaction module 20 is used to generate a platform deployment mode identifier by performing pattern matching according to a preset constraint model based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme; associate and map the platform deployment mode identifier with a preset mode configuration template to obtain a component deployment configuration set; instantiate and expand the component deployment configuration set according to the regional role allocation scheme to generate a platform component deployment list; and perform communication link planning according to a preset communication rule base to generate a cross-regional interaction channel configuration.

[0125] The unified deployment module 30 is used to deploy services in the central management area according to a preset mode based on the platform component deployment list to generate central service instances, to access the central service instances through the cross-regional interaction channel configuration to generate central nodes, to deploy services in each subordinate area according to a preset mode based on the platform component deployment list to generate regional service instances, and to establish a communication link between the regional service instances and the central nodes to generate a multi-cloud, multi-regional data platform.

[0126] As described above, the multi-cloud, multi-region data platform construction system based on a distributed architecture provided in this application can achieve precise layout planning through adaptability analysis and role allocation. It establishes a reliable platform architecture strategy by building a deployment mechanism that combines pattern matching and link planning. Service optimization is introduced through instantiation deployment and channel configuration to ensure efficient platform operation. This method effectively solves the shortcomings of traditional technologies in areas such as regional planning, component deployment, and service instantiation, providing technical support for the construction of multi-cloud, multi-region data platforms.

[0127] In another embodiment of this application, a centralized multi-Region scheme suitable for heterogeneous cloud environments is also provided. Different cloud service providers have different numbers of Regions and distribution strategies, and each Region provides different cloud services and pricing strategies. Therefore, selecting appropriate Regions and cloud service providers can better meet the needs of users in different regions, while achieving high performance, high availability, and high fault tolerance. In this scheme, a multi-cloud data center is built for the enterprise through multiple Regions of different cloud service providers. One Region is selected as the central management Region. Each Region can deploy one or more data centers, which are interconnected through a dedicated network or the Internet to form a distributed system that can provide reliable computing, storage, and network services.

[0128] For different scenarios, the following interaction methods are provided between the center and the Region:

[0129] 1) Data synchronization: Synchronize user / tenant-related business data from the center to each Region. The technical implementation solutions include real-time data subscription synchronization to ensure that each Region can serve all users / tenants at the same time.

[0130] 2) Synchronous Invocation: Enables users to manage and control big data products in various Regions through the central system. Technical implementation solutions include API interface calls, such as resource allocation operations.

[0131] 3) Asynchronous calls: The corresponding messages are broadcast to each Region via message channels. Each big data product performs subsequent processing through message subscription. Technical implementation solutions include message broadcast subscription, such as logical model materialization and metadata reporting.

[0132] Understandably, a multi-regional data platform is a typical distributed system. In a distributed scenario, the goals of global data management and regional tolerance are somewhat contradictory, and system design often requires optimization in one direction. Requiring both simultaneously inevitably leads to higher software development complexity and consequently higher total cost of ownership. In conclusion, global data management, regional tolerance, and total cost of ownership are mutually exclusive; at most, only two of these three characteristics can be satisfied at the same time.

[0133] This embodiment, based on a multi-cloud centralized multi-region architecture and a multi-region data platform capability model, proposes a method and system for constructing a multi-cloud multi-region data platform. It supports multiple balance points across three aspects: global data management, regional tolerance, and total cost of ownership (TCO). Deployment strategies can be flexibly adjusted between these balance points to meet the design goals of multi-region data platforms in different user scenarios. Depending on whether data development and data asset management are deployed centrally or regionally, and combining the design concept of distributed control, three solutions are proposed and implemented: a distributed data platform, a centralized data platform, and a distributed data platform.

[0134] Specifically, distributed data platforms (RT model) include:

[0135] This solution adopts a regionalized deployment model for data development and data asset management. The entire data development, production, and data asset management chain possesses regional fault tolerance capabilities, resulting in lower architectural implementation and maintenance complexity. However, it lacks global centralized data management capabilities. The specific implementation process of this solution is as follows:

[0136] 1. Each Region independently deploys data storage and computing clusters, as well as a complete end-to-end data development and governance platform, covering all capabilities for data development and data asset management. Each region constitutes a complete data platform that can operate independently in a closed loop.

[0137] 2. The center is responsible for managing the data platform deployed in all Regions, mainly including system management operations such as user management, tenant management, permission management, and resource management. It is not responsible for any data management operations. In particular, when there are few deployed Regions and few users / tenants, the center management service may not be built for the time being.

[0138] Central-Region Interaction: In cross-region scenarios, only management and control commands are executed, without data transmission or task scheduling. This requires low network bandwidth and low network latency. For production environments, a dedicated bandwidth of >=10Mbps and a transmission latency of <=100ms are recommended.

[0139] This solution meets the design goals of regional tolerance and total cost of ownership, but it cannot guarantee the design goal of global data management, as each region manages its data independently. It is typically suitable for small and medium-sized enterprises or large corporations where the requirements for data independence and data relevance are high, necessitating decentralized management. For example, in a large corporation, each business line develops independently and builds its own data platform.

[0140] Centralized data platforms (GT model) include:

[0141] This deployment model, employing a centralized data development and data asset management approach, possesses global centralized data management capabilities and has relatively low architectural implementation and maintenance complexity. However, it lacks regional fault tolerance across the entire data development, production, and data asset management chain. The specific implementation process of this solution is as follows:

[0142] 1. Each Region independently deploys data storage and computing clusters, while also deploying a unified engine gateway service to receive and execute various types of data platform tasks from the center. Each region does not provide any platform capabilities related to data development and data asset management.

[0143] 2. The center is responsible for centralized data development and data asset management of all data storage and computing clusters deployed in all regions. The center deploys a complete end-to-end data development and governance platform, covering all capabilities of data development and data asset management, and supporting production scheduling and management of multi-cloud and multi-region data infrastructure.

[0144] Center-Region Interaction: Cross-regional scenarios only involve control command operations such as metadata access, task scheduling and status reporting, and do not involve data transmission operations, so the network bandwidth requirements are low; however, since it involves task production scheduling scenarios of the data platform, it is necessary to ensure network transmission latency and stability; the recommended configuration for the production environment is dedicated bandwidth >= 50Mbps and transmission latency <= 50ms;

[0145] This solution meets the design goals of Global Data Management and Total Cost of Ownership, but it cannot guarantee Regional Tolerance. Data production scheduling in each region depends on the central region. It is typically suitable for medium to large enterprises or group companies where the data independence requirements of each region are low, but the data correlation requirements are high, necessitating unified centralized management; and for enterprise scenarios with multiple cloud and cross-cloud data centers within the same physical region. For example, a group company may have its various business lines uniformly planning and building a data platform, with multiple regional data centers and dedicated network lines between regions; or, due to cost factors, the company may have both IDC data centers and cloud data centers in a physical region, or in cloud migration scenarios, a centralized data platform deployment model can be used to achieve task development and governance, task / data migration, etc., across multiple regions in multiple cloud environments.

[0146] Distributed data platforms (GR model) include:

[0147] This solution adopts a deployment model of regionalized data development and centralized data asset management, possessing global centralized data management capabilities and regional fault tolerance for data development and production. However, its architecture implementation and maintenance are complex, resulting in a high overall cost of ownership. The specific implementation process of this solution is as follows:

[0148] 1. Each Region independently deploys data storage and computing clusters, while also deploying a unified engine gateway service, a unified metadata service, and a data development platform. Each region forms a complete data development platform that can operate independently in a closed loop, and the data development and production process within the region has regional fault tolerance capabilities.

[0149] 2. The center deploys a data asset management platform and management console, responsible for realizing centralized data asset management capabilities for all Region data;

[0150] Center-Region Interaction: Cross-regional scenarios involve not only executing control commands such as resource management, but also data synchronization links such as user synchronization, model distribution, and metadata reporting. This places high demands on network transmission. In real-world scenarios with multiple clouds and regions in a globalized environment, it is difficult to guarantee network bandwidth and transmission latency requirements. Therefore, it is necessary to implement this through a combination of synchronous and asynchronous communication methods.

[0151] This solution meets the design goals of global data management and regional tolerance, but it cannot guarantee the design goal of total cost of ownership (TCO), resulting in a high TCO for the data platform under this model. It is typically suitable for medium-sized or large enterprises or multinational corporations where the requirements for data independence in different regions are low, but the requirements for data correlation are high, necessitating unified and centralized management. Simultaneously, there are high SLA requirements for data production stability in each region. For example, in the overseas expansion scenarios of multinational corporations, data centers need to be established in multiple regions globally. Cross-border and intercontinental network stability is extremely poor; therefore, closed-loop management within regions is necessary to ensure data production stability and to achieve unified management of global data assets across regions.

[0152] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the construction of the multi-cloud, multi-region data platform based on a distributed architecture.

[0153] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned construction of a multi-cloud, multi-region data platform based on a distributed architecture.

[0154] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the aforementioned construction of a multi-cloud, multi-region data platform based on a distributed architecture.

[0155] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0159] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a multi-cloud, multi-region data platform based on a distributed architecture, characterized in that, The method includes: Based on the regional distribution information of multiple cloud service providers, a candidate region set is generated by performing a regional adaptability assessment according to preset regional filtering rules. The candidate region set is then filtered according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table. Each region in the target region configuration table is weighted and scored according to preset center selection rules to generate a regional role allocation scheme. Based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme, a platform deployment mode identifier is generated by pattern matching according to a preset constraint model. The platform deployment mode identifier is associated and mapped with a preset mode configuration template to obtain a component deployment configuration set. The component deployment configuration set is instantiated and expanded according to the regional role allocation scheme to generate a platform component deployment list. The platform component deployment list is used to plan communication links according to a preset communication rule base to generate cross-regional interaction channel configuration. Based on the platform component deployment list, services are deployed in the central management area according to a preset mode to generate central service instances. The central service instances are then connected to the central nodes according to the cross-regional interaction channel configuration. Based on the platform component deployment list, services are deployed in each subordinate region according to a preset mode to generate regional service instances. The regional service instances and the central nodes are then connected to establish a communication link to generate a multi-cloud, multi-regional data platform.

2. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The process of generating a candidate region set by evaluating regional compatibility based on the regional distribution information of multiple cloud service providers according to preset regional filtering rules includes: The system collects regional distribution information of multiple cloud service providers and extracts the geographic location identifier, available service type and resource quota parameters of each region according to preset information parsing rules to generate a set of regional attribute information. The set of regional attribute information is then processed by field mapping and format standardization according to a preset structured template to generate a standardized table of regional attributes. For each region in the standardized regional attribute table, service type matching degree is calculated and resource quota sufficiency is determined according to preset regional filtering rules to generate a regional adaptability score sequence. The regional adaptability score sequence is then filtered according to preset scoring threshold conditions to generate a candidate region set.

3. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The process of filtering the candidate region set according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table, and then generating a region role allocation scheme by weighting each region in the target region configuration table according to a preset center selection rule, includes: For each region in the candidate region set, compliance attribute matching and verification are performed according to data compliance requirements to generate a compliance verification result set. For the regions in the compliance verification result set that pass the verification, bandwidth capacity and transmission delay are determined according to network bandwidth constraints to generate a target region configuration table. The network connectivity index and resource carrying capacity index of each region in the target region configuration table are weighted and scored according to the preset center selection rules to generate a regional comprehensive score sequence. The regional comprehensive score sequence is then used to assign roles to the central management region and subordinate regions according to the preset role division threshold conditions to generate a regional role allocation scheme.

4. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The platform deployment mode identifier is generated by pattern matching based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme according to a preset constraint model. The platform deployment mode identifier is then associated and mapped with a preset mode configuration template to obtain a component deployment configuration set, including: The global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters are extracted from the regional role allocation scheme and vectorized according to the preset parameter encoding rules to generate constraint parameter vectors. The similarity between the constraint parameter vectors and the pattern feature vectors in the preset constraint model is calculated, and the optimal mode is selected according to the preset matching threshold conditions to generate the platform deployment mode identifier. The platform deployment mode identifier is indexed and queried against the preset mode configuration template library to obtain the corresponding mode configuration template. The data development platform deployment location parameters and data asset management platform deployment location parameters in the mode configuration template are bound to the region according to the regional role allocation scheme to generate a component deployment configuration set.

5. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The step of instantiating and expanding the component deployment configuration set according to the regional role allocation scheme to generate a platform component deployment list, and then generating a cross-regional interaction channel configuration by planning communication links according to a preset communication rule base, includes: The deployment location parameters of each component in the component deployment configuration set are instantiated and expanded region by region according to the region identifier in the region role allocation scheme to generate a component instance configuration sequence. The component instance configuration sequence is then structured and assembled according to a preset list encapsulation rule to generate a platform component deployment list. The regional affiliation and functional type of each component instance in the platform component deployment list are extracted, and a communication requirement matrix is ​​generated by analyzing the communication requirements between components according to a preset communication rule library. The communication requirement matrix is ​​then configured with a dual link of synchronous call channel and asynchronous message channel according to a preset link planning rule to generate a cross-regional interaction channel configuration.

6. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The process of deploying services in the central management area based on the platform component deployment list according to a preset mode to generate central service instances, and then configuring the central service instances to access the central nodes according to the cross-regional interaction channel configuration, includes: Extract the component configuration parameters corresponding to the central management area from the platform component deployment list and deploy services according to the preset mode to generate a central service instance. Register the service endpoint and initialize the service status of the central service instance according to the preset service registration rules to generate a registered central service instance. Extract the synchronous call channel configuration and asynchronous message channel configuration related to the registered center service instance from the cross-regional interaction channel configuration, bind the channel port and adapt the communication protocol to generate the channel access configuration, and apply the channel access configuration to the registered center service instance to load cross-regional communication capabilities and generate the center node.

7. The method for constructing a multi-cloud, multi-region data platform based on a distributed architecture according to claim 1, characterized in that, The process of deploying services in each subordinate region according to a preset mode based on the platform component deployment list to generate regional service instances, and establishing communication links between the regional service instances and the central node to generate a multi-cloud, multi-region data platform, includes: Extract the component configuration parameters corresponding to each subordinate region from the platform component deployment list and deploy services to generate regional service instances according to the preset mode. Register the service endpoints and initialize the service status of the regional service instances according to the preset service registration rules to generate registered regional service instances. Extract the channel configuration parameters related to the registered regional service instances from the cross-regional interaction channel configuration, and perform bidirectional channel handshake and link connectivity verification with the central node to generate an established communication link. Then, assemble the established communication link with the central node and each of the registered regional service instances according to the preset platform encapsulation rules to generate a multi-cloud, multi-regional data platform.

8. A multi-cloud, multi-region data platform construction system based on a distributed architecture, characterized in that, The system includes: The region allocation module is used to generate a candidate region set by performing a region adaptability assessment based on the regional distribution information of multiple cloud service providers and according to preset region filtering rules. The candidate region set is then filtered according to data compliance requirements and network bandwidth constraints to obtain a target region configuration table. Each region in the target region configuration table is weighted and scored according to preset center selection rules to generate a region role allocation scheme. The cross-regional interaction module is used to generate a platform deployment mode identifier by performing pattern matching according to a preset constraint model based on the global data management requirement parameters, regional fault tolerance parameters, and total cost of ownership parameters in the regional role allocation scheme. The platform deployment mode identifier is associated and mapped with a preset mode configuration template to obtain a component deployment configuration set. The component deployment configuration set is instantiated and expanded according to the regional role allocation scheme to generate a platform component deployment list. The platform component deployment list is used to generate a cross-regional interaction channel configuration by planning communication links according to a preset communication rule base. The unified deployment module is used to deploy services in the central management area according to a preset mode based on the platform component deployment list to generate central service instances, to access the central service instances through the cross-regional interaction channel configuration to generate central nodes, to deploy services in each subordinate area according to a preset mode based on the platform component deployment list to generate regional service instances, and to establish communication links between the regional service instances and the central nodes to generate a multi-cloud, multi-regional data platform.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for constructing a multi-cloud, multi-region data platform based on a distributed architecture as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cross-region multi-data center resource cloud service modeling method and system

    CN111767139A

  • One-center multi-area deployment system

    CN113824778A

  • Instance deployment method and apparatus, cloud system, computing device, and storage medium

    US20240028415A1