Distributed database deployment planning method, device, equipment, medium and product

By generating and evaluating various regional deployment schemes, the problem of unreasonable resource allocation in distributed database networking planning was solved, the efficiency and rationality of deployment planning were improved, and a balance between performance and stability was achieved.

CN120994746APending Publication Date: 2025-11-21JINZHUAN INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511158408.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, distributed database networking planning relies on human experience, which leads to unreasonable resource allocation, insufficient disaster recovery capabilities, and complex components. The deployment scheme is not scientific or reasonable, resulting in poor deployment planning efficiency and poor performance and stability.

Method used

By acquiring deployment requirement data from demanders, basic deployment schemes with various regional deployment methods are generated, such as single-center, dual-center, and two-site three-center schemes. The schemes are evaluated based on pre-set evaluation dimensions and weights to generate a score, and finally the target deployment scheme is determined.

Benefits of technology

It improves the efficiency and rationality of distributed database deployment planning, ensures a better balance between performance and stability in deployment schemes, and provides diversity and redundancy to adapt to different needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994746A_ABST
    Figure CN120994746A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed database deployment planning method and device, equipment, a medium and a product. The method comprises the following steps: acquiring deployment demand data provided by a demand side; according to the deployment demand data, generating at least three basic deployment schemes with different regional deployment modes; wherein the regional deployment mode comprises a single-center mode, a double-center mode and a two-place three-center mode; according to a preset evaluation dimension and a corresponding dimension weight, evaluating each basic deployment scheme to obtain a scheme score corresponding to each basic deployment scheme; and determining a target deployment scheme of the demander according to the scheme score. According to the technical scheme of the embodiment of the invention, the deployment planning efficiency can be improved; and the deployment scheme is selected on the basis of evaluation of different dimensions, so that the reasonability of the deployment scheme can be improved by referring to facts as much as possible, and the performance and the stability of the deployment scheme further tend to be more balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed database technology, and in particular to a distributed database deployment planning method, apparatus, device, medium and product. Background Technology

[0002] With the continuous development of information technology application innovation, database technology has been widely promoted and advanced. Among them, distributed databases, with their high availability and flexibility, are widely used in various industries to assist in production and daily life in industries related to the Internet and the Internet of Things. However, the ability to quickly and efficiently propose distributed database networking deployment solutions for a specific business scenario remains one of the key research focuses for technical personnel in related fields.

[0003] Currently, the networking planning of distributed databases mainly relies on human experience, which can easily lead to problems such as unreasonable resource allocation, insufficient disaster recovery capabilities, and complex components. This results in unscientific and unreasonable deployment schemes, poor deployment planning efficiency, and poor performance and stability of the deployment planning results. Summary of the Invention

[0004] This application provides a method, apparatus, device, medium, and product for planning distributed database deployment, so as to improve the efficiency and rationality of distributed database network deployment planning and enable the deployment scheme to achieve a better balance between performance and stability.

[0005] According to one aspect of this application, a distributed database deployment planning method is provided, comprising:

[0006] Obtain deployment requirement data provided by the requesting party;

[0007] Based on deployment requirement data, generate at least three basic deployment schemes with different regional deployment methods; among them, the regional deployment methods include single-center, dual-center, and two-site three-center.

[0008] Based on the pre-defined evaluation dimensions and corresponding dimension weights, each basic deployment scheme is evaluated to obtain a scheme score for each basic deployment scheme.

[0009] Based on the evaluation of the proposed solutions, the target deployment plan for the client is determined.

[0010] According to another aspect of this application, a distributed database deployment planning apparatus is provided, comprising:

[0011] The requirement data acquisition module is used to acquire deployment requirement data provided by the requesting party;

[0012] The basic solution generation module is used to generate at least three basic deployment solutions with different regional deployment methods based on deployment requirement data; among them, the regional deployment methods include single-center, dual-center, and two-site three-center.

[0013] The solution scoring module is used to evaluate each basic deployment solution based on pre-set evaluation dimensions and corresponding dimension weights, and obtain the solution score corresponding to each basic deployment solution.

[0014] A target solution determination module is used to determine the target deployment solution for the requester based on solution scoring. According to another aspect of this application, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the distributed database deployment planning method described in any embodiment of this application.

[0018] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the distributed database deployment planning method described in any embodiment of this application.

[0019] According to another aspect of this application, a computer program product is provided, the computer program product including a computer program that, when executed by a processor, implements the distributed database deployment planning method according to any embodiment of this application.

[0020] In the technical solution of this application embodiment, based on the deployment requirement data provided by the demander, multiple basic deployment schemes are generated for different regional deployment methods. This provides the demander with multiple schemes for different regional deployment methods for subsequent screening and selection, improving the diversity and redundancy of solving deployment requirements and enhancing the efficiency of deployment planning. These basic deployment schemes are scored according to different evaluation dimensions and dimensional weights. Based on the scheme scores, the target deployment scheme for the demander to deploy the distributed database is determined. Selecting a deployment scheme based on evaluations of different dimensions can improve the rationality of the deployment scheme as much as possible by referring to facts, thereby further enabling the deployment scheme to be more balanced in terms of performance and stability.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a distributed database deployment planning method according to Embodiment 1 of this application;

[0024] Figure 2 This is a flowchart of a distributed database deployment planning method according to Embodiment 2 of this application;

[0025] Figure 3 This is a schematic diagram of a distributed database deployment planning device according to Embodiment 3 of this application;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the distributed database deployment planning method of the embodiments of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This application provides a flowchart of a distributed database deployment planning method according to Embodiment 1. This embodiment is applicable to the deployment planning of distributed databases. The method can be executed by a distributed database deployment planning device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0031] S110. Obtain deployment requirement data provided by the requester.

[0032] The demand party can be any organization, institution, or individual that needs to deploy a distributed database network; that is, the demand party has a need for the hardware and software required to deploy a distributed database. Accordingly, the deployment requirement data can be the data that the distributed database required by the demand party can fulfill its actual production needs, such as the amount of data the demand party needs to handle, the required network bandwidth, etc. This application embodiment will not exhaustively list these requirements. It is understood that the demand party's deployment requirement data can be directly provided by the demand party, or the application's requirement information can be directly collected based on the status of the database currently used by the demand party.

[0033] For example, if the client is not currently using a distributed database and wants to upgrade to a completely new distributed database, then hardware and network data can be collected from the current source database. This could include data collection such as hardware monitoring data, SQL (Structured Query Language) execution data, and data distribution data.

[0034] S120. Based on deployment requirement data, generate at least three basic deployment schemes with different regional deployment methods; among which, the regional deployment methods include single-center, dual-center, and two-site three-center.

[0035] Geographic deployment can refer to the deployment format of distributed database hardware. As the name suggests, the hardware responsible for different databases within a distributed database can be distributed across different regions. This can include three main forms: single-center, dual-center, and two-site three-center. A single-center deployment can be understood as "one location, one center," meaning all distributed database hardware is deployed in the same data center within the same region. A dual-center deployment can be understood as "one location, two centers," meaning the hardware responsible for different databases within a distributed database is deployed in two different data centers within the same region. A two-site three-center deployment involves the hardware responsible for different databases within a distributed database being deployed in three different data centers across two different regions. It's understood that the more geographically dispersed the hardware deployment and the more data centers involved, the higher the redundancy and the greater the security. Even if an emergency occurs in any one region or data center and is forced to shut down, it will not affect the continued operation of other regions or data centers.

[0036] A basic deployment plan can be a preliminary deployment plan proposed based on the deployment requirements data of the demand side. It is understood that these preliminary plans can be automatically generated directly according to pre-defined rules. For example, based on three different regional deployment methods, different hardware information can be configured according to the demand side's deployment requirements data to form a basic deployment plan. Alternatively, a machine learning model can be trained based on the existing deployment status of different distributed databases. For example, a trained model can generate basic deployment plans, taking the demand side's deployment requirements data and three different regional deployment methods as input, and outputting basic deployment plans for at least three different regional deployment methods. This application does not limit the type of machine learning model or the training and usage process.

[0037] S130. Based on the pre-set evaluation dimensions and corresponding dimension weights, evaluate each basic deployment scheme to obtain the scheme score corresponding to each basic deployment scheme.

[0038] The evaluation dimensions can be different perspectives used to evaluate a basic deployment solution, such as, but not limited to, network performance, hardware configuration, and cost, etc., which are not exhaustively listed here. Dimension weights represent the relative importance of different evaluation dimensions in evaluating a basic deployment solution; the more important the evaluation dimension, the heavier its corresponding weight. Of course, dimension weights can be set by those skilled in the art based on extensive experimentation, specific circumstances, or human experience; this embodiment does not limit this.

[0039] The solution score can be an evaluation score corresponding to each basic deployment solution. Based on the evaluation dimensions, scores are given from different angles, and then weighted and calculated according to the dimension weights to obtain the solution score for each basic deployment solution.

[0040] S140. Based on the scheme scoring, determine the target deployment scheme of the requester.

[0041] After obtaining the respective scheme scores for different basic deployment schemes in the aforementioned steps, the schemes can be directly screened. For example, the basic deployment scheme with the highest scheme score can be selected as the target deployment scheme. Alternatively, some schemes can be eliminated from these basic schemes based on their scheme scores, and then the remaining schemes can be improved to finally determine one scheme as the target deployment scheme. There are various ways to improve the schemes. For example, the deployment schemes can be adapted to the specific industry of the client to better meet the client's interests. This application does not limit the specific methods described herein.

[0042] In the technical solution of this application embodiment, based on the deployment requirement data provided by the demander, multiple basic deployment schemes are generated for different regional deployment methods. This provides the demander with multiple schemes for different regional deployment methods for subsequent screening and selection, improving the diversity and redundancy of solving deployment requirements and enhancing the efficiency of deployment planning. These basic deployment schemes are scored according to different evaluation dimensions and dimensional weights. Based on the scheme scores, the target deployment scheme for the demander to deploy the distributed database is determined. Selecting a deployment scheme based on evaluations of different dimensions can improve the rationality of the deployment scheme as much as possible by referring to facts, thereby further enabling the deployment scheme to be more balanced in terms of performance and stability.

[0043] Example 2

[0044] Figure 2 This is a flowchart illustrating a distributed database deployment planning method provided in Embodiment 2 of this application. This embodiment further refines the scoring operation for each basic deployment scheme based on the aforementioned embodiments. Figure 2 As shown, the method includes:

[0045] S210. Obtain deployment requirement data provided by the requester.

[0046] S220. Based on deployment requirement data, generate at least three basic deployment schemes with different regional deployment methods; among which, the regional deployment methods include single-center, dual-center, and two-site three-center.

[0047] S230. Evaluate each evaluation dimension according to the preset evaluation indicators corresponding to each evaluation dimension, and obtain the dimension score; among which, the evaluation dimensions include network performance, hardware resources, data distribution, disaster recovery capability, component planning and cost-effectiveness.

[0048] The evaluation indicators can be different indicators included in the evaluation dimension. An evaluation dimension may include more than one evaluation indicator, and the comprehensive evaluation result of these multiple evaluation indicators can be reflected through the dimension score of the evaluation dimension. That is, the dimension score is the score of the evaluation dimension, and it is also the comprehensive evaluation result of different evaluation indicators under the evaluation dimension.

[0049] Network performance can be an evaluation dimension of the network capabilities that a distributed database can provide; hardware resources can be an evaluation dimension of the hardware support that a distributed database can provide; data distribution can be an evaluation dimension of the sharding and cross-node query capabilities that a distributed database can provide; disaster recovery capability can be an evaluation dimension of the data security capabilities that a distributed database can provide; component planning can be an evaluation dimension of the hardware coordination capabilities that a distributed database can provide; and cost-effectiveness can be an evaluation dimension of the expenses incurred during the deployment of a distributed database.

[0050] S240. The scores of each dimension are weighted and calculated according to the weight of each dimension to obtain the solution score.

[0051] Understandably, different evaluation indicators can be scored separately, and the scores of each indicator can be aggregated as the dimensional score of the evaluation dimension. Then, a weighted sum is calculated according to the preset dimensional weights to obtain the scheme score of the basic deployment scheme.

[0052] S250. Based on the scheme scoring, determine the target deployment scheme of the requester.

[0053] In the technical solutions of the above embodiments, any basic deployment scheme is evaluated from different evaluation dimensions. The evaluation is conducted from six different evaluation dimensions, including network performance, hardware resources, data distribution, disaster recovery capability, component planning, and cost-effectiveness. This can reflect the performance or cost-effectiveness of the basic deployment scheme from different perspectives, thereby helping to screen and select a suitable deployment scheme in a more practical way, and helping to improve the rationality of distributed database deployment planning.

[0054] In one optional implementation, the step S230 of evaluating each evaluation dimension according to preset evaluation indicators corresponding to each evaluation dimension to obtain a dimension score may include:

[0055] S231. For any evaluation dimension, score each evaluation indicator according to the preset scoring rules corresponding to each evaluation indicator to obtain the indicator score.

[0056] The preset scoring rules can be pre-defined association rules between the data of evaluation indicators and their corresponding evaluation scores. For example, network performance includes two evaluation indicators: network latency and bandwidth utilization. Pre-defined criteria might assign 100 points to network latency less than or equal to 5ms, 80 points to 5-10ms, and 50 points to latency greater than 10ms. Therefore, simulation software can be used to simulate the hardware selection, deployment, and execution of the distributed database in the basic deployment scheme, thereby obtaining the data corresponding to each evaluation indicator. Based on the preset scoring rules and the data obtained from the simulation, the score corresponding to each evaluation indicator is determined, i.e., the indicator score.

[0057] S232. The scores of each indicator are weighted and calculated according to the preset indicator weights corresponding to each evaluation indicator to obtain the dimension score.

[0058] Similar to evaluation dimensions and dimension weights, each evaluation indicator has its corresponding indicator weight. This indicator weight can be preset by relevant technical personnel based on extensive experiments, actual conditions, or human experience; this embodiment does not limit this. The weighted sum of the indicator scores determined based on preset scoring rules in the aforementioned steps, combined with the indicator weights, is used as the dimension score of the evaluation dimension for that indicator.

[0059] For example, continuing from the previous example, network performance includes network latency and bandwidth utilization. According to the preset scoring rules, the score for network latency in the current basic deployment plan is 100 points, and the score for bandwidth utilization is 70 points. The preset weight of each indicator is 0.5. Therefore, the dimension score for network performance is 100×0.5+70×0.5=85 points.

[0060] Furthermore, the evaluation metrics for network performance in the aforementioned evaluation dimensions include latency and bandwidth utilization; the evaluation metrics for hardware resources include peak processor utilization and memory read / write speeds per second; the evaluation metrics for data distribution include sharding balance and cross-node query ratio; the evaluation metrics for disaster recovery capabilities include allowable recovery time, data redundancy, and geographical deployment methods; the evaluation metrics for component planning include resource matching degree, component collaboration efficiency, and elasticity; and the evaluation metrics for cost-effectiveness include hardware procurement costs and operation and maintenance energy consumption costs.

[0061] Among them, the latency metric can be the degree of network latency;

[0062] Bandwidth utilization can be the ratio of actual network bandwidth usage to the theoretical maximum bandwidth.

[0063] Peak processor utilization can be the highest percentage of CPU (Central Processing Unit) load during peak hours;

[0064] The read / write metric for memory can be the number of read / write operations per second (IOPS) or the amount of data per second (MB / s);

[0065] Shard balance can be defined as the evenness of the distribution of data shards among cluster nodes (e.g., standard deviation ≤ 10%).

[0066] Cross-node query percentage can be defined as the proportion of query requests that require collaboration across multiple nodes out of the total query volume.

[0067] The allowed recovery time objective (RTO) can be the maximum tolerable recovery time of a system after a failure.

[0068] Data redundancy can be measured by the number of data replicas.

[0069] The regional deployment method has been introduced earlier and will not be repeated here;

[0070] Resource matching can be the deviation between actual CPU utilization and expected CPU utilization;

[0071] Component collaboration efficiency can refer to the collaboration performance between microservices or middleware, such as query latency;

[0072] Flexibility can be the degree of performance increase after a component is scaled up;

[0073] Hardware procurement costs can be the actual investment in fixed assets such as servers and network equipment used to build a distributed database;

[0074] Operation and maintenance energy costs can include ongoing operating expenses such as electricity and cooling.

[0075] Each of the above evaluation indicators can be assigned a different score value through preset scoring rules, and each indicator has a different weight, which is used to calculate the score for its respective dimension.

[0076] In the two implementation methods described above, specific evaluation indicators are provided for scoring different dimensions. These evaluation indicators show the simulation performance of the basic deployment scheme in different dimensions, providing basic support for determining the target deployment scheme in the future.

[0077] In another optional implementation, the step of S250, determining the target deployment plan of the demander based on the plan score, may include:

[0078] S251. Based on the scheme scoring, screen each basic deployment scheme and determine at least two candidate deployment schemes.

[0079] The foregoing embodiments and implementation schemes described the process of determining the scheme score. Since there are at least three basic deployment schemes, the scheme scores for different basic deployment schemes are generally different. Candidate deployment schemes are selected based on the scheme scores. Of course, at least two basic deployment schemes with the highest scheme scores can be directly selected as candidates, or a scoring threshold can be preset, and basic deployment schemes exceeding the scoring threshold can be considered as candidates.

[0080] S252. Obtain business scenario information provided by the demand side.

[0081] The business scenario information can be the business scenario in which the demander uses the distributed database, such as, but not limited to, the financial, telecommunications, and internet industries. This business scenario information can be obtained based on the industry of the demander.

[0082] S253. Based on business scenario information, or by using a pre-trained reinforcement learning model, adjust the dimensional weights to determine the adjustment score for each candidate deployment scheme.

[0083] The different business scenarios can influence the weights of different evaluation dimensions. It's understandable that the six different evaluation dimensions—network performance, hardware resources, data distribution, disaster recovery capabilities, component planning, and cost-effectiveness—have varying degrees of importance for different business scenarios. Rules for adjusting dimension weights corresponding to different business scenarios can be pre-defined. Alternatively, a reinforcement learning model can be pre-trained, utilizing its state space, action space, and reward function to adjust the dimension weights. For example, the state space could set scores for each evaluation temperature in candidate deployment schemes, the action space could set weights for the evaluation dimensions, and the reward function could represent the degree of optimization of the deployment scheme, which could be compared between the degree of performance improvement and the degree of cost increase. After adjusting the dimension weights, the weighted sum is recalculated according to the adjusted weights, and the result is used as the adjusted score.

[0084] It is important to emphasize that when the dimensional weights change, the hardware and software corresponding to each evaluation dimension in the deployment plan need to be adjusted to accommodate the weight changes, thereby actually altering the deployment plan and testing it through pre-configured topology simulation software.

[0085] S254. Based on the adjusted score, determine the target deployment plan.

[0086] Based on the adjustment score selection of the target deployment scheme determined in the aforementioned steps, preferably, the candidate deployment scheme with the highest adjustment score is selected as the target deployment scheme.

[0087] In the above implementation, on the one hand, by appropriately adjusting the dimension weights based on different business scenarios, it is possible to effectively optimize the corresponding adaptive optimization according to the needs of the business scenarios, which is more in line with the application scenarios of the demanders and improves the rationality of deployment planning; on the other hand, adjusting the dimension weights through reinforcement learning models can improve the efficiency of deployment planning.

[0088] In a further optional implementation, adjusting the dimension weights based on business scenario information as described in S253 may include:

[0089] S2531. In response to business scenario information for financial industry application scenarios, the weight of the dimension corresponding to disaster recovery capability increases, while the weight of the dimension corresponding to cost-effectiveness decreases.

[0090] Understandably, the financial industry has extremely high requirements for data security. For example, in sectors like banking, ensuring data security is worth the extra cost. Therefore, it's appropriate to increase the weight of disaster recovery capabilities to guarantee data security, while simultaneously reducing the weight of cost-effectiveness-related dimensions.

[0091] S2532. In response to business scenario information for telecommunications industry application scenarios, the weight of the dimension corresponding to disaster recovery capability is increased, and the weight of the dimension corresponding to cost-effectiveness is increased.

[0092] Understandably, the electronics and information industry also has very high requirements for data security, but at the same time, it hopes to reduce costs and increase efficiency. Therefore, it is appropriate to increase the weight of disaster recovery capabilities and cost-effectiveness.

[0093] S2533. In response to business scenario information for Internet industry application scenarios, the weight of the dimension corresponding to network performance is increased, and the weight of the dimension corresponding to component planning is decreased.

[0094] Understandably, the internet industry has high requirements for network speed and stability, but its demands on hardware processing and collaboration performance are not as high as in other industries. Therefore, it is appropriate to increase the weight of the dimension corresponding to network performance, while appropriately decreasing the weight of the dimension corresponding to component planning.

[0095] The above implementation provides a practical and effective scheme for adjusting the dimension weights of information for different business scenarios. It can quickly adjust the dimension weights in candidate deployment schemes and then recalculate the adjustment score, which helps to improve the efficiency of determining the target deployment scheme.

[0096] Based on the foregoing embodiments and implementation methods, this application presents a practical example as follows:

[0097] Distributed databases generally involve four types of nodes: Manager Node, Global Transaction Node (GTMNode), Computer Node, and Data Node.

[0098] Management node, used for cluster metadata and high availability control.

[0099] The Global Transaction Manager (GTM) node is used to maintain the lifecycle of global transactions and global sequences, ensuring the consistency and reliability of distributed transactions.

[0100] Compute nodes are used to support horizontal scaling in a stateless environment, are responsible for SQL logic optimization and physical optimization, and generate distributed query plans that satisfy distributed transaction consistency.

[0101] Data nodes are the final storage modules for application data. They are deployed with primary and backup nodes separated. They are responsible for receiving SQL operations from computing nodes, performing logical and physical optimizations, generating and executing the optimal query plan, and efficiently completing data read and write tasks.

[0102] In the specific distributed database deployment planning process, the deployment requirement data is first collected from the source database of the requester. This data may include hardware monitoring data (CPU, memory, IOPS), SQL execution characteristics (high-frequency queries, transaction patterns), and data distribution (table size, relationships).

[0103] Hardware monitoring data collection may include, but is not limited to: CPU utilization (peak / average values ​​collected at the core level), memory usage (including Swap usage), storage IOPS (distinguishing between read and write throughput), and network bandwidth (cross-node traffic statistics).

[0104] SQL execution characteristic analysis may include, but is not limited to: high-frequency query statements (TOP 50 SQL) and their execution plans, transaction patterns (short transaction / long transaction ratio) and lock contention hotspot tables (such as V$LOCK view), etc.

[0105] Data distribution statistics can include, but are not limited to: table size and growth trend (distinguishing between hot and cold data) and the intensity of table join queries (such as JOIN frequency statistics).

[0106] After obtaining the deployment requirements data collected above, the data is preprocessed, including data cleaning and normalization.

[0107] Data cleaning can include outlier removal, such as discarding data with network latency exceeding 1000ms as invalid data; it can also include filling missing values, such as imputing using the mean or nearest neighbor values. Normalization unifies data of different dimensions to the same order of magnitude, facilitating computation.

[0108] Based on these deployment requirement data, and according to different regional deployment methods—single-center, dual-center, and two-site three-center—three different basic deployment schemes are provided to the requesting party. These schemes can be generated directly from preset rules or provided using pre-trained machine learning models; this embodiment of the application does not impose any limitations on this.

[0109] The next step is to evaluate these basic deployment solutions, covering six dimensions: network performance, hardware resources, data distribution, disaster recovery capabilities, component planning, and cost-effectiveness. Each dimension will be scored using evaluation metrics, and the scores will be weighted and calculated based on pre-set dimension weights. The final score will be the overall score for each basic deployment solution.

[0110] For example, the initial preset weights for each evaluation dimension are as follows:

[0111] Network performance 20%, hardware resources 15%, data distribution 25%, disaster recovery capability 20%, drunk driving case planning 10%, cost-effectiveness 10%.

[0112] Then, using pre-configured simulation software, the above-generated at least three basic deployment schemes are simulated to obtain the evaluation indicators for the above six evaluation dimensions in the simulation results, and these indicators are scored according to the preset scoring rules.

[0113] The preset scoring rules pre-define the corresponding index scores within the index range of each indicator. Each evaluation dimension includes more than one evaluation indicator, and each indicator has a preset index weight. The index scores and index weights are weighted and calculated to obtain the dimension score. Then, the scores of each dimension and their corresponding dimension weights are weighted and calculated to obtain the basic deployment scheme score.

[0114] The evaluation indicators included in each evaluation dimension, and the preset scoring rules corresponding to each evaluation indicator are shown in Table 1.

[0115] Table 1

[0116]

[0117]

[0118] Based on the above multi-dimensional evaluation and analysis, the reinforcement learning model optimizes the sharding strategy based on historical data, such as distributing hot data for storage, reducing cross-node queries, matching similar business scenarios with the knowledge base, simulating the operating pressure of various network solutions, and conducting various fault switching drills. It quantifies and displays the actual cluster operating parameters and fault switching, and finally recommends the optimal target deployment solution.

[0119] On the one hand, the weights of dimensions are optimized by combining business scenario information. For example, when the business scenario is the financial industry, the weight of disaster recovery capability is increased to 30%, and the weight of cost-effectiveness is reduced to 5%; when the business scenario is the telecommunications industry, the weight of disaster recovery capability is increased to 30%, and the weight of cost-effectiveness is increased to 20%; when the business scenario is the internet industry, the weight of network performance is increased to 25%, and the weight of component planning is reduced to 10%, etc. Other evaluation dimensions can be adjusted appropriately according to specific circumstances.

[0120] On the other hand, the dimensional weights of the evaluation dimensions are further adjusted using a pre-trained reinforcement learning model. For example, scores for each evaluation temperature in the candidate deployment scheme are set in the state space, weights for the evaluation dimensions are set in the action space, and the reward function represents the degree of optimization of the deployment scheme, which can be compared between the degree of performance improvement and the degree of cost increase. Adjusting the weights means that the selection and deployment of hardware and software in the deployment scheme are changed to adapt to the change in weights. After adjusting the dimensional weights, the weighted sum is recalculated according to the adjusted dimensional weights, and the result is used as the adjusted score.

[0121] Then, the deployment scheme with adjusted weights is re-simulated using pre-deployed simulation software to generate a visual evaluation report, such as a 3D topology map, using color depth to indicate the load status of nodes (e.g., red represents nodes under high load, green represents nodes under balanced load, etc.). The deployment scheme with the highest score after weight adjustment is recommended to the client.

[0122] Example 3

[0123] Figure 3 This is a schematic diagram of a distributed database deployment planning device provided in Embodiment 3 of this application. Figure 3 As shown, the device 300 includes:

[0124] The requirement data acquisition module 310 is used to acquire deployment requirement data provided by the requester.

[0125] The basic solution generation module 320 is used to generate at least three basic deployment solutions with different regional deployment methods based on deployment requirement data; among them, the regional deployment methods include single-center, dual-center, and two-site three-center.

[0126] The solution scoring module 330 is used to evaluate each basic deployment solution according to the pre-set evaluation dimensions and corresponding dimension weights, and obtain the solution score corresponding to each basic deployment solution.

[0127] The target solution determination module 340 is used to determine the target deployment solution for the requester based on the solution scoring.

[0128] In the technical solution of this application embodiment, based on the deployment requirement data provided by the demander, multiple basic deployment schemes are generated for different regional deployment methods. This provides the demander with multiple schemes for different regional deployment methods for subsequent screening and selection, improving the diversity and redundancy of solving deployment requirements and enhancing the efficiency of deployment planning. These basic deployment schemes are scored according to different evaluation dimensions and dimensional weights. Based on the scheme scores, the target deployment scheme for the demander to deploy the distributed database is determined. Selecting a deployment scheme based on evaluations of different dimensions can improve the rationality of the deployment scheme as much as possible by referring to facts, thereby further enabling the deployment scheme to be more balanced in terms of performance and stability.

[0129] In one optional implementation, the scheme scoring and determination module 330 may include:

[0130] The dimension scoring determination unit is used to evaluate each evaluation dimension according to the preset evaluation indicators corresponding to each evaluation dimension, and obtain the dimension score; among which, the evaluation dimensions include network performance, hardware resources, data distribution, disaster recovery capability, component planning and cost-effectiveness.

[0131] The scheme scoring unit is used to weight and calculate the scores of each dimension according to the weight of each dimension to obtain the scheme score.

[0132] In one optional implementation, the dimension scoring determination unit may include:

[0133] The indicator scoring subunit is used to score each evaluation indicator for any evaluation dimension according to the preset scoring rules corresponding to each evaluation indicator, and obtain the indicator score.

[0134] The dimensional scoring subunit is used to weight and calculate the scores of each indicator according to the preset indicator weights corresponding to each evaluation indicator, so as to obtain the dimensional score.

[0135] In one alternative implementation, the evaluation metrics for network performance include latency metrics and bandwidth utilization.

[0136] Evaluation metrics for hardware resources include processor peak utilization and memory read / write speeds per second.

[0137] Evaluation metrics for data distribution include shard balance and cross-node query ratio;

[0138] The evaluation indicators for disaster recovery capabilities include allowable recovery time, data redundancy, and geographical deployment method;

[0139] The evaluation indicators for component planning include resource matching degree, component collaboration efficiency, and elasticity capability;

[0140] Cost-benefit analysis metrics include hardware procurement costs and operation and maintenance energy consumption costs.

[0141] In one alternative implementation, the target scheme determination module 340 may include:

[0142] The candidate solution determination unit is used to screen each basic deployment solution based on the solution score and determine at least two candidate deployment solutions.

[0143] The business scenario acquisition unit is used to acquire business scenario information provided by the demand side.

[0144] The adjustment scoring unit is used to adjust the dimension weights based on business scenario information or by using a pre-trained reinforcement learning model, and to determine the adjustment score for each candidate deployment scheme.

[0145] The target deployment plan determination unit is used to determine the target deployment plan based on the adjusted score.

[0146] In one optional implementation, the adjustment score determination unit may include:

[0147] The financial scenario processing subunit is used to respond to business scenario information as financial industry application scenarios, control the increase of the dimension weight corresponding to disaster recovery capability, and control the decrease of the dimension weight corresponding to cost-effectiveness.

[0148] The telecommunications scenario processing subunit is used to respond to business scenario information as telecommunications industry application scenarios, control the increase of the dimension weight corresponding to disaster recovery capability, and control the increase of the dimension weight corresponding to cost-effectiveness.

[0149] The Internet scenario processing subunit is used to respond to business scenario information as Internet industry application scenarios, control the increase of the dimension weight corresponding to network performance, and control the decrease of the dimension weight corresponding to component planning.

[0150] The distributed database deployment planning apparatus provided in this application embodiment can execute the distributed database deployment planning method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing each distributed database deployment planning method.

[0151] Example 4

[0152] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0153] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0154] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0155] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as distributed database deployment planning methods.

[0156] In some embodiments, the distributed database deployment planning method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the distributed database deployment planning method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the distributed database deployment planning method by any other suitable means (e.g., by means of firmware).

[0157] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0162] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0163] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the distributed database deployment planning method provided in any embodiment of this application. This program product shares the same inventive concept as the distributed database deployment planning method disclosed in the embodiments of this application, and therefore will not be described in detail here.

[0164] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.

[0165] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A distributed database deployment planning method, characterized in that, include: Obtain deployment requirement data provided by the requesting party; Based on the deployment requirement data, at least three different regional deployment schemes are generated; wherein, the regional deployment schemes include single-center, dual-center, and two-site three-center schemes; Based on the pre-set evaluation dimensions and corresponding dimension weights, each of the basic deployment schemes is evaluated to obtain a scheme score for each of the basic deployment schemes. Based on the scoring of the proposed solutions, the target deployment plan for the requesting party is determined.

2. The method according to claim 1, characterized in that, The step of evaluating each basic deployment scheme according to pre-set evaluation dimensions and corresponding dimension weights to obtain a scheme score for each basic deployment scheme includes: Each evaluation dimension is evaluated based on the preset evaluation indicators corresponding to each evaluation dimension, and a dimension score is obtained; wherein, the evaluation dimensions include network performance, hardware resources, data distribution, disaster recovery capability, component planning and cost-effectiveness; The scores for each dimension are weighted and calculated according to the weights of each dimension to obtain the score for the scheme.

3. The method according to claim 2, characterized in that, The step of evaluating each evaluation dimension according to the preset evaluation indicators corresponding to each evaluation dimension to obtain a dimension score includes: For any of the evaluation dimensions, each evaluation indicator is scored according to the preset scoring rules corresponding to each evaluation indicator to obtain the indicator score; The scores of each indicator are weighted and calculated according to the preset indicator weights corresponding to each evaluation indicator to obtain the dimension score.

4. The method according to claim 3, characterized in that, The evaluation metrics for network performance include latency and bandwidth utilization. The evaluation metrics for the hardware resources include processor peak utilization and memory read / write per second. The evaluation metrics corresponding to the data distribution include shard balance and cross-node query ratio. The evaluation indicators corresponding to the disaster recovery capability include the allowable recovery time, data redundancy, and the geographical deployment method; The evaluation indicators corresponding to the component planning include resource matching degree, component collaboration efficiency, and elasticity capability; The cost-benefit evaluation indicators include hardware procurement costs and operation and maintenance energy consumption costs.

5. The method according to claim 2, characterized in that, The step of determining the target deployment plan for the requester based on the plan score includes: Based on the scoring of the proposed solutions, each of the basic deployment solutions is screened to determine at least two candidate deployment solutions. Obtain the business scenario information provided by the requesting party; Based on the business scenario information, or by using a pre-trained reinforcement learning model, the dimensional weights are adjusted to determine the adjustment score for each of the candidate deployment schemes; Based on the adjusted score, the target deployment plan is determined.

6. The method according to claim 5, characterized in that, The step of adjusting the dimension weights based on the business scenario information includes: In response to the fact that the business scenario information is a financial industry application scenario, the weight of the dimension corresponding to the disaster recovery capability is increased, and the weight of the dimension corresponding to the cost-effectiveness is decreased. In response to the fact that the business scenario information is a telecommunications industry application scenario, the weight of the dimension corresponding to the disaster recovery capability is increased, and the weight of the dimension corresponding to the cost-effectiveness is increased. In response to the business scenario information being an internet industry application scenario, the weight of the dimension corresponding to network performance is increased, and the weight of the dimension corresponding to component planning is decreased.

7. A distributed database deployment planning device, characterized in that, include: The requirement data acquisition module is used to acquire deployment requirement data provided by the requesting party; The basic solution generation module is used to generate at least three basic deployment solutions with different regional deployment methods based on the deployment requirement data; wherein, the regional deployment methods include single-center, dual-center, and two-site three-center; The scheme scoring determination module is used to evaluate each of the basic deployment schemes according to the pre-set evaluation dimensions and corresponding dimension weights, and obtain the scheme score corresponding to each of the basic deployment schemes. The target solution determination module is used to determine the target deployment solution for the requester based on the solution score.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the distributed database deployment planning method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the distributed database deployment planning method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the distributed database deployment planning method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Distributed database deployment method and device, electronic equipment and storage medium

    CN119226262A

  • Multi-database deployment method based on server cluster

    CN119718358A

  • Operation system health assessment method and system based on multi-dimensional index dynamic weighting

    CN120353679A

  • Techniques for conditional deployment of application artifacts

    US20120066674A1

  • Evaluating adoption of computing deployment solutions

    US20170109685A1

Cited By

  • MaaS platform large model management method and system based on intelligent orchestration engine

    CN122470180A