Data access performance optimization strategy based on distributed database
By building a weighted graph model and dynamically updating the topological node execution link, optimizing the data access path of the distributed database, the problems of high cross-node query latency and unbalanced node load are solved, and data access performance is improved.
Patent Information
- Application Number
- CN202510471475.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
Existing data access based on distributed databases has performance bottlenecks, especially in Join and aggregation operations, high cross-node query latency, and excessive loading of specific nodes in local data access, and traditional load balancing strategies are weak.
By calculating the link bias coefficient of the data task to be accessed, a weighted graph model is constructed, a network access cost factor for the topology node execution link is generated, the best topology node execution link is filtered, and dynamically updated during the execution process to optimize the data access path, considering the load fluctuations of nodes in the link.
It effectively reduces the impact of excessive load in local data access centers and specific nodes on the data access process, and improves data access performance and load balancing effect.
Smart Images

Figure CN120336381A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and specifically to a data access performance optimization strategy based on a distributed database. Background Technique
[0002] With the explosive growth of cloud computing, the Internet of Things, and large-scale Internet applications, the amount of data has increased exponentially. Traditional single-machine databases have bottlenecks in scalability, high availability, and disaster tolerance. Distributed databases achieve horizontal expansion through core technologies such as data sharding, replication, and distributed transactions, and have become the underlying infrastructure in fields such as finance, e-commerce, and social networking (such as MongoDB, Cassandra, TiDB, etc.).
[0003] However, existing distributed databases have data access performance bottlenecks. There are significant network overheads in operations that require cross-node communication such as Join and aggregation, and the cross-node query latency is relatively high. In some special scenarios (such as the product flash sale scenario), local data access is relatively concentrated, resulting in a situation where a specific node has an overly high load. In this case, the effect of traditional load balancing strategies is weak. Therefore, there is an urgent need for a data access performance optimization strategy based on distributed databases. Summary of the Invention
[0004] The purpose of the present invention is to provide a data access performance optimization strategy based on a distributed database to solve the problems raised in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: A data access performance optimization strategy based on a distributed database, including:
[0006] Step S100: Obtain the data access task to be accessed, calculate the link bias coefficient of the data access task to be accessed and bind it to the data access task to be accessed;
[0007] Step S200: Real-time collect the topological information of adjacent nodes in the distributed database cluster, and construct a weighted graph model based on the topological information;
[0008] Step S300: According to the position of the distributed database cluster node to which the access information corresponding to the data access task to be accessed belongs, construct different topological node execution links corresponding to the data access task to be accessed; Combine the weighted graph model to generate network access cost factors corresponding to each topological node execution link of the data access task to be accessed; And construct a priority sequence of topological node execution links corresponding to the data access task to be accessed based on the network access cost factors;
[0009] Step S400: Based on the execution link priority sequence of the topology nodes corresponding to the data task to be accessed, filter the best topology node execution link for the data task to be accessed; and during the execution process, dynamically update and adjust the best topology node execution link for the data task to be accessed based on the load and network status of the topology nodes within the link.
[0010] Further, the formula for calculating the link bias coefficient of the data task to be accessed in step S100 is as follows:
[0011]
[0012] Among them, P represents the link bias coefficient of the data task to be accessed;
[0013] Extract the preset keywords of the data task to be accessed to obtain the keyword set of the data task to be accessed; record the total number of elements in the keyword set of the data task to be accessed as N; record the number of historical tasks to which the nth element in the keyword set of the data task to be accessed belongs as Mn;
[0014] Record the data-intensive task tendency coefficient of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs as DT (n,m) ;
[0015] DT (n,m) = r1·DTB (n,m) + r2·DTT (n,m)
[0016] DTB (n,m) represents the bandwidth utilization rate of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; DTT (n,m) represents the I / O throughput of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs;
[0017] Record the coordination-intensive task tendency coefficient of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs as CT (n,m) ;
[0018] CT (n,m) = r3·CTU (n,m) + r4·CTD (n,m)
[0019] CTU (n,m) represents the CPU utilization rate of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; CTD (n,m)It represents the reciprocal of the latency corresponding to the m-th historical task to which the n-th element in the keyword set of the data task to be accessed belongs; r1, r2, r3, and r4 respectively represent preset normalization coefficients.
[0020] The present invention analyzes the historical tasks matching the keywords of the data task to be accessed in the historical data from the perspectives of bandwidth utilization rate, I / O throughput, CPU utilization rate, network latency, etc., and divides the data task to be accessed into data-intensive tasks and coordination-intensive tasks. Data-intensive tasks mainly process large-scale data (such as: transmission, conversion, aggregation), and their resource consumption types mainly occupy bandwidth, disk I / O, and memory; while coordination-intensive tasks mainly manage distributed states (such as: transaction coordination), and their resource consumption types mainly occupy CPU computing, network latency, and lock contention. For different task types, there are also differences in the performance optimization directions they focus on; therefore, the present invention realizes the quantification of the bias of the task type to which the data task to be accessed belongs through the link bias coefficient, providing data support for constructing the weighted graph model in the subsequent steps.
[0021] Further, the step S200 includes:
[0022] The topological information of the adjacent nodes includes the distance between the physical positions corresponding to the adjacent nodes, the average network latency of data transmission between the adjacent nodes in the most recent unit time, and the maximum bandwidth occupied by data transmission between the adjacent nodes in the most recent unit time. The connection relationship between the topological nodes is preset, and each topological node corresponds to a distributed database storage unit;
[0023] Calculate the weighted value between the adjacent nodes with topological connections in the weighted graph model, and the involved calculation formula is as follows:
[0024] Q = F{QL, QBL}·[e1·QT + P·e2·QD]
[0025] Among them, Q represents the weighted value between the adjacent nodes with topological connections in the weighted graph model based on the data task to be accessed; QT represents the average network latency of data transmission between the adjacent nodes with topological connections in the most recent unit time; QD represents the maximum bandwidth occupied by data transmission between the adjacent nodes with topological connections in the most recent unit time; QL represents the distance between the physical positions corresponding to the adjacent nodes with topological connections; QBL represents the preset topological node distance threshold; when QL ≤ QBL, it is determined that F{QL, QBL} = 1; otherwise, it is determined that F{QL, QBL} = QL / QBL; e1 and e2 respectively represent preset weight factors.
[0026] In the weighted graph model constructed by the present invention, the weight values between topologically connected adjacent nodes change dynamically under the influence of multiple factors. It is not only affected by factors such as network latency for data transmission between nodes, bandwidth occupied by transmitted data, and physical location, but also interfered by the link bias coefficient of the data access task to be visited. It is the result of comprehensive quantitative analysis of multiple factors.
[0027] Further, the step S300 includes:
[0028] Step S301, obtain the location of the currently logged-in node and the location of the distributed database cluster node to which the access information corresponding to the data access task to be visited belongs, and construct different topological node execution links corresponding to the data access task to be visited; the starting point of each topological node execution link of the constructed data access task to be visited is the currently logged-in node, and the end point is the location of the distributed database cluster node to which the access information corresponding to the data access task to be visited belongs; the relationship between any two adjacent nodes in each topological node execution link of the constructed data access task to be visited in the distributed database cluster is topological connection;
[0029] Step S302, obtain the weighted graph model based on the data access task to be visited, and generate the network access cost factors corresponding to each topological node execution link of the data access task to be visited;
[0030] Step S303, sort each topological node execution link corresponding to the data access task to be visited in ascending order of the network access cost factor, and generate a topological node execution link priority sequence corresponding to the data access task to be visited.
[0031] Further, the calculation formula for generating the network access cost factors corresponding to each topological node execution link of the data access task to be visited in the step S302 is as follows:
[0032]
[0033] Among them, VC i represents the network access cost factor corresponding to the i-th topological node execution link of the data access task to be visited; Q (i,k) represents the weight value corresponding to the k-th link node and the next link node in the i-th topological node execution link of the data access task to be visited in the weighted graph model; K represents the total number of link nodes in the i-th topological node execution link of the data access task to be visited; k ∈ [1, K - 1];
[0034] Divide the difference between the maximum and minimum access volumes at each time point in the preset time interval to which the current time belongs and the next preset time interval within the same cycle in the historical data for the corresponding link node by the preset access volume, and denote the quotient as the access volume fluctuation coefficient for the corresponding link node within the preset time interval to which the current time belongs and the next preset time interval in the corresponding cycle in the historical data; X (i,k) Denote the maximum value of the access volume fluctuation coefficient of the k-th link node in the execution link of the i-th topology node corresponding to the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; X (i,k+1) Denote the maximum value of the access volume fluctuation coefficient of the (k + 1)-th link node in the execution link of the i-th topology node corresponding to the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; max{X (i,k) ,X (i,k+1)} represents the maximum value of X (i,k) and X (i,k+1) ; G() represents a decision function;
[0035] When max{X (i,k) ,X (i,k+1)} ≤ X0, then determine that G(max{X (i,k) ,X (i,k+1)}, X0) = 1; otherwise, determine that G(max{X (i,k) ,X (i,k+1)}, X0) = α · [max{X (i,k) ,X (i,k+1)} - X0] + 1; where α is a preset constant.
[0036] Further, the best topology node execution link of the data task to be accessed in step S400 is the topology node execution link with the smallest network access cost factor in the topology node execution link priority sequence corresponding to the data task to be accessed.
[0037] Further, during the execution process, when the network access delay of the link exceeds the preset delay or the load change rate of the unexecuted topology node in the link is greater than the preset change value, then adjust the best topology node execution link of the data task to be accessed; otherwise, keep the best topology node execution link of the data task to be accessed unchanged; the load change rate of the unexecuted topology node in the link is equal to the quotient of the difference between the maximum and minimum access volumes of the corresponding unexecuted topology node within the time interval formed by the start execution time to the current execution time of the corresponding link divided by the preset access volume tolerance threshold of the corresponding topology node.
[0038] Further, in the process of dynamically updating and adjusting the optimal topological node execution link of the data task to be accessed, extract the topological node execution link segments that have not been executed in the current optimal topological node execution link; denote the summary set of the starting link nodes and the topological nodes with a load change rate greater than the preset change value in the unexecuted topological node execution link segments as the marking set, calculate the correlation coefficient between each topological node execution link in the topological node execution link priority sequence corresponding to the data task to be accessed and the marking set, and use the topological node execution link with the smallest correlation coefficient as the optimal topological node execution link of the adjusted data task to be accessed; the correlation coefficient is equal to the ratio of the number of link nodes belonging to the marking set in the corresponding topological node execution link to the total number of link nodes in the corresponding topological node execution link.
[0039] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0040] (1) The present invention can quantify the corresponding link bias coefficient in combination with the data task type of the data task to be accessed, and dynamically construct a weighted graph model in combination with the topological information of adjacent nodes in the distributed database cluster, providing data support for subsequent screening of the optimal topological node execution link of the data task to be accessed;
[0041] (2) In the present invention, not only the latency of cross-node queries is considered, but also the load fluctuations of the link nodes passing through in the execution link during the data access process are taken into account, realizing real-time dynamic update of the optimal topological node execution link of the data task to be accessed, and effectively reducing the impact of the situation of concentrated local data access and excessive load on specific nodes on the data access process of the data task to be accessed. Description of the Drawings
[0042] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0043] Figure 1 is a schematic flow chart of a data access performance optimization strategy based on a distributed database according to the present invention. Detailed Embodiments
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] The present invention provides a technical solution: as Figure 1As shown in the figure, in this embodiment, a data access performance optimization strategy based on a distributed database includes:
[0046] Step S100: Obtain the data access task to be accessed, calculate the link bias coefficient of the data access task to be accessed, and bind it to the data access task to be accessed;
[0047] The formula for calculating the link bias coefficient of the data access task to be accessed in step S100 is as follows:
[0048]
[0049] Where P represents the link bias coefficient of the data access task to be accessed;
[0050] Extract the preset keywords of the data access task to be accessed to obtain the keyword set of the data access task to be accessed; record the total number of elements in the keyword set of the data access task to be accessed as N; record the number of historical tasks to which the nth element in the keyword set of the data access task to be accessed belongs as Mn;
[0051] Record the data-intensive task tendency coefficient of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs as DT (n,m) ;
[0052] DT (n,m) = r1·DTB (n,m) + r2·DTT (n,m)
[0053] DTB (n,m) represents the bandwidth utilization rate of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs; DTT (n,m) represents the I / O throughput of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs;
[0054] Record the coordination-intensive task tendency coefficient of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs as CT (n,m) ;
[0055] CT (n,m) = r3·CTU (n,m) + r4·CTD (n,m)
[0056] CTU (n,m) represents the CPU utilization rate of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs; CTD (n,m) represents the reciprocal of the corresponding delay of the mth historical task to which the nth element in the keyword set of the data access task to be accessed belongs; r1, r2, r3, and r4 respectively represent the preset normalization coefficients.
[0057] Step S200: Collect the topology information of adjacent nodes in the distributed database cluster in real time, and construct a weighted graph model based on the topology information.
[0058] The step S200 includes:
[0059] The topology information of the adjacent nodes includes the distance between the physical positions corresponding to the adjacent nodes, the average network delay of data transmission between the adjacent nodes in the most recent unit time, and the maximum bandwidth occupied by data transmission between the adjacent nodes in the most recent unit time. The connection relationship between topology nodes is preset, and each topology node corresponds to a distributed database storage unit.
[0060] Calculate the weighted value between adjacent nodes with topological connections in the weighted graph model. The involved calculation formula is as follows:
[0061] Q = F{QL, QBL}·[e1·QT + P·e2·QD]
[0062] Among them, Q represents the weighted value between adjacent nodes with topological connections in the weighted graph model based on the data access task to be accessed; QT represents the average network delay of data transmission between adjacent nodes with topological connections in the most recent unit time; QD represents the maximum bandwidth occupied by data transmission between adjacent nodes with topological connections in the most recent unit time; QL represents the distance between the physical positions corresponding to adjacent nodes with topological connections; QBL represents a preset topological node distance threshold; when QL ≤ QBL, it is determined that F{QL, QBL} = 1; otherwise, it is determined that F{QL, QBL} = QL / QBL; e1 and e2 respectively represent preset weight factors.
[0063] Step S300: According to the position of the distributed database cluster node to which the access information corresponding to the data access task to be accessed belongs, construct different topological node execution links corresponding to the data access task to be accessed; combine the weighted graph model to generate network access cost factors corresponding to each topological node execution link of the data access task to be accessed; and construct a priority sequence of topological node execution links corresponding to the data access task to be accessed based on the network access cost factors.
[0064] The step S300 includes:
[0065] Step S301: Obtain the current logged-in node location and the location of the distributed database cluster node to which the access information corresponding to the data task to be accessed belongs, and construct different topological node execution links corresponding to the data task to be accessed; the starting point of each topological node execution link of the constructed data task to be accessed is the current logged-in node and the end point is the location of the distributed database cluster node to which the access information corresponding to the data task to be accessed belongs; the relationship between any two adjacent nodes in each topological node execution link of the constructed data task to be accessed in the distributed database cluster is topological connection;
[0066] Step S302: Obtain the weighted graph model based on the data task to be accessed, and generate network access cost factors corresponding to each topological node execution link of the data task to be accessed;
[0067] The calculation formula for generating the network access cost factors corresponding to each topological node execution link of the data task to be accessed in Step S302 is as follows:
[0068]
[0069] Among them, VC i represents the network access cost factor corresponding to the i-th topological node execution link of the data task to be accessed; Q (i,k) represents the weighted value corresponding to the k-th link node and the next link node in the i-th topological node execution link of the data task to be accessed in the weighted graph model; K represents the total number of link nodes in the i-th topological node execution link of the data task to be accessed; k ∈ [1, K - 1];
[0070] Divide the difference between the maximum access volume and the minimum access volume of each time point for the corresponding link node in the preset time interval to which the current time belongs and the next preset time interval within the same cycle in the historical data by the preset access volume, and denote the quotient as the access volume fluctuation coefficient of the corresponding link node in the preset time interval to which the current time belongs and the next preset time interval within the corresponding cycle in the historical data; X (i,k) represents the maximum value of the access volume fluctuation coefficient of the k-th link node in the i-th topological node execution link of the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; X (i,k+1) represents the maximum value of the access volume fluctuation coefficient of the k + 1-th link node in the i-th topological node execution link of the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; max{X (i,k) , X (i,k+1)} represents the maximum value of X (i,k) and X (i,k+1) ; G() represents a decision function;
[0071] When max{X (i,k) , X (i,k+1)} ≤ X0, then it is determined that G(max{X (i,k) , X (i,k+1)}, X0) = 1; otherwise, it is determined that G(max{X (i,k) , X (i,k+1)}, X0) = α · [max{X (i,k) , X (i,k+1)} - X0] + 1; where α is a preset constant.
[0072] Step S303: Sort the execution links of each topology node corresponding to the data task to be accessed in ascending order of the network access cost factor, and generate a priority sequence of the execution links of the topology nodes corresponding to the data task to be accessed.
[0073] Step S400: Based on the priority sequence of the execution links of the topology nodes corresponding to the data task to be accessed, screen the best execution link of the topology node of the data task to be accessed; and during the execution process, dynamically update and adjust the best execution link of the topology node of the data task to be accessed based on the load and network status of the topology nodes within the link.
[0074] The best execution link of the topology node of the data task to be accessed in Step S400 is the execution link of the topology node with the smallest network access cost factor in the priority sequence of the execution links of the topology nodes corresponding to the data task to be accessed.
[0075] During the execution process, when the network access delay of the link exceeds the preset delay or the load change rate of the unexecuted topology nodes within the link is greater than the preset change value, then adjust the best execution link of the topology node of the data task to be accessed; otherwise, keep the best execution link of the topology node of the data task to be accessed unchanged; the load change rate of the unexecuted topology nodes within the link is equal to the quotient obtained by dividing the difference between the maximum access volume and the minimum access volume of the corresponding unexecuted topology nodes within the time interval formed by the start execution time to the current execution time of the corresponding link by the preset access volume tolerance threshold of the corresponding topology node.
[0076] During the process of dynamically updating and adjusting the best topological node execution link of the data task to be accessed, extract the topological node execution link segment that has not been executed in the current best topological node execution link; denote the summary set of the starting link node and the topological nodes with a load change rate greater than the preset change value in the unexecuted topological node execution link segment as the marking set, calculate the correlation coefficient between each topological node execution link in the topological node execution link priority sequence corresponding to the data task to be accessed and the marking set, and use the topological node execution link with the smallest correlation coefficient as the best topological node execution link of the adjusted data task to be accessed; the correlation coefficient is equal to the ratio of the number of link nodes belonging to the marking set in the corresponding topological node execution link to the total number of link nodes in the corresponding topological node execution link.
[0077] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0078] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data access performance optimization strategy based on a distributed database, characterized in that, Including: Step S100: Obtain the data access task to be accessed, calculate the link bias coefficient of the data access task to be accessed and bind it to the data access task to be accessed; Step S200: Real-time collect the topology information of adjacent nodes in the distributed database cluster, and construct a weighted graph model based on the topology information; Step S300: According to the position of the distributed database cluster node to which the access information corresponding to the data access task to be accessed belongs, construct different topology node execution links corresponding to the data access task to be accessed; Combine the weighted graph model to generate network access cost factors corresponding to each topology node execution link of the data access task to be accessed; And construct a topology node execution link priority sequence corresponding to the data access task to be accessed based on the network access cost factors; Step S400: Based on the topology node execution link priority sequence corresponding to the data access task to be accessed, screen the best topology node execution link of the data access task to be accessed; And during the execution process, dynamically update and adjust the best topology node execution link of the data access task to be accessed based on the load and network status of the topology nodes within the link.
2. The data access performance optimization strategy based on a distributed database according to claim 1, characterized in that: The formula for calculating the link bias coefficient of the data access task to be accessed in step S100 is as follows: Among them, P represents the link bias coefficient of the data task to be accessed; perform preset keyword extraction on the data task to be accessed to obtain the keyword set of the data task to be accessed; denote the total number of elements in the keyword set of the data task to be accessed as N; denote the number of historical tasks to which the nth element in the keyword set of the data task to be accessed belongs as Mn; DTB (n,m) represents the bandwidth utilization rate of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; DTT (n,m) represents the I / O throughput of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; CTU (n,m) represents the CPU utilization rate of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; CTD (n,m) represents the reciprocal of the corresponding delay of the mth historical task to which the nth element in the keyword set of the data task to be accessed belongs; r1, r2, r3, and r4 respectively represent preset normalization coefficients.
3. The data access performance optimization strategy based on a distributed database according to claim 2, characterized in that: In step S200, it includes: The topology information of the adjacent nodes includes the distance between the physical positions corresponding to the adjacent nodes, the average network delay of data transmission between the adjacent nodes in the most recent unit time, and the maximum bandwidth occupied by data transmission between the adjacent nodes in the most recent unit time. The connection relationship between the topology nodes is preset, and each topology node corresponds to a distributed database storage unit; Calculating the weight value between adjacent nodes with topological connections in the weighted graph model involves the following calculation formula: Q = F{QL, QBL}·[e1·QT + P·e2·QD] Where, Q represents the weight value between adjacent nodes with topological connections in the weighted graph model based on the data access task to be accessed; QT represents the average network delay of data transmission between adjacent nodes with topological connections in the most recent unit time; QD represents the maximum bandwidth occupied by data transmission between adjacent nodes with topological connections in the most recent unit time; QL represents the distance between the physical positions corresponding to the adjacent nodes with topological connections; QBL represents the preset topology node distance threshold; When QL ≤ QBL, it is determined that F{QL, QBL} = 1; Otherwise, it is determined that F{QL, QBL} = QL / QBL; e1 and e2 respectively represent preset weight factors.
4. A data access performance optimization strategy based on a distributed database according to claim 1, characterized in that: Step S300 includes: Step S301: Obtain the current login node position and the position of the distributed database cluster node to which the access information corresponding to the data access task to be accessed belongs, and construct different topology node execution links corresponding to the data access task to be accessed; The starting point of each topology node execution link of the constructed data access task to be accessed is the current login node, and the end point is the position of the distributed database cluster node to which the access information corresponding to the data access task to be accessed belongs; The relationship between any two adjacent nodes in each topology node execution link of the constructed data access task to be accessed in the distributed database cluster is topological connection; Step S302: Obtain a weighted graph model based on the data task to be accessed, and generate network access cost factors corresponding to each topological node execution link of the data task to be accessed. Step S303: Sort each topological node execution link corresponding to the data task to be accessed in ascending order of the network access cost factor, and generate a topological node execution link priority sequence corresponding to the data task to be accessed.
5. A data access performance optimization strategy based on a distributed database according to claim 4, characterized in that: The calculation formula for generating the network access cost factors corresponding to each topological node execution link of the data task to be accessed in step S302 is as follows: Among them, VC i represents the network access cost factor corresponding to the execution link of the i-th topology node for the data task to be accessed; Q (i,k) represents the weighted value corresponding to the k-th link node and the next link node in the execution link of the i-th topology node for the data task to be accessed in the weighted graph model; K represents the total number of link nodes in the execution link of the i-th topology node for the data task to be accessed; k ∈ [1, K - 1]; Divide the difference between the maximum and minimum access volumes at each time point in the preset time interval to which the current time belongs and the next preset time interval within the same cycle in the historical data by the preset access volume, and denote the quotient as the access volume fluctuation coefficient for the corresponding link node in the preset time interval to which the current time belongs and the next preset time interval within the corresponding cycle in the historical data; X (i,k) Denote the maximum value of the access volume fluctuation coefficient of the k-th link node in the execution link of the i-th topology node corresponding to the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; X (i,k+1) Denote the maximum value of the access volume fluctuation coefficient of the (k + 1)-th link node in the execution link of the i-th topology node corresponding to the data task to be accessed in the historical data within the preset time interval to which the current time belongs and the next preset time interval; max{X (i,k) , X (i,k+1)} denotes the maximum value of X (i,k) and X (i,k+1) ; G() represents a decision function; When max{X (i,k) ,X (i,k+1)} ≤ X0, then it is determined that G(max{X (i,k) ,X (i,k+1)}, X0) = 1; otherwise, it is determined that G(max{X (i,k) ,X (i,k+1)}, X0) = α·[max{X (i,k) ,X (i,k+1)}-X0]+1; where α is a preset constant.
6. The data access performance optimization strategy based on a distributed database according to claim 1, characterized in that: In step S400, the optimal topological node execution link of the data task to be accessed is the topological node execution link with the smallest network access cost factor in the topological node execution link priority sequence corresponding to the data task to be accessed.
7. A data access performance optimization strategy based on a distributed database according to claim 1, characterized in that: During the execution process, when the network access delay of the link exceeds the preset delay or the load change rate of the unexecuted topological nodes in the link is greater than the preset change value, the optimal topological node execution link of the data task to be accessed is adjusted; otherwise, the optimal topological node execution link of the data task to be accessed remains unchanged; the load change rate of the unexecuted topological nodes in the link is equal to the quotient obtained by dividing the difference between the maximum access volume and the minimum access volume of the corresponding unexecuted topological nodes within the time interval formed by the start execution time to the current execution time of the corresponding link by the preset access volume threshold of the corresponding topological node.
8. A data access performance optimization strategy based on a distributed database according to claim 7, characterized in that: During the process of dynamically updating and adjusting the optimal topological node execution link of the data task to be accessed, extract the unexecuted topological node execution link segments in the current optimal topological node execution link; denote the summary set of the start link nodes and the topological nodes with a load change rate greater than the preset change value in the unexecuted topological node execution link segments as the marked set, calculate the correlation coefficient between each topological node execution link in the topological node execution link priority sequence corresponding to the data task to be accessed and the marked set, and use the topological node execution link with the smallest correlation coefficient as the adjusted optimal topological node execution link of the data task to be accessed; the correlation coefficient is equal to the ratio of the number of link nodes belonging to the marked set in the corresponding topological node execution link to the total number of link nodes in the corresponding topological node execution link.