A data governance method, apparatus, device, and storage medium
By deploying monitoring agents at the database engine layer, collecting and analyzing data reading information in real time, calculating data reading possible coefficients and performing hierarchical early warnings, and dynamically adjusting database caches and connection pools, the problems of inefficiency, waste of resources and poor adaptability under database reading problems in the existing technology are solved, and data governance for high-performance and efficient resource utilization is achieved.
Patent Information
- Application Number
- CN202510519509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing technology has problems of inefficiency, waste of resources and poor adaptability when solving database illusion problems. Especially in high concurrency and large data volume environments, it is difficult for existing methods to dynamically adjust resource allocation to cope with load changes.
By deploying a monitoring agent at the database engine layer, collecting data reading information in real time, and calculating data reading possible coefficients, hierarchical warnings are performed based on the set judgment rules, and then dynamically adjusting the database cache amount and connection pool size.
It realizes that while ensuring data consistency, dynamically adjust resource configuration, improve system performance and resource utilization, reduce calculation volume and database read and write pressure, and adapt to resource allocation under different load conditions.
Smart Images

Figure CN120030060B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer data management, and relates to a data governance method, apparatus, device and storage medium. Background Art
[0002] In modern database management, the data phantom read problem is a challenge that needs to be solved urgently. A phantom read refers to a situation in a database transaction where, during the execution of a transaction, due to insertions or deletions by other transactions, the transaction gets different result sets when it executes the same query again. To solve this problem, it is crucial to increase the database cache size. An increase in the cache can significantly improve data access speed, reduce disk I / O operations, and thus relieve the read and write pressure on the system. Especially when dealing with high concurrency and large amounts of data, cache optimization can effectively reduce transaction conflicts and lock waiting times caused by phantom reads, improving data processing efficiency and system throughput.
[0003] However, the prior art still has many deficiencies and drawbacks in solving the database phantom read problem. Common solutions include increasing the transaction isolation level and using lock mechanisms. Increasing the transaction isolation level, such as using serializable isolation, can effectively prevent phantom reads, but it also brings significant performance overhead. A high isolation level leads to increased transaction processing latency and reduced system concurrency. In addition, although lock mechanisms can ensure data consistency, in a high-concurrency environment, lock contention can lead to extended transaction waiting times and even deadlock problems. These methods solve the data consistency problem to a certain extent, but at the cost of the system's response speed and resource utilization efficiency.
[0004] Moreover, the prior art usually lacks sensitivity and adaptability to real-time data changes. Fixed cache settings cannot dynamically respond to changes in system load, resulting in unbalanced resource utilization. In low-load situations, excessive cache allocation may cause resource waste; while in high-load situations, insufficient cache leads to performance bottlenecks. This static configuration lacks flexibility and cannot be dynamically adjusted according to actual data traffic and transaction complexity, thus limiting the optimization potential of the system.
[0005] Therefore, although the prior art provides some solutions, they still face problems such as low efficiency, resource waste, and poor adaptability when dealing with phantom read problems. To overcome these deficiencies, there is an urgent need to develop an innovative method that can monitor the database status in real time and dynamically adjust the cache and connection pool to maximize system performance and resource utilization while ensuring data consistency. This method can not only solve the phantom read problem, but also improve the overall operating efficiency of the database, meet the requirements of modern data management systems for high performance and flexibility, effectively reduce the computational load and the read and write pressure on the database, and enable limited resources to serve more users. Summary of the Invention
[0006] In view of the problems existing in the above prior art, the present invention provides a data governance method, device, equipment and storage medium for solving the above technical problems.
[0007] In order to achieve the above and other purposes, the technical solutions adopted by the present invention are as follows:
[0008] The first aspect of the present invention provides a data governance method, which includes the following steps:
[0009] S1. Deploy a monitoring agent in the database engine layer of the data management system to collect data phantom read information in real time within a specific time period;
[0010] S2. Calculate the data phantom read probability coefficient η in the database within a specific time period, and based on the set judgment rules, perform a hierarchical early warning operation on the data in combination with the data phantom read probability coefficient;
[0011] S3. Based on the corresponding early warning level, comprehensively adjust the cache volume of the database and the adjustment amount of the connection pool.
[0012] The data phantom read information within the specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the data change amount △D, the number of transaction conflicts C, the number of active connections N, and the occurrence time points of each phantom read event.
[0013] Calculating the data phantom read probability coefficient η in the database within a specific time period includes the following steps:
[0014] Subtract the occurrence time point of the first phantom read event from the occurrence time point of the last phantom read event within the specific time period to obtain the total phantom read observation time window within the specific time period ;
[0015] Thus, calculate the data phantom read probability coefficient in the database within a specific time period ; where α1, α2, and α3 respectively represent the set data change weight factor, transaction conflict weight factor, and time sensitivity factor, and satisfy α1 + α2 + α3 = 1;
[0016] is the standard deviation of the time interval between two consecutive phantom read events.
[0017] Performing a hierarchical early warning operation on the data in combination with the data phantom read probability coefficient includes the following steps:
[0018] Trigger a three-level early warning mechanism based on the data phantom read probability coefficient η in the database within a specific time period:
[0019] If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a low - risk blue warning;
[0020] If the data phantoming probability coefficient η in the database satisfies within a specific time period , it is marked as a medium - risk yellow warning;
[0021] If the data phantoming probability coefficient η in the database satisfies within a specific time period , it is marked as a high - risk red warning;
[0022] where τ is a set judgment constant used to prevent it from being zero;
[0023] respectively represent the set judgment lower - limit threshold and judgment upper - limit threshold. The specific acquisition process is as follows:
[0024] Determine the judgment lower - limit reference threshold and judgment upper - limit reference threshold through historical data regression analysis, and denote them as ;
[0025] Thus, the set judgment lower - limit threshold and judgment upper - limit threshold are calculated as ;
[0026] L is the system load factor of the data management system within a specific time period, and its specific calculation formula is:
[0027] .
[0028] When the identified warning level is a low - risk blue warning, the process of comprehensively adjusting the cache volume of the database and the adjustment volume of the connection pool is as follows:
[0029] Using the calculation formula , calculate the estimated cache volume of the database at the low - risk blue warning; is the maximum cache volume of the database, L is the system load factor of the data management system within a specific time period, is the current cache volume of the database corresponding; e is the natural constant, and t is the number of consecutive risk warning cycles;
[0030] According to the calculation formula , calculate the adjustment volume of the connection pool corresponding to the database at the low - risk blue warning, respectively represent the maximum number of connections supported by the database connection pool and the current active connection number of the database connection pool corresponding, and P is the number of phantoming detection times within a specific time period;
[0031] Feed back the estimated cache volume of the database and the adjustment volume of the corresponding connection pool at the low - risk blue warning to the background terminal of the data management system to make corresponding adjustments to the database.
[0032] When the corresponding early warning level is identified as the medium - risk yellow early warning, the process of comprehensively adjusting the cache volume of the database and the adjustment volume of the connection pool is as follows:
[0033] Calculate the expected cache volume of the database when the medium - risk yellow early warning occurs and the adjustment volume of the corresponding connection pool ;
[0034] ;
[0035] ;
[0036] In the above calculation formula, T is the total number of transactions within a specific time period;
[0037] Feed back the expected cache volume of the database and the adjustment volume of the corresponding connection pool when the medium - risk yellow early warning occurs to the back - end terminal of the data management system, and make corresponding adjustments to the database.
[0038] When the early warning level is identified as the high - risk red early warning, the process of comprehensively adjusting the cache volume of the database and the adjustment volume of the connection pool is as follows:
[0039] Calculate the expected cache volume of the database when the high - risk red early warning occurs and the adjustment volume of the corresponding connection pool ;
[0040] ;
[0041] ;
[0042] In the above calculation formula, exp(·) represents the exponential function with the real number e as the base, △D represents the data change amount within a specific time period, and k is a set conversion coefficient used to convert the unit of the data change amount to the unit of the cache volume;
[0043] Feed back the expected cache volume of the database and the adjustment volume of the corresponding connection pool when the high - risk red early warning occurs to the back - end terminal of the data management system, and make corresponding adjustments to the database.
[0044] The second aspect of the present invention provides a data governance device, including: a data acquisition module, a phantom read classification module, a phantom read adjustment module, and a back - end terminal of the data management system. Each module is connected by wired and / or wireless connection methods to realize data transmission between each module;
[0045] Data acquisition module: Deploy a monitoring agent at the database engine layer in the data management system to collect data phantom read information in real - time within a specific time period;
[0046] Phantom read grading module: Calculate the possible coefficient of data phantom reads in the database within a specific time period, and based on the set judgment rules, perform grading and early warning operations on the data in combination with the possible coefficient of data phantom reads;
[0047] Phantom read adjustment module: Based on the corresponding early warning level, comprehensively adjust the cache volume of the database and the adjustment volume of the connection pool;
[0048] Data management system background terminal: Used to receive the adjustment volume of the cache volume and connection pool of the database corresponding to each early warning level, and perform corresponding adjustment operations on the database in combination with the early warning level within a specific time period.
[0049] The third aspect of the present invention provides a computer device, including a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of a data governance method as described in the present invention are implemented.
[0050] The fourth aspect of the present invention provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the steps of a data governance method as described in the present invention are implemented.
[0051] As described above, a data governance method, device, equipment and storage medium provided by the present invention have at least the following beneficial effects:
[0052] A data governance method, device, equipment and storage medium provided by the present invention, by calculating the possible coefficient of data phantom reads in the database within a specific time period and performing grading and early warning operations based on the set judgment rules, is a key step in realizing intelligent data governance. The possible coefficient of phantom reads quantifies the potential risks faced by the database and can help the system identify and classify different degrees of phantom read problems. By combining specific judgment rules, the system can automatically perform risk assessment and grading and early warning. This grading and early warning mechanism not only improves the system's response ability, but also makes resource allocation more accurate and efficient. The division of early warning levels can guide the system to take appropriate countermeasures under different risk levels to ensure that the database can quickly adjust its strategy to maintain data consistency in high-risk situations.
[0053] Based on the corresponding warning levels, comprehensively adjusting the cache volume and connection pool of the database is an effective means to achieve resource optimization and performance improvement. By dynamically adjusting the cache and connection pool, the system can flexibly allocate resources under different load conditions, improving the concurrent processing ability and data access speed. This dynamic adjustment strategy ensures that the system can quickly respond and adjust resource allocation in high-risk situations, avoiding performance bottlenecks caused by insufficient resources. At the same time, this flexible resource management mode can also avoid resource waste, reasonably reducing resource occupancy in low-risk situations and enhancing the overall efficiency of the system.
[0054] The implementation of these steps not only enhances the stability and reliability of the database but also significantly improves the resource utilization rate and response speed of the system. In modern data-intensive applications, the complexity and dynamics of data operations require the system to have a higher level of intelligence and automation. It not only meets the current requirements of data management systems for high performance and high reliability but also provides a sustainable development path for future data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 Schematic diagram of the connection of each step of the method of the present invention.
[0057] Figure 2 Schematic diagram of the structure of the data governance device shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The above content is only an example and explanation of the concept of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by this claims, they should fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] Please refer to Figure 1 shown, a data governance method, which includes the following steps:
[0061] S1. Deploy a monitoring agent at the database engine layer in the data management system to collect data phantom read information in real time during a specific time period; where the database stores, including but not limited to, spatial surveying and mapping data, urban and rural surveying and mapping data, etc.;
[0062] The data phantom read information within a specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the data change amount △D, the number of transaction conflicts C, the number of active connections N, and the occurrence time points of each phantom read event.
[0063] The number of phantom read detections indicates the number of times when, in a transaction, the results are inconsistent when querying the same data range multiple times. This is usually because other transactions insert or delete data during the execution of the transaction; the number of phantom reads occurring in each transaction can be recorded through a database monitoring tool;
[0064] The total number of transactions represents the total number of all transactions processed by the database within a specific time period, and the total number of transactions started within a certain time period can be counted from the database transaction log;
[0065] The data change amount represents the total number of data rows inserted or deleted during the execution of a transaction, and the insert and delete operations on the data in each transaction can be counted by analyzing the transaction log;
[0066] The number of transaction conflicts represents the number of lock waits or conflicts generated when multiple transactions attempt to access the same data resource simultaneously. The number of lock wait events can be recorded through the database's lock management system;
[0067] The number of active connections represents the number of client applications that are currently connected to the database and performing operations at a certain moment. The number of current active connections can be obtained from the database connection pool.
[0068] It should be added that the above-mentioned number of phantom read detections P, total number of transactions T, data change amount △D, number of transaction conflicts C, and number of active connections N are the total numbers of data phantom read information detected at all time points within a specific time period.
[0069] S2. Calculate the data phantom read probability coefficient η in the database within a specific time period, and based on the set judgment rules, perform a hierarchical early warning operation on the data in combination with the data phantom read probability coefficient;
[0070] Calculating the data phantom read probability coefficient η in the database within a specific time period includes the following steps:
[0071] Subtract the occurrence time point of the first phantom read event from the occurrence time point of the last phantom read event within a specific time period to obtain the total phantom read observation time window within the specific time period ;
[0072] Thereby calculate the data phantom read probability coefficient in the database within a specific time period ; where α1, α2, and α3 represent the set data change weight factor, transaction conflict weight factor, and time sensitivity factor respectively, and satisfy α1 + α2 + α3 = 1; α1 can be obtained through historical data analysis and regression, reflecting the contribution of data changes to phantoms; α2 is determined based on the correlation analysis of lock waiting time and the number of transaction rollbacks; α3 is dynamically adjusted according to the characteristics of the business peak period and defaults to 0.15;
[0073] is the standard deviation of the time interval between two consecutive phantom read events, with the unit of ms.
[0074] The above calculation formula is designed to calculate the possible coefficient of data phantoms in the database within a specific time period, and evaluate the risk of phantoms occurring within a specific time period through the weighted sum of multiple parameters. These parameters reflect different aspects of database operations, and can provide a reasonable estimate of the phantom risk after comprehensive consideration;
[0075] The first term of the above calculation formula combines the phantom read frequency and the impact of data changes, and uses the natural logarithm function to smooth the impact of data changes to prevent over-amplification. The second term amplifies the impact of conflicts by squaring, emphasizing the inconsistent risk that conflicts may cause in a high-concurrency environment. The third term evaluates the time instability of the occurrence of phantom read events. The larger the standard deviation, the higher the risk.
[0076] Perform a hierarchical early warning operation on the data in combination with the possible coefficient of data phantoms, including the following steps:
[0077] Trigger a three-level early warning mechanism based on the possible coefficient of data phantoms η in the database within a specific time period:
[0078] If the possible coefficient of data phantoms η in the database within a specific time period satisfies , it is marked as a low-risk blue early warning;
[0079] If the possible coefficient of data phantoms η in the database within a specific time period satisfies , it is marked as a medium-risk yellow early warning;
[0080] If the possible coefficient of data phantoms η in the database within a specific time period satisfies , it is marked as a high-risk red early warning;
[0081] where τ is a set judgment constant used to prevent it from being zero;
[0082] respectively represent the set judgment lower threshold and judgment upper threshold, and the specific acquisition process is as follows:
[0083] Determine the judgment lower limit reference threshold and judgment upper limit reference threshold through historical data regression analysis, and denote them as ;
[0084] Calculate the mean and standard deviation of the historical data phantom read probability coefficient over a number of sampling periods, and denote them as μ and σ respectively;
[0085] Then the calculation formulas for the lower limit reference threshold and the upper limit reference threshold are as follows:
[0086] ;
[0087] Thus, the set lower limit threshold and upper limit threshold are calculated. ;
[0088] L is the system load factor of the data management system for a specific time period, and its specific calculation formula is:
[0089] .
[0090] S3. Based on the corresponding warning level, comprehensively adjust the cache size of the database and the adjustment amount of the connection pool.
[0091] When the identified corresponding warning level is a low-risk blue warning, the process of comprehensively adjusting the cache size of the database and the adjustment amount of the connection pool is as follows:
[0092] Use the calculation formula , and calculate the estimated cache size of the database when there is a low-risk blue warning ; is the maximum cache size of the database, L is the system load factor of the data management system for a specific time period, is the current cache size of the database; e is the natural constant, t is the number of consecutive risk warning cycles;
[0093] It should be noted that the number of consecutive risk warning cycles t is used to dynamically adjust the cache and connection pool sizes of the database. This parameter is used to track the warning situation of the system in consecutive cycles, so as to take more flexible and appropriate adjustment measures when continuous problems occur;
[0094] Specifically, the number of consecutive warning cycles t represents the number of consecutive warnings of the system within a certain time range. The warning logic for this occurrence is the same as the warning logic within a specific time period. For example, if there was also a warning in the previous time period within a specific time period, then the number of consecutive warning cycles t = 2; in the dynamic adjustment strategy, the role of t is to introduce the time dimension, enabling the system to consider the persistence of the warning signal. By monitoring the number of consecutive warning cycles t, the system can better judge the severity of the current problem, and thus take corresponding adjustment measures to optimize the performance and stability of the database.
[0095] In the above calculation formulas is a fixed parameter used to limit the upper bound of cache adjustment. η reflects the degree of phantom reads. When the phantom read frequency is high and the system load is low, the value of η will increase; L is the system load factor, which is an indicator that comprehensively considers CPU utilization, memory occupancy, and I / O throughput. L reflects the current system load. When the system load is low, the value of L is small; the increase of t can be used to achieve the time accumulation effect, that is, the more consecutive warning cycles, the greater the adjustment amplitude.
[0096] According to the operations of the above parameters, the formula calculates represents the cache adjustment amount calculated based on the current data phantom read probability coefficient and the system load condition. 0.1 multiplied in the formula is to reserve at least 10% of the available margin to ensure that the system has enough cache resources for use. The latter part is calculated based on the data phantom read probability coefficient and the system load factor to achieve the purpose of dynamically adjusting the cache according to the phantom read risk degree and the system load condition.
[0097] Based on the calculation formula , the calculated adjustment amount of the database corresponding connection pool during the low-risk blue warning is , respectively represent the maximum number of connections supported by the database connection pool and the current active connections of the database corresponding connection pool. P is the number of phantom read detections within a specific time period;
[0098] In a database management system, a connection pool is used to manage database connections to improve resource utilization efficiency and response speed. Specifically, represents the number of connections currently in use, that is, those connections in the active state. These connections are executing transactions or queries and occupying database resources. Therefore, during the dynamic adjustment process, can be used to measure the usage of the connection pool and determine whether to adjust the size of the connection pool based on its value to optimize the performance of the database. represents the current system load status, which helps to judge the available expansion space.
[0099] The purpose of the above calculation formula is to adjust the size of the database connection pool according to the current system state and the data phantom read probability coefficient to achieve a reasonable adjustment of the database connection pool during the low-risk blue warning, where Indicates the maximum number of concurrent connections supported by the database, with the unit being the number; the number of phantom read detections indicates the number of times the results of repeated range queries within the same transaction are inconsistent. This parameter reflects the possibility of phantom reads. The unit is the number of times. The second part of the above calculation formula is to adjust the size of the connection pool according to the current number of phantom read detections. The more phantom read detections, the higher the possibility of phantom reads. Therefore, a larger connection pool is required to handle concurrent transactions to reduce the occurrence of phantom reads. At the same time, the current number of connections is used to control the adjustment range of the connection pool size by multiplying by a proportionality coefficient of 0.05, with the unit being the number. The first part of the above calculation formula is to limit the upper limit of the connection pool expansion to ensure that the number of database connections does not exceed 15% of the maximum number of connections, with the unit being the number. Finally, the smaller value of the above two calculated values is taken as the adjustment value of the connection pool to ensure that the size of the connection pool is adjusted within a reasonable range to adapt to the current phantom read situation of the system.
[0100] When a low-risk blue warning occurs, the estimated cache volume of the database and the corresponding adjustment amount of the connection pool will be fed back to the background terminal of the data management system to make corresponding adjustments to the database.
[0101] When the identified warning level is a medium-risk yellow warning, the process of comprehensively adjusting the cache volume and the adjustment amount of the connection pool of the database is as follows:
[0102] Calculate respectively the estimated cache volume of the database and the corresponding adjustment amount of the connection pool ;
[0103] ;
[0104] The logic and basis of the above calculation formula design are as follows:
[0105] First of all, considering the medium-risk state, the system may face a certain risk of phantom reads, and it is necessary to adjust the cache of the database to cope with possible data consistency problems. In the formula, represents the current cache capacity, which is the object to be adjusted. And η is the possible coefficient of data phantom reads, which reflects the degree of phantom reads in the current system. The higher it is, the greater the risk. Through this term, the hyperbolic tangent function is introduced to smooth the change of phantom read risk, making the adjustment amount more continuous and stable;
[0106] In addition, the parameter L represents the load factor of the system. The part in the formula adjusts the adjustment amount according to the system load situation. When the system load approaches or exceeds the threshold, the adjustment amount will be reduced accordingly to avoid over-adjustment resulting in a decline in system performance;
[0107] Taken together, the design of this formula takes into account the phantom read risk, the current cache capacity, and the system load. Through the comprehensive calculation of these parameters, the cache amount to be adjusted in the medium-risk state is obtained. This design enables the adjustment amount to be more targeted and flexible, dynamically adjusting the cache size according to the current state of the system to cope with the phantom read risk while avoiding unnecessary impacts on system performance.
[0108] ;
[0109] In a database system, the size of the connection pool directly affects the processing ability of concurrent transactions. Appropriate adjustment of the connection pool can optimize resource utilization and improve system throughput and response speed. η has no unit. The non-linear impact of the phantom read risk is reflected by the square of this coefficient, indicating that when the risk increases, a more significant adjustment is required. The number of phantom read detections, in units of times. It is used to evaluate the frequency of phantom reads occurring in a transaction. The total number of transactions, in units of times. It provides a benchmark to measure the proportion of phantom read events in the overall transactions.
[0110] η^2 is used to amplify the impact of the phantom read risk. The square relationship means that an increase in the risk coefficient will lead to a larger adjustment of the connection pool, reflecting a sensitive response to the medium-risk state; multiplying the ratio of P and T by 0.2 is used to quantify the frequency of phantom read events. By combining the ratio of the number of phantom reads to the total number of transactions, a relative indicator is provided to reflect the stability of the system when processing transactions; subtracting represents the unused connection space in the current system. Through this difference, it is ensured that the adjustment will not exceed the carrying capacity of the system, and at the same time, the unused resources are utilized to improve performance.
[0111] In the above calculation formula, T is the total number of transactions within a specific time period;
[0112] When there is a medium-risk yellow warning, the estimated cache amount of the database and the adjustment amount of the corresponding connection pool will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
[0113] The process of identifying a high-risk red warning and comprehensively adjusting the cache amount and connection pool adjustment amount of the database is as follows:
[0114] Calculate respectively the estimated cache amount of the database and the adjustment amount of the corresponding connection pool ;
[0115] ;
[0116] The design logic of the above calculation formula is based on the following key ideas:
[0117] A high - risk warning means that the system faces a relatively large phantom read pressure. This pressure is quantified by η, and the term in the formula is used to reflect the impact of the risk coefficient on the cache requirement. The characteristics of the exponential function make it such that when η is high, the adjustment amplitude increases significantly, reflecting that in high - risk situations, the cache needs to be rapidly increased to cope with the pressure.
[0118] The difference between the current cache capacity and the maximum capacity reflects the potential for system expansion. When the risk coefficient indicates the need for more cache, the formula ensures that it does not exceed the maximum value allowed by the hardware, preventing resource waste caused by over - allocation.
[0119] The term k times △D takes into account the direct impact of the data change volume on the cache requirement. The data change volume △D refers to the number of rows of data inserted or deleted during the execution of a transaction, with the unit of row. The more data changes, the more data needs to be accessed quickly. Therefore, it is reasonable to increase the cache to improve access efficiency. The coefficient k can be determined based on the average size of each row of data in the system.
[0120] ;
[0121] The above formula is used to dynamically adjust the size of the database connection pool in high - risk situations. High risk usually means that the system faces a relatively large load and potential performance bottlenecks. Therefore, it is necessary to flexibly adjust the connection pool to improve the concurrent processing ability.
[0122] The denominator term in the fraction in the formula is used to dynamically reflect the impact of the risk level on the connection pool requirement. The higher the risk coefficient, the greater the read - write conflicts and inconsistencies faced by the system, and more connections are needed to distribute the load.
[0123] The second term in the formula embodies an inverse - proportion control mechanism. As the risk coefficient and the current number of connections increase, the adjustment amount gradually increases, but due to the inverse - proportion characteristic, the growth rate of the adjustment gradually slows down. This design ensures that in high - risk situations, the adjustment is stable and does not lead to over - allocation of resources. By using the inverse - proportion function, the formula realizes the control of the adjustment amplitude. Even in high - risk situations, the adjustment is gradual, avoiding the impact of drastic fluctuations on system stability.
[0124] In the above calculation formula, exp(·) represents the exponential function with the real number e as the base, △D represents the data change volume within a specific time period, and k is a set conversion coefficient used to convert the unit of the data change volume to the unit of the cache volume;
[0125] When there is a high - risk red warning, the estimated cache volume of the database and the adjustment amount of the corresponding connection pool are fed back to the back - end terminal of the data management system to make corresponding adjustments to the database.
[0126] Example 2
[0127] Please refer to Figure 2 As shown, a data governance device includes: a data acquisition module, a phantom read classification module, a phantom read adjustment module, and a back-end terminal of the data management system. Each module is connected by wired and / or wireless connection to achieve data transmission between modules;
[0128] Data acquisition module: Deploy a monitoring agent at the database engine layer in the data management system to collect data phantom read information in real time within a specific time period;
[0129] Phantom read classification module: Calculate the possible coefficient of data phantom read in the database within a specific time period, and based on the set judgment rules, combine the possible coefficient of data phantom read to perform a classification warning operation on the data;
[0130] Phantom read adjustment module: Based on the corresponding warning level, comprehensively adjust the cache volume of the database and the adjustment volume of the connection pool;
[0131] Back-end terminal of the data management system: Used to receive the adjustment volume of the cache volume and connection pool of the database corresponding to each warning level, and combine the warning level within a specific time period to perform corresponding adjustment operations on the database.
[0132] Embodiment 3
[0133] A computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of a data governance method as described in the present invention are implemented.
[0134] Embodiment 4
[0135] A computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of a data governance method as described in the present invention are implemented.
[0136] It should be understood that in various embodiments of the present application, the adjustment amount of the serial numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0137] It should be understood that determining B based on A does not mean determining B only based on A, but also B can be determined based on A and / or other information.
[0138] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0139] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data governance method, characterized in that: include: S1. Deploy monitoring agents at the database engine layer in the data management system to collect data phantom read information in real time within a specific time period; The data phantom read information in the specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the amount of data changes ΔD, the number of transaction conflicts C, the number of active connections N, and the time point of occurrence of each phantom read event; S2. Calculate the data phantom reading possibility coefficient η in the database within a specific time period, and perform graded warning operations on the data based on the set judgment rules and the data phantom reading possibility coefficient; Calculating the data phantom read possibility coefficient η in a database within a specific time period includes the following steps: Subtract the time point of the first phantom read event from the time point of the last phantom read event in a specific time period to obtain the total phantom read observation time window in the specific time period. ; This calculates the probability coefficient of data phantom reading in the database within a specific time period ; Where α1, α2 and α3 represent the set data change weight factor, transaction conflict weight factor and time sensitivity factor respectively, and satisfy α1+α2+α3=1; The standard deviation of the time interval between two consecutive phantom read events; S3. Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are comprehensively adjusted.
2. A data management method according to claim 1, characterized in that: Based on the data phantom read possibility coefficient, a hierarchical warning operation is performed on the data, including the following steps: The three-level warning mechanism is triggered based on the probability coefficient η of data phantom reading in the database within a specific time period: If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a low-risk blue warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a medium-risk yellow warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a high-risk red warning; Where τ is a set judgment constant, which is used to prevent it from being zero; They represent the set lower and upper thresholds for judgment respectively. The specific acquisition process is as follows: The lower limit benchmark threshold and the upper limit benchmark threshold are determined by historical data regression analysis and are denoted as ; The set lower and upper thresholds are calculated. ; L is the system load factor of the digital management system corresponding to a specific time period, and its specific calculation formula is: 。 3. A data management method according to claim 1, characterized in that: If the corresponding warning level is identified as a low-risk blue warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Using the calculation formula , calculate the estimated cache size of the database when the low-risk blue warning occurs ; is the maximum cache size of the database, L is the system load factor of the data management system corresponding to a specific time period, is the current cache size of the database; e is a natural constant, and t is the number of consecutive risk warning cycles; According to the calculation formula , calculate the adjustment amount of the database corresponding connection pool when the low risk blue warning is in effect , They represent the maximum number of connections supported by the connection pool of the database and the current number of active connections in the connection pool corresponding to the database, respectively. P represents the number of phantom read detections within a specific time period. When a low-risk blue warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the backend terminal of the data management system to make corresponding adjustments to the database.
4. A data governance method according to claim 3, characterized in that: If the corresponding warning level is identified as a medium-risk yellow warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Calculate the estimated cache size of the database when the medium-risk yellow warning is in effect And the corresponding connection pool adjustment ; ; ; In the above calculation formula, T is the total number of transactions in a specific time period; When a medium-risk yellow warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
5. A data governance method according to claim 4, characterized in that: Identify the warning level as a high-risk red warning, and make comprehensive adjustments to the database cache and connection pool adjustments as follows: Calculate the estimated cache size of the database when a high-risk red alert occurs And the corresponding connection pool adjustment ; ; ; In the above calculation formula, exp(·) represents an exponential function with the real number e as the base, △D represents the amount of data change in a specific time period, and k is a set conversion coefficient used to convert the unit of the data change amount into the unit of the cache amount; When a high-risk red alert is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
6. A data management device, characterized in that: It is implemented based on a data governance method according to any one of claims 1 to 5, including: Data collection module: Deploy monitoring agents at the database engine layer in the data management system to collect data phantom reading information in real time within a specific time period; Phantom reading classification module: calculates the data phantom reading probability coefficient in the database within a specific time period, and performs graded warning operations on the data based on the set judgment rules and the data phantom reading probability coefficient; Phantom read adjustment module: Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are adjusted comprehensively; Data management system backend terminal: used to receive the cache amount of the database and the adjustment amount of the connection pool corresponding to each warning level, and perform corresponding adjustment operations on the database based on the warning level within a specific time period.
7. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of a data governance method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of a data governance method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Database read-write control method and device and computer readable storage medium
CN118861080A
Construction method and system for membrane material database
CN119248749A