Data management method and device, equipment and storage medium
By collecting and analyzing database reading information in real time, calculating data reading possible coefficients and performing hierarchical warnings, and dynamically adjusting caches and connection pools, solving the problems of inefficiency, waste of resources and poor adaptability under database reading problems in the existing technology, achieving efficient data governance and system performance improvement.
Patent Information
- Application Number
- CN202510519509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art has problems of inefficiency, waste of resources and poor adaptability when solving database illusion problems, especially in environments with high concurrency and large data volumes.
By deploying a monitoring agent at the database engine layer, collecting data reading information in real time, calculating data reading possible coefficients, and performing hierarchical warnings based on the set judgment rules, and then dynamically adjusting the database cache amount and connection pool size.
It achieves the maximization of system performance and resource utilization while ensuring data consistency, improves the overall operation efficiency of the database, and reduces the calculation amount and the read and write pressure of the database.
Smart Images

Figure CN120030060A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer data management and relates to a data management method, device, equipment and storage medium. Background Art
[0002] In modern database management, data phantom reads are a challenge that needs to be solved urgently. Phantom reads refer to the situation in which a transaction gets a different result set when it executes the same query again due to insert or delete operations of other transactions during the execution of a transaction in database transaction processing. In order to solve this problem, it is crucial to increase the amount of database cache. Increasing the cache can significantly increase data access speed and reduce disk I / O operations, thereby reducing the read and write pressure of the system. Especially when dealing with high concurrency and large data volumes, cache optimization can effectively reduce transaction conflicts and lock waiting times caused by phantom reads, and improve data processing efficiency and system throughput.
[0003] However, the existing technologies still have many defects and drawbacks in solving the problem of database phantom reads. Common solutions include increasing the transaction isolation level and using a lock mechanism. Increasing the transaction isolation level, such as using serialized isolation, can effectively prevent phantom reads, but it also brings significant performance overhead. A high isolation level will increase the latency of transaction processing and reduce the concurrency of the system. In addition, although the lock mechanism can ensure data consistency, in a high-concurrency environment, lock competition will cause the transaction waiting time to be extended and may even cause a deadlock problem. These methods solve the data consistency problem to a certain extent, but at the expense of the system's response speed and resource utilization efficiency.
[0004] In addition, existing technologies usually lack sensitivity and adaptability to real-time data changes. Fixed cache settings cannot dynamically respond to changes in system load, resulting in uneven resource utilization. Under low load conditions, excessive cache allocation may cause resource waste; while under high load conditions, insufficient cache will lead to performance bottlenecks. This static configuration lacks flexibility and cannot be dynamically adjusted according to actual data traffic and transaction complexity, thus limiting the optimization potential of the system.
[0005] Therefore, although existing technologies provide some solutions, they still face problems such as low efficiency, waste of resources and poor adaptability when dealing with phantom reads. In order to overcome these shortcomings, it is urgent to develop an innovative method that can monitor the database status in real time and dynamically adjust the cache and connection pool to maximize system performance and resource utilization while ensuring data consistency. This method can not only solve the phantom read problem, but also improve the overall operation efficiency of the database, meet the requirements of modern data management systems for high performance and flexibility, and effectively reduce the amount of calculation and the read and write pressure of the database, so that limited resources can serve more users. Summary of the invention
[0006] In view of the above problems existing in the prior art, the present invention provides a data governance method, device, equipment and storage medium for solving the above technical problems.
[0007] In order to achieve the above purpose and other purposes, the technical solution adopted by the present invention is as follows: A first aspect of the present invention provides a data governance method, the method comprising the following steps: S1. Deploy monitoring agents at the database engine layer in the data management system to collect data phantom read information in real time within a specific time period; S2, calculating the data phantom reading possibility coefficient η in the database within a specific time period, and performing a graded warning operation on the data based on the set judgment rules and the data phantom reading possibility coefficient; S3. Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are comprehensively adjusted.
[0008] The data phantom read information in the specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the data change amount ΔD, the number of transaction conflicts C, the number of active connections N, and the time point of occurrence of each phantom read event.
[0009] Calculating the data phantom read possibility coefficient η in a database within a specific time period includes the following steps: Subtract the time point of the first phantom read event from the time point of the last phantom read event in a specific time period to obtain the total phantom read observation time window in the specific time period. ; This calculates the probability coefficient of data phantom reading in the database within a specific time period ; Where α1, α2 and α3 represent the set data change weight factor, transaction conflict weight factor and time sensitivity factor respectively, and satisfy α1+α2+α3=1; The standard deviation of the time interval between two consecutive phantom read events.
[0010] Based on the data phantom read possibility coefficient, a hierarchical warning operation is performed on the data, including the following steps: The three-level warning mechanism is triggered based on the probability coefficient η of data phantom reading in the database within a specific time period: If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a low-risk blue warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a medium-risk yellow warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a high-risk red warning; Where τ is a set judgment constant, which is used to prevent it from being zero; They represent the set lower and upper thresholds for judgment respectively. The specific acquisition process is as follows: The lower limit benchmark threshold and the upper limit benchmark threshold are determined by historical data regression analysis and are denoted as ; The set lower and upper thresholds are calculated. ; L is the system load factor of the digital management system corresponding to a specific time period, and its specific calculation formula is: .
[0011] If the corresponding warning level is identified as a low-risk blue warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Using the calculation formula , calculate the estimated cache size of the database when the low-risk blue warning occurs ; is the maximum cache size of the database, L is the system load factor of the data management system corresponding to a specific time period, is the current cache size of the database; e is a natural constant, and t is the number of consecutive risk warning cycles; According to the calculation formula , calculate the adjustment amount of the database corresponding connection pool when the low risk blue warning is in effect , They represent the maximum number of connections supported by the connection pool of the database and the current number of active connections in the connection pool corresponding to the database, respectively. P represents the number of phantom read detections within a specific time period. When a low-risk blue warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the backend terminal of the data management system to make corresponding adjustments to the database.
[0012] If the corresponding warning level is identified as a medium-risk yellow warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Calculate the estimated cache size of the database when the medium-risk yellow warning is in effect And the corresponding connection pool adjustment ; ; ; In the above calculation formula, T is the total number of transactions in a specific time period; When a medium-risk yellow warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
[0013] Identify the warning level as a high-risk red warning, and make comprehensive adjustments to the database cache and connection pool adjustments as follows: Calculate the estimated cache size of the database when a high-risk red alert occurs And the corresponding connection pool adjustment ; ; ; In the above calculation formula, exp(·) represents an exponential function with the real number e as the base, △D represents the amount of data change in a specific time period, and k is a set conversion coefficient used to convert the unit of the data change amount into the unit of the cache amount; When a high-risk red alert is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
[0014] The second aspect of the present invention provides a data management device, including: a data acquisition module, a phantom reading classification module, a phantom reading adjustment module and a data management system background terminal, each module is connected by wired and / or wireless connection to achieve data transmission between each module; Data collection module: Deploy monitoring agents at the database engine layer in the data management system to collect data phantom reading information in real time within a specific time period; Phantom reading classification module: calculates the data phantom reading probability coefficient in the database within a specific time period, and performs graded warning operations on the data based on the set judgment rules and the data phantom reading probability coefficient; Phantom read adjustment module: Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are adjusted comprehensively; Data management system backend terminal: used to receive the cache amount of the database and the adjustment amount of the connection pool corresponding to each warning level, and perform corresponding adjustment operations on the database based on the warning level within a specific time period.
[0015] A third aspect of the present invention provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of a data governance method as described in the present invention.
[0016] A fourth aspect of the present invention provides a computer-readable storage medium having computer-readable instructions stored thereon, and when the computer-readable instructions are executed by a processor, the steps of a data governance method as described in the present invention are implemented.
[0017] As described above, the data governance method, device, equipment and storage medium provided by the present invention have at least the following beneficial effects: The data governance method, device, equipment and storage medium provided by the present invention are key steps to realize intelligent data governance by calculating the data phantom read possibility coefficient in the database within a specific time period and performing graded warning operations based on the set judgment rules. The phantom read possibility coefficient quantifies the potential risks faced by the database and can help the system identify and classify phantom read problems of different degrees. By combining specific judgment rules, the system can automatically perform risk assessment and graded warning. This graded warning mechanism not only improves the responsiveness of the system, but also makes resource allocation more accurate and efficient. The division of warning levels can guide the system to take appropriate countermeasures at different risk levels, ensuring that the database can quickly adjust its strategy to maintain data consistency in high-risk situations.
[0018] Comprehensively adjusting the database cache and connection pool based on the corresponding warning level is an effective means to achieve resource optimization and performance improvement. By dynamically adjusting the cache and connection pool, the system can flexibly allocate resources under different load conditions, improve concurrent processing capabilities and data access speed. This dynamic adjustment strategy ensures that the system can quickly respond and adjust resource configuration in high-risk situations, avoiding performance bottlenecks caused by insufficient resources. At the same time, this flexible resource management mode can also avoid resource waste, reasonably reduce resource usage in low-risk situations, and improve the overall efficiency of the system.
[0019] The implementation of these steps not only enhances the stability and reliability of the database, but also significantly improves the resource utilization and response speed of the system. In modern data-intensive applications, the complexity and dynamism of data operations require the system to have a higher level of intelligence and automation. It not only meets the current data management system's requirements for high performance and high reliability, but also provides a sustainable development path for future data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0021] Figure 1It is a schematic diagram of the connection of each step of the method of the present invention.
[0022] Figure 2 It is a structural schematic diagram of the data management device shown in the present invention. DETAILED DESCRIPTION
[0023] The above contents in combination with the implementation of the present invention are merely examples and explanations of the concept of the present invention. The technical personnel in the relevant technical field may make various modifications or supplements to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the claims, they shall all fall within the protection scope of the present invention.
[0024] Example 1 See also Figure 1 As shown, a data governance method comprises the following steps: S1. Deploy monitoring agents at the database engine layer in the data management system to collect data phantom information in real time within a specific time period; the database stores data including but not limited to spatial mapping data, urban and rural mapping data, etc.; The data phantom read information in the specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the data change amount ΔD, the number of transaction conflicts C, the number of active connections N, and the time point of occurrence of each phantom read event.
[0025] The number of phantom read detections indicates the number of times the same data range is queried multiple times in a transaction, resulting in inconsistent results. This is usually because other transactions insert or delete data during the transaction execution. The number of phantom reads that occur in each transaction can be recorded through database monitoring tools. The total number of transactions indicates the total number of all transactions processed by the database in a specific time period. All transactions started in a certain time period can be counted from the database transaction log; The amount of data changes refers to the total number of data rows inserted or deleted during the execution of a transaction. By analyzing the transaction log, the insertion and deletion operations of data in each transaction can be counted. The transaction conflict count indicates the number of lock waits or conflicts that occur when multiple transactions attempt to access the same data resource at the same time. The number of lock wait events can be recorded through the database lock management system; The number of active connections indicates the number of clients that are connecting to the database and performing operations at a certain moment. The number of current active connections can be obtained from the database connection pool.
[0026] It should be added that the above-mentioned phantom read detection times P, total transaction number T, data change amount △D, transaction conflict times C, and active connection number N are the sum of data phantom read information detected at all time points within a specific time period.
[0027] S2. Calculate the data phantom reading possibility coefficient η in the database within a specific time period, and perform graded warning operations on the data based on the set judgment rules and the data phantom reading possibility coefficient; Calculating the data phantom read possibility coefficient η in a database within a specific time period includes the following steps: Subtract the time point of the first phantom read event from the time point of the last phantom read event in a specific time period to obtain the total phantom read observation time window in the specific time period. ; This calculates the probability coefficient of data phantom reading in the database within a specific time period ; α1, α2 and α3 represent the set data change weight factor, transaction conflict weight factor and time sensitivity factor respectively, and satisfy α1+α2+α3=1; α1 can be obtained through historical data analysis and regression, reflecting the contribution of data changes to phantom reads; α2 is determined based on the correlation analysis between lock waiting time and transaction rollback times; α3 is dynamically adjusted according to the characteristics of business peak hours, and the default value is 0.15; The standard deviation of the time interval between two consecutive phantom read events, in ms.
[0028] The above calculation formula is designed to calculate the probability coefficient of phantom reading of data in the database within a specific time period, and to evaluate the risk of phantom reading within a specific time period by weighted summation of multiple parameters. These parameters reflect different aspects of database operations, and after comprehensive consideration, they can provide a reasonable estimate of the risk of phantom reading; The first term of the above calculation formula combines the frequency of phantom reads and the impact of data changes, and uses the natural logarithm function to smooth the impact of data changes to prevent over-amplification. The second term amplifies the impact of conflicts by squaring, emphasizing the risk of inconsistency that conflicts may cause in a high-concurrency environment. The third term evaluates the instability of the time when phantom read events occur. The larger the standard deviation, the higher the risk.
[0029] Based on the data phantom read possibility coefficient, a hierarchical warning operation is performed on the data, including the following steps: The three-level warning mechanism is triggered based on the probability coefficient η of data phantom reading in the database within a specific time period: If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a low-risk blue warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a medium-risk yellow warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a high-risk red warning; Where τ is a set judgment constant, which is used to prevent it from being zero; They represent the set lower and upper thresholds for judgment respectively. The specific acquisition process is as follows: The lower limit benchmark threshold and the upper limit benchmark threshold are determined by historical data regression analysis and are denoted as ; The mean and standard deviation of the phantom read probability coefficient of historical data are calculated through several sampling periods, and are denoted as μ and σ respectively; The calculation formulas for judging the lower limit reference threshold and the upper limit reference threshold are as follows: ; The set lower and upper thresholds are calculated. ; L is the system load factor of the digital management system corresponding to a specific time period, and its specific calculation formula is: .
[0030] S3. Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are comprehensively adjusted.
[0031] If the corresponding warning level is identified as a low-risk blue warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Using the calculation formula , calculate the estimated cache size of the database when the low-risk blue warning occurs ; is the maximum cache size of the database, L is the system load factor of the data management system corresponding to a specific time period, is the current cache size of the database; e is a natural constant, and t is the number of consecutive risk warning cycles; It should be explained that the number of consecutive risk warning cycles t is used to dynamically adjust the database cache and connection pool size. This parameter is used to track the situation of the system warning in several consecutive cycles, so as to take more flexible and appropriate adjustment measures when continuity problems occur; Specifically, the number of consecutive warning cycles t indicates the number of consecutive warnings in the system within a certain time range. The logic of the warning is the same as the warning logic in a specific time period. For example, if a warning also appears in the previous time period within a specific time period, the number of consecutive warning cycles t=2; in the dynamic adjustment strategy, the role of t is to introduce the time dimension so that the system can consider the continuity of the warning signal. By monitoring the number of consecutive warning cycles t, the system can better judge the severity of the current problem and take corresponding adjustment measures to optimize the performance and stability of the database.
[0032] In the above calculation formula is a fixed parameter used to limit the upper limit of cache adjustment. η reflects the degree of phantom reading. When the phantom reading frequency is high and the system load is low, the value of η will increase. L is the system load factor, which is an indicator that comprehensively considers CPU utilization, memory occupancy, and I / O throughput. L reflects the current system load. When the system load is low, the value of L is small. The increase of t can be used to achieve the time accumulation effect, that is, the more consecutive warning cycles, the greater the adjustment range. According to the operation of the above parameters, the formula calculates Indicates the cache adjustment amount calculated based on the current data phantom read probability coefficient and system load. The 0.1 in the formula is multiplied by The purpose is to reserve at least 10% of the storage margin to ensure that the system has sufficient cache resources available. The latter part is calculated based on the data phantom read possibility coefficient and the system load factor to achieve the purpose of dynamically adjusting the cache according to the degree of phantom read risk and system load conditions.
[0033] According to the calculation formula , calculate the adjustment amount of the database corresponding connection pool when the low risk blue warning is in effect , They represent the maximum number of connections supported by the connection pool of the database and the current number of active connections in the connection pool corresponding to the database, respectively. P represents the number of phantom read detections within a specific time period. In database management systems, connection pools are used to manage database connections to improve resource utilization efficiency and response speed. Specifically, Indicates the number of connections currently in use, that is, those that are active. These connections are executing transactions or queries and occupying database resources. Therefore, during the dynamic adjustment process, It can be used to measure the usage of the connection pool and decide whether to adjust the size of the connection pool based on its value to optimize the performance of the database, indicate the current system load status, and help determine the available expansion space.
[0034] The purpose of the above calculation formula is to adjust the database connection pool size according to the current system status and the data phantom read probability coefficient, so as to achieve reasonable adjustment of the database connection pool during low-risk blue warning. Indicates the maximum number of concurrent connections supported by the database, in units; phantom read detection times indicates the number of inconsistent results of repeated range queries in the same transaction. This parameter reflects the possibility of phantom reads. In units of times. The second part of the above calculation formula is to adjust the size of the connection pool according to the current number of phantom read detection times. The more phantom read detection times, the higher the possibility of phantom reads, so a larger connection pool is needed to handle concurrent transactions to reduce the occurrence of phantom reads. At the same time, the current number of connections is multiplied by a proportional coefficient of 0.05 to control the adjustment range of the connection pool size, in units. The first part of the above calculation formula is to limit the upper limit of the connection pool expansion to ensure that the number of database connections does not exceed 15% of the maximum number of connections, in units. Finally, the smaller value of the above two calculated values is taken as the connection pool adjustment value to ensure that the size of the connection pool is adjusted within a reasonable range to adapt to the phantom read situation of the current system.
[0035] When a low-risk blue warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the backend terminal of the data management system to make corresponding adjustments to the database.
[0036] If the corresponding warning level is identified as a medium-risk yellow warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Calculate the estimated cache size of the database when the medium-risk yellow warning is in effect And the corresponding connection pool adjustment ; ; The logic and basis of the above calculation formula are as follows: First, considering that the system may face a certain risk of phantom reading in the medium-risk state, it is necessary to adjust the database cache to deal with possible data consistency issues. Indicates the current cache capacity, which is the object that needs to be adjusted. η is the data phantom read probability coefficient, which reflects the phantom read degree of the current system. The higher the phantom read, the greater the risk. In this item, a hyperbolic tangent function is introduced to smooth the changes in phantom read risk, making the adjustment amount more continuous and stable; In addition, the parameter L represents the load factor of the system. Some of them adjust the adjustment amount according to the system load. When the system load approaches or exceeds the threshold, the adjustment amount will be reduced accordingly to avoid excessive adjustment and thus degrade the system performance. Taken together, the design of this formula takes into account the phantom read risk, the current cache capacity, and the system load. Through the comprehensive calculation of these parameters, the cache amount to be adjusted under the medium-risk state is obtained. This design enables the adjustment amount to be more targeted and flexible, dynamically adjusting the cache size according to the current state of the system to cope with the phantom read risk while avoiding unnecessary impacts on system performance.
[0037] ; In a database system, the size of the connection pool directly affects the processing ability of concurrent transactions. Appropriate adjustment of the connection pool can optimize resource utilization and improve system throughput and response speed. η has no unit. The non-linear impact of the phantom read risk is reflected by the square of this coefficient, indicating that when the risk increases, a more significant adjustment is required. The number of phantom read detections, in units of times. It is used to evaluate the frequency of phantom reads occurring in a transaction. The total number of transactions, in units of times. It provides a benchmark to measure the proportion of phantom read events in the overall transactions.
[0038] η^2 is used to amplify the impact of the phantom read risk. The square relationship means that an increase in the risk coefficient will lead to a larger adjustment of the connection pool, reflecting a sensitive response to the medium-risk state; multiplying the ratio of P and T by 0.2 is used to quantify the frequency of phantom read events. By combining the ratio of the number of phantom reads to the total number of transactions, a relative indicator is provided to reflect the stability of the system when processing transactions; subtract represents the unused connection space in the current system. Through this difference, it is ensured that the adjustment will not exceed the bearing capacity of the system, and at the same time, the unused resources are utilized to improve performance.
[0039] In the above calculation formula, T is the total number of transactions within a specific time period; When a medium-risk yellow warning occurs, the predicted cache amount of the database and the adjustment amount of the corresponding connection pool will be fed back to the background terminal of the data management system to make corresponding adjustments to the database.
[0040] The process of comprehensively adjusting the cache amount and the connection pool of the database when the warning level is identified as a high-risk red warning is as follows: Calculate respectively the predicted cache amount of the database and the adjustment amount of the corresponding connection pool ; ; The design logic of the above calculation formula is based on the following key ideas: A high-risk warning means that the system faces a large phantom read pressure. This pressure is quantified by η, and in the formula, This term reflects the impact of the risk factor on cache requirements. The characteristics of the exponential function make the adjustment amplitude increase significantly when η is high, reflecting the need to quickly increase the cache to cope with the pressure in high-risk situations.
[0041] The difference between the current cache capacity and the maximum capacity reflects the system's potential for expansion. When the risk factor indicates that more cache is needed, the formula ensures that the maximum value allowed by the hardware is not exceeded to prevent over-allocation and waste of resources.
[0042] The term k multiplied by △D takes into account the direct impact of data changes on cache requirements. Data changes △D refers to the number of data rows inserted or deleted during the execution of a transaction, in units of rows. More data changes means more data needs to be accessed quickly, so it is reasonable to increase the cache to improve access efficiency. The coefficient k can be determined based on the average size of each row of data in the system.
[0043] ; The above formula is used to dynamically adjust the size of the database connection pool in high-risk situations. High risk usually means that the system faces a large load and potential performance bottlenecks, so the connection pool needs to be flexibly adjusted to improve concurrent processing capabilities.
[0044] The denominator in the fraction in the formula is used to dynamically reflect the impact of the risk level on the connection pool requirements. The higher the risk factor, the greater the read-write conflicts and inconsistencies faced by the system, and more connections are needed to distribute the load.
[0045] The second term in the formula reflects an inverse proportional control mechanism. As the risk factor and the current number of connections increase, the adjustment amount gradually increases, but due to the inverse proportionality, the growth rate of the adjustment gradually slows down. This design ensures that in high-risk situations, the adjustment is robust and does not lead to over-allocation of resources. By using an inverse proportional function, the formula achieves control over the adjustment range. Even in high-risk situations, the adjustment is gradual, avoiding the impact of drastic fluctuations on system stability.
[0046] In the above calculation formula, exp(·) represents an exponential function with the real number e as the base, △D represents the amount of data change in a specific time period, and k is a set conversion coefficient used to convert the unit of the data change amount into the unit of the cache amount; When a high-risk red alert is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
[0047] Example 2 See also Figure 2As shown, a data management device includes: a data acquisition module, a phantom reading classification module, a phantom reading adjustment module and a data management system background terminal, each module is connected by wired and / or wireless connection to achieve data transmission between each module; Data collection module: Deploy monitoring agents at the database engine layer in the data management system to collect data phantom reading information in real time within a specific time period; Phantom reading classification module: calculates the data phantom reading probability coefficient in the database within a specific time period, and performs graded warning operations on the data based on the set judgment rules and the data phantom reading probability coefficient; Phantom read adjustment module: Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are adjusted comprehensively; Data management system backend terminal: used to receive the cache amount of the database and the adjustment amount of the connection pool corresponding to each warning level, and perform corresponding adjustment operations on the database based on the warning level within a specific time period.
[0048] Example 3 A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of a data governance method as described in the present invention when executing the computer-readable instructions.
[0049] Example 4 A computer-readable storage medium having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by a processor, implement the steps of a data governance method as described in the present invention.
[0050] It should be understood that in various embodiments of the present application, the adjustment amount of the serial number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0051] It should be understood that determining B based on A does not mean determining B only based on A. B can also be determined based on A and / or other information.
[0052] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0053] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data governance method, characterized in that: include: S1. Deploy monitoring agents at the database engine layer in the data management system to collect data phantom read information in real time within a specific time period; S2. Calculate the data phantom reading possibility coefficient η in the database within a specific time period, and perform graded warning operations on the data based on the set judgment rules and the data phantom reading possibility coefficient; S3. Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are comprehensively adjusted.
2. A data management method according to claim 1, characterized in that: The data phantom read information in the specific time period in step S1 includes the number of phantom read detections P, the total number of transactions T, the data change amount ΔD, the number of transaction conflicts C, the number of active connections N, and the time point of occurrence of each phantom read event.
3. A data management method according to claim 2, characterized in that: Calculating the data phantom read possibility coefficient η in a database within a specific time period includes the following steps: Subtract the time point of the first phantom read event from the time point of the last phantom read event in a specific time period to obtain the total phantom read observation time window in the specific time period. ; This calculates the probability coefficient of data phantom reading in the database within a specific time period ; Where α1, α2 and α3 represent the set data change weight factor, transaction conflict weight factor and time sensitivity factor respectively, and satisfy α1+α2+α3=1; The standard deviation of the time interval between two consecutive phantom read events.
4. A data governance method according to claim 1, characterized in that: Based on the data phantom read possibility coefficient, a hierarchical warning operation is performed on the data, including the following steps: The three-level warning mechanism is triggered based on the probability coefficient η of data phantom reading in the database within a specific time period: If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a low-risk blue warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a medium-risk yellow warning; If the data phantom read probability coefficient η in the database within a specific time period satisfies , it is marked as a high-risk red warning; Where τ is a set judgment constant, which is used to prevent it from being zero; They represent the set lower and upper thresholds for judgment respectively. The specific acquisition process is as follows: The lower limit benchmark threshold and the upper limit benchmark threshold are determined by historical data regression analysis and are denoted as ; The set lower and upper thresholds are calculated. ; L is the system load factor of the digital management system corresponding to a specific time period, and its specific calculation formula is: 。 5. A data governance method according to claim 1, characterized in that: If the corresponding warning level is identified as a low-risk blue warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Using the calculation formula , calculate the estimated cache size of the database when the low-risk blue warning occurs ; is the maximum cache size of the database, L is the system load factor of the data management system corresponding to a specific time period, is the current cache size of the database; e is a natural constant, and t is the number of consecutive risk warning cycles; According to the calculation formula , calculate the adjustment amount of the database corresponding connection pool when the low risk blue warning is in effect , They represent the maximum number of connections supported by the connection pool of the database and the current number of active connections in the connection pool corresponding to the database, respectively. P represents the number of phantom read detections within a specific time period. When a low-risk blue warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the backend terminal of the data management system to make corresponding adjustments to the database.
6. A data governance method according to claim 5, characterized in that: If the corresponding warning level is identified as a medium-risk yellow warning, the process of comprehensively adjusting the database cache amount and the connection pool adjustment amount is as follows: Calculate the estimated cache size of the database when the medium-risk yellow warning is in effect And the corresponding connection pool adjustment ; ; ; In the above calculation formula, T is the total number of transactions in a specific time period; When a medium-risk yellow warning is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the backend terminal of the data management system to make corresponding adjustments to the database.
7. A data governance method according to claim 6, characterized in that: Identify the warning level as a high-risk red warning, and make comprehensive adjustments to the database cache and connection pool adjustments as follows: Calculate the estimated cache size of the database when a high-risk red alert occurs And the corresponding connection pool adjustment ; ; ; In the above calculation formula, exp(·) represents an exponential function with the real number e as the base, △D represents the amount of data change in a specific time period, and k is a set conversion coefficient used to convert the unit of the data change amount into the unit of the cache amount; When a high-risk red alert is issued, the estimated cache volume of the database and the corresponding connection pool adjustment amount will be fed back to the back-end terminal of the data management system to make corresponding adjustments to the database.
8. A data management device, characterized in that: It is implemented based on a data governance method according to any one of claims 1 to 7, including: Data collection module: Deploy monitoring agents at the database engine layer in the data management system to collect data phantom reading information in real time within a specific time period; Phantom reading classification module: calculates the data phantom reading probability coefficient in the database within a specific time period, and performs graded warning operations on the data based on the set judgment rules and the data phantom reading probability coefficient; Phantom read adjustment module: Based on the corresponding warning level, the cache amount of the database and the adjustment amount of the connection pool are adjusted comprehensively; Data management system backend terminal: used to receive the cache amount of the database and the adjustment amount of the connection pool corresponding to each warning level, and perform corresponding adjustment operations on the database based on the warning level within a specific time period.
9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of a data governance method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of a data governance method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Database read-write control method and device and computer readable storage medium
CN118861080A
Construction method and system for membrane material database
CN119248749A
High land tourism safety risk warning method based on reinforcement learning
US20230072985A1