Database partition creation method and device, electronic equipment and storage medium

Through time series analysis and model updates, we predict the future data volume of database business and create partitions in advance, solving the problem of frequent database partition creation inefficiently and achieving more efficient and stable database operation.

CN120386822APending Publication Date: 2025-07-29PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584738.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the database frequently triggers partition creation during data insertion operation, resulting in low partition creation efficiency and affecting the database operation stability and performance.

Method used

By obtaining the sample business data volume of the target database, using time series analysis and model updates, predicting the business data volume in the future period, creating database partitions in advance, and avoiding frequent partition creation operations.

Benefits of technology

Improves the efficiency of database partition creation, reduces performance jitter and data insertion latency, and ensures the stability and performance consistency of database operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386822A_ABST
    Figure CN120386822A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a database partition creation method and device, electronic equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps of obtaining a sample service data volume of a target database, the sample service data volume comprising a first service data volume of the target database in a first historical period and a second service data volume of the target database in a second historical period, and performing service volume prediction on the first service data volume through an original service model, obtaining a predicted business data volume of the target database in a second historical time period, calculating a model business residual error of the original business model according to the second business data volume and the predicted business data volume, and performing model updating on the original business model according to the model business residual error to obtain a target business model, and the business volume prediction is performed on the preset business data volume through the target business model to obtain the target business data volume, and partition creation is performed on the target database according to the target business data volume, so that the database partition creation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and is applied to the fields of fintech and digital healthcare. In particular, it relates to a method, apparatus, electronic device, and storage medium for creating database partitions. Background Art

[0002] With the development of business, the volume of business data will continuously increase. To enable a database to process large-scale business data, it is necessary to create database partitions. For example, in a fintech scenario, a banking business database is used to manage and store a large amount of banking business information such as personal information, account details, and transaction histories of a large number of users. As the number of customers increases, the volume of banking business information will also increase. To enable the banking business database to have sufficient capacity to store banking business information, it is necessary to create partitions for the banking business database. Another example is in a digital healthcare scenario, where a digital healthcare database is used to manage and store a large amount of medical information such as electronic medical records and health test results of patients. As the number of patients increases and the treatment status of patients is updated, the volume of medical information will increase significantly. By creating partitions for the digital healthcare database, the ability of the digital healthcare database to process large-scale medical information can be enhanced.

[0003] In related technologies, a database creates partitions using an on-demand creation strategy. When the database performs a data insertion operation, if it detects that the current partition cannot accommodate the new data, a partition creation operation will be triggered. A database is a data-intensive system that needs to perform a large number of data insertion operations on a large amount of data, so that each insertion operation may trigger a partition creation operation. And partition creation involves a series of complex operations such as metadata update, storage allocation, and index adjustment, resulting in low efficiency of database partition creation. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method, apparatus, electronic device, and storage medium for creating database partitions, aiming to improve the efficiency of database partition creation.

[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a method for creating database partitions, the method including:

[0006] Obtain the sample business data volume of a target database; wherein, the sample business data volume is an ordered sequence of stationary non-white noise, and the sample business data volume includes the first business data volume of the target database in a first historical period and the second business data volume in a second historical period, and the second historical period is after the first historical period;

[0007] Perform business volume prediction on the first business data volume through a pre-obtained original business model to obtain the predicted business data volume of the target database in the second historical period;

[0008] Calculate the model business residual of the original business model according to the second business data volume and the predicted business data volume;

[0009] Update the original business model according to the model business residual to obtain a target business model, and predict the business volume of a preset business data volume through the target business model to obtain the target business data volume of the target database in a future period;

[0010] Create partitions for the target database according to the target business data volume.

[0011] In some embodiments, obtaining the sample business data volume of the target database includes:

[0012] Obtain the original business data volume of the target database;

[0013] Perform a sequence conversion on the original business data volume to obtain a candidate business data volume; wherein, the candidate business data volume is a stationary time series;

[0014] Perform a white noise detection on the candidate business data volume to obtain a first noise detection result;

[0015] If the first noise detection result indicates that the candidate business data volume is non-white noise, then use the candidate business data volume as the sample business data volume.

[0016] In some embodiments, before predicting the business volume of the first business data volume through a pre-obtained original business model to obtain the predicted business data volume of the target database in the second historical period, the method further includes:

[0017] Perform a graph detection on the sample business data volume to obtain an autocorrelation coefficient graph and a partial autocorrelation coefficient graph; wherein, the autocorrelation coefficient graph is used to indicate a first coefficient form of the autocorrelation coefficient, and the partial autocorrelation coefficient graph is used to indicate a second coefficient form of the partial autocorrelation coefficient;

[0018] Screen a preset business model according to the first coefficient form and the second coefficient form to obtain the original business model.

[0019] In some embodiments, the preset business model includes an autoregressive model, a moving average model, and an autoregressive integrated moving average model. The screening of the preset business model according to the first coefficient form and the second coefficient form to obtain the original business model includes:

[0020] If the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing pattern and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a truncated pattern, then select the autoregressive model from the preset service models to obtain the original service model;

[0021] If the first coefficient pattern indicates that the autocorrelation coefficient is in a truncated pattern and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing pattern, then select the moving average model from the preset service models to obtain the original service model;

[0022] If the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing pattern and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing pattern, then select the autoregressive integrated moving average model from the preset service models to obtain the original service model.

[0023] In some embodiments, the step of selecting the autoregressive integrated moving average model from the preset service models to obtain the original service model includes:

[0024] Select the autoregressive integrated moving average model from the preset service models;

[0025] Initialize the model parameters of the autoregressive integrated moving average model to obtain at least two sub-service models; wherein, the sub-service model has an autoregressive order and a moving average order;

[0026] Calculate the Bayesian score of the sub-service model according to the autoregressive order and the moving average order;

[0027] Select the sub-service model with the smallest Bayesian score as the original service model.

[0028] In some embodiments, the step of updating the original service model according to the model service residuals to obtain the target service model includes:

[0029] Perform white noise detection on the model service residuals to obtain a second noise detection result;

[0030] If the second noise detection result indicates that the model service residuals are non-white noise, then update the original service model to obtain the target service model.

[0031] In some embodiments, the step of creating partitions for the target database according to the target service data volume includes:

[0032] Obtain the partition capacity;

[0033] Calculate the number of partitions according to the target service data volume and the partition capacity;

[0034] Partition the target database according to the number of partitions described above.

[0035] To achieve the above object, a second aspect of the embodiments of the present application provides a database partition creation device, which includes:

[0036] A data volume acquisition module, configured to acquire the sample service data volume of the target database; wherein, the sample service data volume is an ordered sequence of stationary non-white noise, and the sample service data volume includes the first service data volume of the target database in the first historical period and the second service data volume in the second historical period, and the second historical period is after the first historical period;

[0037] A data volume prediction module, configured to perform service volume prediction on the first service data volume through a pre-acquired original service model to obtain the predicted service data volume of the target database in the second historical period;

[0038] A calculation module, configured to calculate the model service residual of the original service model according to the second service data volume and the predicted service data volume;

[0039] A model update module, configured to update the original service model according to the model service residual to obtain a target service model, and perform service volume prediction on a preset service data volume through the target service model to obtain the target service data volume of the target database in a future period;

[0040] A partition creation module, configured to partition the target database according to the target service data volume.

[0041] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0042] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0043] The database partition creation method, database partition creation device, electronic device, and computer-readable storage medium proposed in the embodiments of the present application obtain the sample service data volume of the target database to determine the model parameters of the service model based on the sample service data volume. To improve the accuracy of service volume prediction, the sample service data volume is an ordered sequence of stationary non-white noise, so as to capture the implicit data growth pattern in the ordered sequence based on the service model. The sample service data volume includes the first service data volume of the target database in the first historical period and the second service data volume in the second historical period. The second historical period is after the first historical period, and there is a correlation between the service volume in the second historical period and the service volume in the first historical period. The service volume of the target database in the second historical period is predicted by using the pre-obtained original service model for the first service data volume. According to the second service data volume and the predicted service data volume, the model service residual of the original service model is calculated to determine the difference between the service volume prediction value and the service volume true value based on the model service residual, thereby evaluating the service volume prediction performance of the service model. To obtain a service model with better prediction performance, the original service model is updated according to the model service residual to optimize the service model and obtain the target service model. The target service model is used to predict the service volume of the preset service data volume to obtain the target service data volume of the target database in the future period. The target database is partitioned based on the target service data volume to create database partitions in advance, avoiding triggering the partition creation operation for each insert operation when the partition capacity is insufficient, and improving the low efficiency of database partition creation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flowchart of the database partition creation method provided by the embodiments of the present application;

[0045] Figure 2 is Figure 1 a flowchart of step S110 in

[0046] Figure 3 is another flowchart of the database partition creation method provided by the embodiments of the present application;

[0047] Figure 4 is Figure 3 a flowchart of step S320 in

[0048] Figure 5 is Figure 4 a flowchart of step S430 in

[0049] Figure 6 is Figure 1 a flowchart of step S140 in

[0050] Figure 7 isFigure 1 Flowchart of step S150 in

[0051] Figure 8 Schematic structural diagram of a database partition creation device provided by an embodiment of the present application;

[0052] Figure 9 Schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] It should be noted that although functional module division is performed in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different sequence in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific sequence or order.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0056] With the development of business, the volume of business data will continuously increase. To enable a database to process large-scale business data, it is necessary to create partitions for the database. For example, in the fintech scenario, a banking business database is used to manage and store a large amount of banking business information such as personal information, account details, and transaction histories of a large number of users. As the number of customers increases, the volume of banking business information will increase accordingly. To enable the banking business database to have sufficient capacity to store banking business information, it is necessary to create partitions for the banking business database. Another example is in the digital healthcare scenario, a digital healthcare database is used to manage and store a large amount of medical information such as electronic medical records and health test results of patients. As the number of patients increases and the patient diagnosis and treatment status is updated, the volume of medical information will increase significantly. By creating partitions for the digital healthcare database, the ability of the digital healthcare database to process large-scale medical information can be enhanced.

[0057] In the related art, the database creates partitions using a on-demand creation strategy. When the database performs a data insertion operation, if it detects that the current partition cannot accommodate new data, a partition creation operation is triggered. The database is a data-intensive system, and it needs to perform a large number of data insertion operations on a large amount of data, so that each insertion operation may trigger a partition creation operation. However, partition creation involves a series of complex operations such as metadata update, storage allocation, and index adjustment, resulting in low efficiency of database partition creation.

[0058] Based on this, the embodiments of the present application provide a database partition creation method, a database partition creation device, an electronic device, and a computer-readable storage medium, aiming to improve the efficiency of database partition creation.

[0059] The database partition creation method, database partition creation device, electronic device, and computer-readable storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the database partition creation method in the embodiments of the present application is described.

[0060] The database partition creation method provided by the embodiments of the present application relates to the field of computer technology. The database partition creation method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the database partition creation method, etc., but is not limited to the above forms.

[0061] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0062] Figure 1 is an alternative flowchart of the database partition creation method provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S110 to S150.

[0063] Step S110, obtain the sample business data volume of the target database; wherein, the sample business data volume is an ordered sequence of stationary non-white noise, and the sample business data volume includes the first business data volume of the target database in the first historical period and the second business data volume in the second historical period, and the second historical period is after the first historical period;

[0064] Step S120, perform business volume prediction on the first business data volume through a pre-obtained original business model to obtain the predicted business data volume of the target database in the second historical period;

[0065] Step S130, calculate the model business residual of the original business model according to the second business data volume and the predicted business data volume;

[0066] Step S140, update the original business model according to the model business residual to obtain the target business model, and perform business volume prediction on the preset business data volume through the target business model to obtain the target business data volume of the target database in the future period;

[0067] Step S150, create partitions for the target database according to the target business data volume.

[0068] Steps S110 to S150 shown in the embodiments of the present application perform time series analysis on the business data volume in the historical period, predict the target business data volume in the future period, and create database partitions in advance according to the target business data volume, avoiding triggering the partition creation operation every time the actual insertion operation is executed when the partition capacity is insufficient, and improving the efficiency of database partition creation.

[0069] In the related art, the database creates partitions using an on-demand creation strategy. When the database executes a data insertion operation, if it detects that the current partition cannot accommodate the data to be inserted, a partition creation operation will be triggered. Although this strategy can ensure the dynamic scalability of the partition, for a data-intensive system, in the actual operation process of the database, this strategy will affect the running stability of the database. For example, when the database needs to insert a large amount of data, the on-demand creation strategy may cause a partition creation operation to be triggered for each insertion operation. The partition creation involves complex operations such as metadata update, storage allocation, and index adjustment, resulting in low efficiency of partition creation and a significant increase in the total duration of data insertion. In a high-concurrency scenario, this performance jitter will have a serious impact on the stability of the database. For example, the performance jitter will cause the response time of the system to become unpredictable. Users may observe a sudden increase in the delay when the database processes data insertion, which will not only affect the user experience but also may have a chain reaction on the business logic that depends on the database. In addition, frequent partition creation operations will cause additional load pressure on the database, especially in an environment with limited storage resources or low disk I / O performance, and this impact will be more obvious. To improve the running stability of the database and the efficiency of partition creation, the embodiments of the present application perform time series analysis on the sample business data volume, predict the target business data volume in the future period, and create database partitions according to the target business data volume.

[0070] Please refer to Figure 2 , in some embodiments, step S110 may include but is not limited to steps S210 to S240:

[0071] Step S210, obtaining the original business data volume of the target database;

[0072] Step S220, performing sequence conversion on the original business data volume to obtain candidate business data volume; wherein, the candidate business data volume is a stationary time series;

[0073] Step S230, performing white noise detection on the candidate business data volume to obtain the first noise detection result;

[0074] Step S240, if the first noise detection result indicates that the candidate business data volume is non-white noise, then using the candidate business data volume as the sample business data volume.

[0075] In step S210 of some embodiments, the original business data volume of the target database is obtained. The target database is a database with a need for partition creation. The original business data volume is an ordered sequence of business data volumes arranged in chronological order, including the first business data volume of the target database in the first historical period and the second business data volume in the second historical period. The second historical period is after the first historical period. Both the first business data volume and the second business data volume are the scales of business data generated by the target database when performing business operations, such as the data volumes of transaction records, customer account information, etc. The data volume size can be measured in bytes (Byte) as the basic unit.

[0076] In step S220 of some embodiments, a scatter plot of the original business data volume is obtained, and the original business data volume is preliminarily detected for stationarity based on the scatter plot. The scatter plot visually shows the changing trend of the business data volume over time. If the business data volume fluctuates around a specific fixed level, and there is no obvious upward or downward trend and no periodic changes, the original business data volume may be a stationary time series. To improve the accuracy of the stationarity detection, the original business data volume is re-detected for stationarity by the unit root test method. If the original business data volume has a unit root, the original business data volume is a non-stationary time series. If the original business data volume does not have a unit root, the original business data volume is a stationary time series.

[0077] If the original business data volume is a non-stationary time series, then a difference calculation is performed on the original business data volume to convert the original business data volume from a non-stationary time series to a stationary time series, obtaining a candidate business data volume, and the order of the difference calculation is recorded to obtain the difference order. The difference order is the minimum number of difference times required to convert a non-stationary time series to a stationary time series.

[0078] In step S230 of some embodiments, in order to ensure that there is a predictable pattern in the candidate business data volume, making modeling and prediction possible, a white noise detection is performed on the candidate business data volume to obtain a first noise detection result. The first noise detection result is used to indicate whether the candidate business data volume is white noise or non-white noise. Specifically, the mean and variance of the candidate business data volume are calculated. If the mean is 0 and the variance is a constant, the candidate business data volume is white noise. If the mean is not 0 or the variance is not a constant, the candidate business data volume is non-white noise.

[0079] In step S240 of some embodiments, white noise is a completely random signal with no correlation between the values at each time point. If the candidate service data volume is white noise, then the candidate service data volume has no predictable pattern, and no useful information can be extracted from it for prediction. If the first noise detection result indicates that the candidate service data volume is non-white noise, then the candidate service data volume is used as the sample service data volume. The sample service data volume is an ordered sequence of stationary non-white noise, and the sample service data volume includes the first service data volume of the target database in the first historical period and the second service data volume in the second historical period, where the second historical period is after the first historical period. The first service data volume of the sample service data volume can be obtained by performing a difference operation on the first service data volume of the original service data volume. The second service data volume of the sample service data volume can be obtained by performing a difference operation on the second service data volume of the original service data volume. The first historical period and the second historical period can be collectively referred to as the historical period, and the first service data volume and the second service data volume can be collectively referred to as the service data volume. If the first noise detection result indicates that the candidate service data volume is white noise, then service volume prediction cannot be performed based on the candidate service data volume.

[0080] Through the above steps S210 to S240, a sample service data volume of stationary non-white noise can be obtained to capture the data growth pattern based on the sample service data volume.

[0081] Please refer to Figure 3 , in some embodiments, before step S120, the database partitioning creation method may further include, but is not limited to, steps S310 to S320:

[0082] Step S310, perform a graph detection on the sample service data volume to obtain an autocorrelation coefficient graph and a partial autocorrelation coefficient graph; wherein, the autocorrelation coefficient graph is used to indicate the first coefficient form of the autocorrelation coefficient, and the partial autocorrelation coefficient graph is used to indicate the second coefficient form of the partial autocorrelation coefficient;

[0083] Step S320, perform a model screening on the preset service model according to the first coefficient form and the second coefficient form to obtain the original service model.

[0084] In step S310 of some embodiments, the lag order is obtained. The lag order is the number of time period units delayed backward from the current time period. If the lag order is k and the current time period is t, the lag order can delay the current time period t backward by k time period units. Calculate the autocorrelation coefficient of the sample service data volume according to the lag order, and calculate the partial autocorrelation coefficient according to the lag order and the autocorrelation coefficient. The autocorrelation coefficient is used to reflect the correlation between the service data volume in the current historical time period t and the service data volume k time period units lagged in the sample service data volume. The partial autocorrelation coefficient is used to reflect the correlation between the service data volume in the current historical time period t and the service data volume k time period units lagged after controlling the service data volume in the intermediate historical time periods. The calculation formula of the autocorrelation coefficient is defined as:

[0085]

[0086] where k is the lag order; cov represents covariance; X t is the service data volume in the historical time period t; X t+k is the service data volume of the current historical time period t lagged by k historical time period units; σ 2 is the variance of the sample service data volume.

[0087] The calculation formula of the partial autocorrelation coefficient is defined as:

[0088]

[0089] where ρ k is the autocorrelation coefficient with a lag order of k; j is an index variable with a value in the range of [1, k - 1]; is the j-th coefficient of the partial autocorrelation coefficient with a lag order of k - 1.

[0090] Take the lag order as the abscissa and the autocorrelation coefficient as the ordinate to generate an autocorrelation coefficient graph. Take the lag order as the abscissa and the partial autocorrelation coefficient as the ordinate to generate a partial autocorrelation coefficient graph. The autocorrelation coefficient graph is used to indicate the first coefficient form of the autocorrelation coefficient changing with the lag order, and the partial autocorrelation coefficient graph is used to indicate the second coefficient form of the partial autocorrelation coefficient changing with the lag order.

[0091] In step S320 of some embodiments, the preset business models include an Autoregressive model (AR), a Moving Average model (MA), and an Autoregressive Integrated Moving Average model (ARIMA). An original business model is selected from the autoregressive model, the moving average model, and the autoregressive integrated moving average model according to the first coefficient pattern and the second coefficient pattern.

[0092] Through the above steps S310 to S320, a business model suitable for processing the sample business data volume can be selected to accurately capture the data growth pattern hidden in the sample business data volume based on the business model, thereby improving the accuracy of business volume prediction.

[0093] Please refer to Figure 4 , in some embodiments, step S320 may include but is not limited to step S410, step S420, or step S430:

[0094] Step S410, if the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing form and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a truncated form, then select the autoregressive model from the preset business models to obtain the original business model;

[0095] Step S420, if the first coefficient pattern indicates that the autocorrelation coefficient is in a truncated form and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing form, then select the moving average model from the preset business models to obtain the original business model;

[0096] Step S430, if the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing form and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing form, then select the autoregressive integrated moving average model from the preset business models to obtain the original business model.

[0097] In step S410 of some embodiments, if the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing form, that is, the autocorrelation coefficient gradually decays as the lag order increases but does not rapidly approach 0, and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a truncated form, that is, the partial autocorrelation coefficient suddenly becomes 0 at a specific lag order, then select the autoregressive model from the preset business models and use the specific lag order as the autoregressive order of the autoregressive model to obtain the original business model.

[0098] In step S420 of some embodiments, if the first coefficient pattern indicates that the autocorrelation coefficient is in a truncated form, it means that the autocorrelation coefficient suddenly becomes 0 at a specific lag order, and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing form, it means that the partial autocorrelation coefficient gradually decays as the lag order increases. Then, a moving average model is selected from the preset service models, and this specific lag order is used as the moving average order of the moving average model to obtain the original service model.

[0099] In step S430 of some embodiments, if the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing form and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing form, it means that both the autocorrelation coefficient and the partial autocorrelation coefficient gradually decay as the lag order increases. Then, an autoregressive integrated moving average model is selected from the preset service models, and the autoregressive order and the moving average order of the autoregressive integrated moving average model are estimated to obtain the original service model.

[0100] The above steps S410 to S430 can quickly select a service model suitable for processing this time series of sample service data volume by using the graphical features of the autocorrelation coefficient graph and the partial autocorrelation coefficient graph, improving the efficiency of model selection.

[0101] Please refer to Figure 5 , in some embodiments, step S430 may include but is not limited to steps S510 to S540:

[0102] Step S510, select an autoregressive integrated moving average model from the preset service models;

[0103] Step S520, initialize the model parameters of the autoregressive integrated moving average model to obtain at least two sub-service models; among them, the sub-service model has an autoregressive order and a moving average order;

[0104] Step S530, calculate the Bayesian score of the sub-service model according to the autoregressive order and the moving average order;

[0105] Step S540, select the sub-service model with the smallest Bayesian score as the original service model.

[0106] In step S510 of some embodiments, if the first coefficient pattern indicates that the autocorrelation coefficient is in a trailing form and the second coefficient pattern indicates that the partial autocorrelation coefficient is in a trailing form, then select an autoregressive integrated moving average model from the preset service models.

[0107] In step S520 of some embodiments, multiple possible values of the autoregressive order of the autoregressive integrated moving average model are obtained to get multiple first values, multiple possible values of the moving average order of the autoregressive integrated moving average model are obtained to get multiple second values, and the differencing order obtained by differencing calculation is acquired. The combination of the first value, the second value, and the differencing order is the model parameter. According to multiple combinations constructed by one first value, one second value, and the differencing order, the model parameters of the autoregressive integrated moving average model are initialized, and multiple sub-business models can be obtained. The first value is the autoregressive order of the sub-business model, and the second value is the moving average order of the sub-business model. If the first value is p, the second value is q, and the differencing order is d, then the sub-business model is ARIMA(p, q, d).

[0108] In step S530 of some embodiments, the sample business data volume is fitted by the sub-business model to obtain the estimated autoregressive coefficient, the estimated moving average coefficient, and the estimated constant term of the sub-business model. The likelihood function value is calculated based on the estimated autoregressive coefficient, the estimated moving average coefficient, the estimated constant term, and the sample business data volume. The above steps are repeated until the likelihood function value reaches the maximum. According to the maximum likelihood function value, the autoregressive order, the moving average order, the differencing order, and the number of samples in the sample business data volume, the Bayesian score of the sub-business model is calculated. The calculation formula of the Bayesian score is defined as:

[0109] BIC = -2Ln(L) + kLn(n),

[0110] where BIC is the Bayesian score, L is the maximum likelihood function value; k is the number of model parameters, expressed as k = p + q + d, where p, q, and d are the autoregressive order, the moving average order, and the differencing order of the sub-business model respectively; n is the number of samples; Ln represents the natural logarithm operation.

[0111] In step S540 of some embodiments, the Bayesian scores of the sub-business models are compared, and the sub-business model with the smallest Bayesian score is selected as the original business model. The smaller the Bayesian score, the better the model fitting goodness of the sub-business model, and at the same time, the complexity of the model is also controlled.

[0112] Through the above steps S510 to S540, the autoregressive integrated moving average model with the best fitting degree is selected as the original business model, thereby improving the business volume prediction performance of the original business model.

[0113] In step S120 of some embodiments, the original business model is used to predict the business volume of the first business data volume of the target database in the first historical period, and the predicted business data volume of the target database in the second historical period is obtained.

[0114] In step S130 of some embodiments, calculate the difference between the second business data volume of the target database in the second historical period and the predicted business data volume in the second historical period to obtain the model business residual of the original business model.

[0115] Please refer to Figure 6 , in some embodiments, step S140 may include but is not limited to steps S610 to S620:

[0116] Step S610, perform white noise detection on the model business residual to obtain a second noise detection result;

[0117] Step S620, if the second noise detection result indicates that the model business residual is non-white noise, update the original business model to obtain a target business model.

[0118] In step S610 of some embodiments, in order to judge the model performance of the original business model, perform white noise detection on the model business residual to obtain a second noise detection result. The second noise detection result is used to indicate that the model business residual is white noise or non-white noise.

[0119] In step S620 of some embodiments, if the second noise detection result indicates that the model business residual is non-white noise, it means that the original business model cannot accurately predict the business volume, then update the model parameters of the original business model such as the autoregressive order, moving average order, autoregressive coefficient, moving average coefficient, and constant term, etc., to obtain a target business model. If the second noise detection result indicates that the model business residual is white noise, it means that the original business model can accurately learn the data growth pattern from the ordered time series, so as to accurately predict the business volume, then use the original business model as the target business model.

[0120] The preset business data volume is a stationary non-white noise ordered sequence, and the preset business data volume includes the business data volume of the target historical period. The target historical period can be the first historical period, the second historical period, or other historical periods. Use the target business model to predict the business volume of the preset business data volume, determine the business data volume of the target database in the future period, and obtain the target business data volume.

[0121] Through the above steps S610 to S620, the business model can be optimized to improve the business volume prediction performance of the business model.

[0122] Please refer to Figure 7 , in some embodiments, step S150 may include but is not limited to steps S710 to S730:

[0123] Step S710, obtain the partition capacity;

[0124] Step S720: Calculate the number of partitions based on the target business data volume and the partition capacity.

[0125] Step S730: Create partitions for the target database according to the number of partitions.

[0126] In step S710 of some embodiments, obtain the partition capacity. The partition capacity is the amount of data that a partition can hold, and the partition capacity can be set according to the actual situation, such as 256 GB, where GB represents gigabyte.

[0127] In step S720 of some embodiments, the preset business data volume is an ordered sequence of stationary non-white noise. The preset business data volume includes the business data volume of the target historical period. Obtain the business data volume of the last target historical period in the preset business data volume. Obtain the number of time periods in the future period. If the number of time periods is 1, then calculate the difference between the target business data volume and the business data volume of the last target historical period. If the difference is greater than 0, it means that the difference is the data growth amount, and calculate the ratio between the data growth amount and the partition capacity to obtain the number of partitions. If the number of time periods is greater than 1, then calculate the difference between the target business data volumes of every two adjacent future time periods. Add the absolute values of each difference to obtain the estimated data volume. Calculate the ratio between the estimated data volume and the partition capacity to obtain the number of partitions.

[0128] In step S730 of some embodiments, create partitions for the target database according to the number of partitions to create multiple partitions in the target database in advance. For example, if the number of partitions is 4, then create 4 partitions on the target database.

[0129] Through the above steps S710 to S730, it is possible to create a preset number of partitions in advance, improving the stability of database operation and the efficiency of partition creation.

[0130] The database partition creation method according to the embodiments of the present application predicts the target service data volume in a future period and creates partitions based on the target service data volume. It can establish the partitions of the database in advance, without waiting until the database performs a data insertion operation and discovers that the partition capacity is insufficient to establish partitions, thus avoiding performance jitter of the database. By establishing partitions in advance, the insertion operation only needs to locate the predefined partitions, without dynamically creating partitions, which significantly speeds up the data insertion into the database, avoids the overhead of dynamically creating partitions during data insertion, reduces performance jitter and data insertion latency, and makes the operation of the database system more stable. This static partition strategy does not need to handle operations of dynamically creating partitions and deleting partitions, is easier to manage, reduces maintenance complexity, and eliminates the additional calculations and resource consumption during dynamic partition creation, ensuring performance consistency during the data insertion process and improving the predictability and stability of the system. The database partition creation method according to the embodiments of the present application has significant advantages in optimizing the dynamic partition strategy, reducing performance jitter, improving data insertion and query performance, simplifying management, and improving system stability.

[0131] Please refer to Figure 8 , the embodiments of the present application further provide a database partition creation device, which can implement the above database partition creation method. The database partition creation device includes:

[0132] A data volume acquisition module 810, configured to acquire the sample service data volume of the target database; wherein, the sample service data volume is an ordered sequence of stationary non-white noise, and the sample service data volume includes the first service data volume of the target database in the first historical period and the second service data volume in the second historical period, and the second historical period is after the first historical period;

[0133] A data volume prediction module 820, configured to perform service volume prediction on the first service data volume through a pre-acquired original service model to obtain the predicted service data volume of the target database in the second historical period;

[0134] A calculation module 830, configured to calculate the model service residual of the original service model according to the second service data volume and the predicted service data volume;

[0135] A model update module 840, configured to update the original service model according to the model service residual to obtain a target service model, and perform service volume prediction on a preset service data volume through the target service model to obtain the target service data volume of the target database in a future period;

[0136] A partition creation module 850, configured to create partitions for the target database according to the target service data volume.

[0137] The specific implementation of the database partition creation device is basically the same as the specific embodiments of the above database partition creation method, and will not be elaborated here.

[0138] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above database partition creation method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0139] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0140] A processor 910, which can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0141] A memory 920, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 920 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 920, and the processor 910 is called to execute the database partition creation method of the embodiments of the present application;

[0142] An input / output interface 930, which is used to implement information input and output;

[0143] A communication interface 940, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);

[0144] A bus 950, which transmits information between various components of the device (such as the processor 910, the memory 920, the input / output interface 930, and the communication interface 940);

[0145] Among them, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are communicatively connected to each other inside the device through the bus 950.

[0146] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned database partition creation method is implemented.

[0147] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0148] The database partition creation method, database partition creation device, electronic device, and computer storage medium provided by the embodiments of the present application perform time series analysis on the historical period business data volume, predict the target business data volume in the future period, and create database partitions according to the target business data volume, avoiding triggering the partition creation operation for each insertion operation when the partition capacity is insufficient, and improving the efficiency of database partition creation.

[0149] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0150] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0153] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0154] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0155] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical or other form.

[0156] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0159] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for creating database partitions, characterized in that, The method includes: Obtaining the sample business data volume of the target database; wherein, the sample business data volume is an ordered sequence of stationary non-white noise, and the sample business data volume includes the first business data volume of the target database in the first historical period and the second business data volume in the second historical period, and the second historical period is after the first historical period; Performing business volume prediction on the first business data volume through a pre-acquired original business model to obtain the predicted business data volume of the target database in the second historical period; Calculating the model business residual of the original business model according to the second business data volume and the predicted business data volume; Updating the original business model according to the model business residual to obtain a target business model, and performing business volume prediction on a preset business data volume through the target business model to obtain the target business data volume of the target database in a future period; Creating partitions for the target database according to the target business data volume.

2. The method according to claim 1, characterized in that, The obtaining of the sample business data volume of the target database includes: Obtaining the original business data volume of the target database; Performing sequence conversion on the original business data volume to obtain candidate business data volume; wherein, the candidate business data volume is a stationary time series; Performing white noise detection on the candidate business data volume to obtain a first noise detection result; If the first noise detection result indicates that the candidate business data volume is non-white noise, then using the candidate business data volume as the sample business data volume.

3. The method according to claim 1, characterized in that Before performing business volume prediction on the first business data volume through a pre-acquired original business model to obtain the predicted business data volume of the target database in the second historical period, the method further includes: Performing graph detection on the sample business data volume to obtain an autocorrelation coefficient graph and a partial autocorrelation coefficient graph; wherein, the autocorrelation coefficient graph is used to indicate the first coefficient form of the autocorrelation coefficient, and the partial autocorrelation coefficient graph is used to indicate the second coefficient form of the partial autocorrelation coefficient; Screening a preset business model according to the first coefficient form and the second coefficient form to obtain the original business model.

4. The method according to claim 3, wherein The preset business models include an autoregressive model, a moving average model, and an autoregressive integrated moving average model. The screening of the preset business model according to the first coefficient form and the second coefficient form to obtain the original business model includes: If the first coefficient form indicates that the autocorrelation coefficient is in a trailing form and the second coefficient form indicates that the partial autocorrelation coefficient is in a truncated form, then screening out the autoregressive model from the preset business models to obtain the original business model; If the first coefficient form indicates that the autocorrelation coefficient is in a truncated form and the second coefficient form indicates that the partial autocorrelation coefficient is in a trailing form, then screening out the moving average model from the preset business models to obtain the original business model; If the first coefficient pattern indicates that the autocorrelation coefficient is a trailing pattern and the second coefficient pattern indicates that the partial autocorrelation coefficient is a trailing pattern, then select the autoregressive integrated moving average model from the preset service models to obtain the original service model.

5. The method according to claim 4, wherein The step of selecting the autoregressive integrated moving average model from the preset service models to obtain the original service model includes: Select the autoregressive integrated moving average model from the preset service models; Initialize the model parameters of the autoregressive integrated moving average model to obtain at least two sub-service models; wherein, the sub-service model has an autoregressive order and a moving average order; Calculate the Bayesian score of the sub-service model according to the autoregressive order and the moving average order; Select the sub-service model with the smallest Bayesian score as the original service model.

6. The method according to any one of claims 1 to 5, characterized in that The step of updating the original service model according to the model service residuals to obtain the target service model includes: Perform white noise detection on the model service residuals to obtain a second noise detection result; If the second noise detection result indicates that the model service residuals are non-white noise, then update the original service model to obtain the target service model.

7. The method according to any one of claims 1 to 5, characterized in that, The step of creating partitions for the target database according to the target service data volume includes: Obtain the partition capacity; Calculate the number of partitions according to the target service data volume and the partition capacity; Create partitions for the target database according to the number of partitions.

8. A database partition creation device, characterized in that, The device includes: A data volume acquisition module, configured to acquire a sample service data volume of a target database; wherein, the sample service data volume is an ordered sequence of stationary non-white noise, and the sample service data volume includes a first service data volume of the target database in a first historical period and a second service data volume in a second historical period, and the second historical period is after the first historical period; A data volume prediction module, configured to perform service volume prediction on the first service data volume through a pre-acquired original service model to obtain a predicted service data volume of the target database in the second historical period; A calculation module, configured to calculate the model service residuals of the original service model according to the second service data volume and the predicted service data volume; A model update module, configured to update the original service model according to the model service residuals to obtain a target service model, and perform service volume prediction on a preset service data volume through the target service model to obtain a target service data volume of the target database in a future period; A partition creation module, configured to create partitions for the target database according to the target service data volume.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.