Method, apparatus, electronic device, and storage medium for scattering database data

By performing business logic analysis and partitioning of database data tables, the high cost problem caused by single-point deployment of Oracle databases is solved, distributed storage is realized, storage costs are reduced and efficiency is improved.

CN114297185BActive Publication Date: 2025-07-25WEBANK (CHINA)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111646222.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-25
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, when the loan core system uses Oracle database, single-point deployment leads to a bottleneck in data volume, which cannot be expanded, and is expensive.

Method used

By determining the business logic of the data table in the database, selecting the appropriate breaking strategy, partitioning the data tables and breaking them, and realizing distributed storage.

Benefits of technology

It reduces data storage costs, improves the efficiency and reliability of data storage, and adapts to higher data volume requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114297185B_ABST
    Figure CN114297185B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, electronic device and storage medium for data scattering in a database. The method includes: determining the business logic of a data table in the database; determining the target scattering strategy of the data table in a preset scattering strategy according to the business logic of the data table; storing and partitioning the data in the data table according to the target scattering strategy, and scattering the data of the data table according to the storage partition. The method of the present application realizes distributed storage of the data in the database, and compared with the single-point deployment in the prior art, reduces the cost of data storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to communication technologies, and in particular, to a method, apparatus, electronic device, and storage medium for data scattering in a database. Background Art

[0002] With the development of computer technologies, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech), and database technologies are no exception. However, due to the security and real-time requirements of the financial industry, higher requirements are also imposed on technologies.

[0003] The loan core system needs to process hundreds of millions of data flows every day. Its functions are relatively complex, and the timeliness requirements are strict. Many functions rely on the data of the core system. In the existing loan core system, an Oracle database is usually used. The Oracle database adopts single-point deployment and has high performance. However, the data volume of a single data center node (DCN) has reached a bottleneck, and it is impossible to continue to expand the database of the DCN.

[0004] However, the cost of using an Oracle database server in the existing loan core system is expensive. And for higher data volume storage requirements, it can only be achieved through more expensive servers, resulting in higher costs. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for data scattering in a database, so as to solve the technical problem of high cost of database servers when the data volume for storage is large in the prior art.

[0006] In a first aspect, this application provides a method for data scattering in a database, including:

[0007] Determine the business logic of data tables in the database; determine the target scattering strategy of the data tables according to the business logic of the data tables in the preset scattering strategy; perform storage partitioning on the data in the data tables according to the target scattering strategy, and scatter the data of the data tables according to the storage partitioning.

[0008] In the embodiments of this application, by selecting a suitable scattering strategy as the target scattering strategy from the preset scattering strategies according to the business logic of the data tables in the database, and then performing storage partitioning and scattering on the data in the data tables according to the target scattering strategy, distributed storage of the data in the database is achieved, and compared with single-point deployment in the prior art, the cost of data storage is reduced.

[0009] In a possible implementation manner, for the method for data scattering in a database provided in the embodiments of this application, after scattering the data of the data tables according to the storage partitioning, it further includes:

[0010] Analyze the data distribution of the data table within each partition; if the difference between the amounts of data stored in any two regions exceeds a preset quantity or a preset ratio, update the target scattering strategy by using the data of the data table and the determination method of the target scattering strategy.

[0011] In the embodiment of the present application, after scattering the data of the data table, analyze the data distribution in the storage partition. When the distribution does not meet the uniform expectation, update the target scattering strategy by using the data of the data table, ensuring the reliability of the target scattering strategy.

[0012] In a possible implementation manner, for the method for scattering database data provided by the embodiment of the present application, the preset scattering strategies include a scattering strategy based on the distribution of stock data, a scattering strategy based on the distribution of upstream data, and a scattering strategy based on the distribution of yesterday's data; determining the target scattering strategy of the data table from the preset scattering strategies according to the business logic of the data table includes:

[0013] If the similarity between the data distribution of the data table and the data distribution of the stock data reaches a first preset similarity, determine the scattering strategy based on the distribution of stock data as the target scattering strategy from the preset scattering strategies.

[0014] If the data table has upstream data and there is a proportional relationship between the upstream data and the amount of data in the data table, determine the scattering strategy based on the distribution of upstream data as the target scattering strategy from the preset scattering strategies.

[0015] If the data table has yesterday's data and the similarity between yesterday's data and the data in the data table reaches a second preset similarity, determine the scattering strategy based on the distribution of yesterday's data as the target scattering strategy from the preset scattering strategies.

[0016] In the embodiment of the present application, multiple preset scattering strategies are set, and the appropriate scattering strategy is selected from the multiple preset scattering strategies to scatter the data of the data table according to the business logic of the data table, improving the reliability of data scattering.

[0017] In a possible implementation manner, for the method for scattering database data provided by the embodiment of the present application, the scattering strategy based on the distribution of stock data includes:

[0018] Deduplicate the stock data according to the index to obtain the deduplicated index quantity of the index;

[0019] Determine the ratio of the number of rows of the stock data to the number of deduplicated indexes as the factor multiple of the index; generate a new array using the number of deduplicated indexes and the factor multiple; sort the new array according to the index, and determine the number of indexes and the index range that can be stored in each area using the number of rows of the new array and the preset number of areas.

[0020] In a possible implementation manner, the method for scattering database data provided by the embodiments of the present application, based on the scattering strategy of the upstream data distribution, includes:

[0021] Determine the amount of index data that can be stored in each area according to the amount of index data and the number of areas of the index in the upstream data; determine the number of areas of the data table according to the number of rows of the data table and the amount of index data that can be stored in each area; perform deduplication processing on the upstream data according to the index to generate the number of deduplicated indexes of the index; determine the number of indexes and the index range stored in each area according to the number of deduplicated indexes and the number of areas.

[0022] In a possible implementation manner, the method for scattering database data provided by the embodiments of the present application, based on the scattering strategy of the data distribution of the previous day, includes:

[0023] Perform deduplication processing on the data of the previous day according to the index to obtain the number of deduplicated indexes of the index; determine the number of indexes and the index range that can be stored in each area according to the number of deduplicated indexes and the preset number of areas.

[0024] In a possible implementation manner, before selecting the target scattering strategy of the data table from the preset scattering strategies according to the business logic of the data table in the method for scattering database data provided by the embodiments of the present application, it further includes:

[0025] Determine the primary key type of the data table; if the primary key type is a numeric type, divide the data table into multiple different primary key interval segments in the order of numeric auto-increment.

[0026] Next, the device, electronic device, computer-readable storage medium, and computer program product for scattering database data provided by the embodiments of the present application are introduced. The content and effects can refer to the method for scattering database data provided by the embodiments of the present application and will not be elaborated here.

[0027] In a second aspect, the present application provides a device for scattering database data, including:

[0028] A determination module, configured to determine the business logic of the data table in the database.

[0029] The determination module is further configured to determine the target scattering strategy of the data table from the preset scattering strategies according to the business logic of the data table.

[0030] A processing module, configured to perform storage partitioning on the data in the data table according to a target fragmentation strategy, and fragment the data of the data table according to the storage partitioning.

[0031] In a possible implementation manner, the apparatus for fragmenting database data provided by an embodiment of the present application further includes an update module.

[0032] The update module is configured to: analyze the distribution of the data in the data table within each partition; if the difference between the amounts of data stored in any two regions exceeds a preset quantity or a preset ratio, update the target fragmentation strategy by using the data of the data table and the determination method of the target fragmentation strategy.

[0033] In a possible implementation manner, the preset fragmentation strategies of the apparatus for fragmenting database data provided by an embodiment of the present application include a fragmentation strategy based on the distribution of stock data, a fragmentation strategy based on the distribution of upstream data, and a fragmentation strategy based on the distribution of yesterday's data.

[0034] The determination module is specifically configured to:

[0035] If the similarity between the data distribution of the data table and the data distribution of the stock data reaches a first preset similarity, determine the fragmentation strategy based on the distribution of the stock data in the preset fragmentation strategies as the target fragmentation strategy.

[0036] If there is upstream data in the data table and there is a proportional relationship between the upstream data and the amount of data in the data table, determine the fragmentation strategy based on the distribution of the upstream data in the preset fragmentation strategies as the target fragmentation strategy.

[0037] If there is yesterday's data in the data table and the similarity between yesterday's data and the data in the data table reaches a second preset similarity, determine the fragmentation strategy based on the distribution of yesterday's data in the preset fragmentation strategies as the target fragmentation strategy.

[0038] In a possible implementation manner, the fragmentation strategy based on the distribution of stock data of the apparatus for fragmenting database data provided by an embodiment of the present application includes:

[0039] Perform duplicate removal processing on the stock data according to the index to obtain the number of de-duplicated indexes of the index;

[0040] Determine the ratio of the number of rows of the stock data to the number of de-duplicated indexes as the factor multiple of the index; generate a new array by using the number of de-duplicated indexes and the factor multiple; sort the new array according to the index, and determine the number of indexes and the index range that can be stored in each region by using the number of rows of the new array and the preset number of regions.

[0041] In a possible implementation manner, the fragmentation strategy based on the distribution of upstream data of the apparatus for fragmenting database data provided by an embodiment of the present application includes:

[0042] Determine the amount of index data that can be stored in each region according to the amount of index data and the number of regions in the upstream data; determine the number of regions of the data table according to the number of rows of the data table and the amount of index data that can be stored in each region; perform deduplication processing on the upstream data according to the index to generate the deduplicated index quantity of the index; determine the number of indexes and the index range stored in each region according to the deduplicated index quantity and the number of regions.

[0043] In a possible implementation manner, the device for scattering database data provided by the embodiments of the present application, based on the scattering strategy of yesterday's data distribution, includes:

[0044] Perform deduplication processing on yesterday's data according to the index to obtain the deduplicated index quantity of the index; determine the number of indexes and the index range that can be stored in each region according to the deduplicated index quantity and the preset number of regions.

[0045] In a possible implementation manner, the determining module of the device for scattering database data provided by the embodiments of the present application is further configured to:

[0046] Determine the primary key type of the data table; if the primary key type is a numeric type, divide the data table into multiple different primary key interval segments in the order of numeric auto-increment.

[0047] In a third aspect, the embodiments of the present application provide an electronic device, including:

[0048] A processor, and a memory communicatively connected to the processor;

[0049] The memory stores computer-executable instructions;

[0050] The processor executes the computer-executable instructions stored in the memory to implement the method for scattering database data provided in the first aspect or the implementable manner of the first aspect.

[0051] In a fourth aspect, the embodiments of the present application provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method for scattering database data provided in the first aspect or the implementable manner of the first aspect.

[0052] In a fifth aspect, the embodiments of the present application provide a computer program product, including computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method for scattering database data provided in the first aspect or the implementable manner of the first aspect.

[0053] The method, device, electronic device and storage medium for database data scattering provided by the present application first determine the business logic of the data table in the database; then determine the target scattering strategy of the data table according to the business logic of the data table in the preset scattering strategy; finally, perform storage partitioning on the data in the data table according to the target scattering strategy, and scatter the data of the data table according to the storage partitioning. Since a suitable scattering strategy is selected as the target scattering strategy from the preset scattering strategy according to the business logic of the data table in the database, and then the data in the data table is stored and partitioned and scattered according to the target scattering strategy, the distributed storage of the database data is realized, and compared with the single-point deployment in the prior art, the cost of data storage is reduced. Description of the Drawings

[0054] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0055] Figure 1 is an exemplary application scenario architecture diagram provided by an embodiment of the present application;

[0056] Figure 2 is a schematic flowchart of a method for database data scattering provided by an embodiment of the present application;

[0057] Figure 3 is a schematic flowchart of a method for database data scattering provided by another embodiment of the present application;

[0058] Figure 4 is a schematic flowchart of a scattering strategy based on the distribution of stock data provided by an embodiment of the present application;

[0059] Figure 5 is a schematic flowchart of a scattering strategy based on the distribution of upstream data provided by an embodiment of the present application;

[0060] Figure 6 is a schematic flowchart of a scattering strategy based on the distribution of yesterday's data provided by an embodiment of the present application;

[0061] Figure 7 is a schematic structural diagram of a device for database data scattering provided by an embodiment of the present application;

[0062] Figure 8 is a schematic structural diagram of a device for database data scattering provided by another embodiment of the present application;

[0063] Figure 9 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0064] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by reference to specific embodiments. Detailed Description of the Invention

[0065] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0066] The terms "first" and "second" in the specification, claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0067] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech), and database technology is no exception. However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for technology. The loan core system needs to process hundreds of millions of data streams every day, with complex functions and strict timeliness requirements. In the prior art, the hundreds of millions-level loan core system usually uses an Oracle database. The Oracle database adopts single-point deployment with high performance. The data volume of a single data center node (DCN) has reached the bottleneck, and it is impossible to continue to expand the database of the DCN. For higher data volume storage requirements, it can only be achieved through more expensive servers, resulting in higher costs.

[0068] To solve the above technical problems, the inventive concept of the method, apparatus, electronic device, and storage medium for database data scattering provided by the embodiments of the present application lies in selecting a suitable scattering strategy as the target scattering strategy from the preset scattering strategies according to the business logic of the data tables in the database, and then storing and scattering the data in the data tables according to the target scattering strategy, realizing the distributed storage of the database data. Compared with the single-point deployment adopted in the prior art, the cost of data storage is reduced.

[0069] Hereinafter, an exemplary application scenario of the embodiments of the present application will be introduced.

[0070] The method for database data scattering provided by the embodiments of the present application can be executed by the apparatus for database data scattering provided by the embodiments of the present application. The apparatus for database data scattering provided by the embodiments of the present application can be integrated on a server or a server cluster, or the apparatus for database data scattering can be the server or the server cluster itself. It can also be an application (APP) installed on a terminal device, etc. In a possible implementation manner, the server in the embodiments of the present application can be a server of a platform that generates hot data and needs to process these hot data. The embodiments of the present application do not limit the specific type, quantity, etc. of the server.

[0071] Exemplarily, Figure 1 is an exemplary application scenario architecture diagram provided by the embodiments of the present application. As Figure 1 shown, the architecture mainly includes: a terminal device 11 (personal computer) and a server 12. Data streams are obtained through the web page or client of the terminal device 11. The data streams can be bank loan business data, deposit and withdrawal cash flow data, wealth management product transaction data, etc. The embodiments of the present application do not limit this. The terminal device 11 sends the data streams to the server 12 and stores them in the TIDB database in the server 12. The embodiments of the present application do not limit the specific type of the terminal device 11. For example, the terminal device can be a personal computer as Figure 1 shown, and can also be other terminals such as a tablet computer or a smart phone. The embodiments of the present application are not limited thereto.

[0072] Next, specific embodiments will be used to detail the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These specific embodiments below can be combined with each other. Concepts or processes that are the same or similar may not be repeated in some embodiments. Next, the embodiments of the present application will be described with reference to the accompanying drawings.

[0073] Figure 2FIG. 0 is a schematic flowchart of a method for data dispersion in a database provided by an embodiment of the present application. This method can be executed by a device for data dispersion in a database, and the device can be implemented in software and / or hardware. Hereinafter, the method for data dispersion in a database will be described with a server as the execution entity. As Figure 2 shown, the method for data dispersion in a database provided by an embodiment of the present application may include:

[0074] Step S101: Determine the business logic of the data tables in the database.

[0075] The data in the database is usually stored in multiple data tables. The data stored in the data tables includes the primary key and index of the data table. A record in the data table can be uniquely determined by the primary key, and the data query speed can be improved by the index.

[0076] There may be a certain business logic between the data stored in the data table and the inventory data, upstream data, and yesterday's data. For example, the distribution of the data in the data table is basically the same as the distribution of the inventory data, and the newly added data is less every day; for another example, there is clear upstream data in the data table, and there is a proportional relationship between the data volume of the upstream data and the downstream data; for another example, the data in the data table is similar to yesterday's data, etc. This is only an example in the embodiment of the present application and is not limited thereto. For example, there may be other business logics in the data table, etc., which can be specifically set according to the actual situation.

[0077] Step S102: Determine the target dispersion strategy of the data table in the preset dispersion strategy according to the business logic of the data table.

[0078] After determining the business logic of the data table, for different business logics of the data table, a suitable dispersion strategy can be selected from the preset dispersion strategies as the target dispersion strategy of the data table. Among them, the preset dispersion strategy may be a dispersion strategy for the index of the data table. In a possible implementation manner, if the data table in the database includes multiple indexes, the preset dispersion strategy is used to disperse the data table according to different indexes of the data table respectively.

[0079] The embodiment of the present application does not limit the specific strategy and form of the preset dispersion strategy. In a possible implementation manner, the preset dispersion strategy is a dispersion strategy formulated for different business logics of the data table respectively. When selecting the target dispersion strategy, the target dispersion strategy of the business logic can be determined through the correspondence between the business logic and the dispersion strategy. In a possible implementation manner, the preset dispersion strategy can be implemented by multiple scripts. Then, according to the business logic of the data table, determining the target dispersion strategy of the data table in the preset dispersion strategy can be achieved by the server calling the script of the target dispersion strategy corresponding to the business logic of the data table.

[0080] Step S103: Perform storage partitioning on the data in the data table according to the target fragmentation strategy, and fragment the data in the data table according to the storage partitions.

[0081] After determining the target fragmentation strategy of the data table in the preset fragmentation strategy according to the business logic of the data table, the server performs storage partitioning on the data in the data table according to the target fragmentation strategy, and fragments the data in the data table according to the storage partitions.

[0082] The target fragmentation strategy includes the number of regions of the data table and the number and range of indexes in each region. By performing storage partitioning on the data in the data table according to the number of regions in the target fragmentation strategy, and writing the data in the data table into each region respectively according to the number and range of indexes in each region, the fragmentation of the data table in the database is realized.

[0083] Exemplarily, Table 1 is an exemplary partition table in an embodiment of the present application. By pre-estimating the data volume of the data table, for example, if the data volume exceeds 1 million (W), 1000 regions are allocated for the data table, and the indexes of each region are divided. The specific division is shown in Table 1.

[0084] Table 1 is an exemplary partition table in an embodiment of the present application

[0085] Region number Index minimum value Index maximum value 1 1 10000 2 10001 20000 3 20001 30000 … … … … … … N (N-1)*10000+1 N*10000 998 9970001 9980000 999 9980001 9990000 1000 9990001 10000000

[0086] As shown in Table 1, the number of regions is 1000, and each region is numbered with Arabic numerals respectively. The number of indexes in each region is 10000. Taking Region 1 as an example, the minimum index value of Region 1 is 1, and the maximum index value is 10000. Then the index range of Region 1 is 1 - 10000. Other regions are similar to Region 1. For specific content, reference can be made to Table 1, and details will not be elaborated here. When the program starts, perform storage partitioning on the data in the data table according to Table 1, and fragment the data in the data table according to the number and range of indexes in each region.

[0087] The preset fragmentation strategy is the fragmentation method of the indexes of the data table. In a possible implementation manner, the data table can also be fragmented according to the primary key of the data table. The method for fragmenting database data provided in the embodiment of the present application may further include before selecting the target fragmentation strategy of the data table in the preset fragmentation strategy according to the business logic of the data table:

[0088] Determine the type of the primary key of the data table; if the primary key type is a numeric type, divide the data table into multiple different primary key interval segments in the order of numeric auto-increment.

[0089] For different primary key types, different primary key scattering methods can be adopted. When the primary key type of the data table is a numeric type, the numeric auto-increment method can be used to divide the data table into multiple different primary key range segments, where each primary key range segment is independent and different from each other. When the primary key type of the data table is a string, the SHADOW_ROW_BITS algorithm can be used to scatter the string, divide the data table into multiple different primary key range segments, and each primary key range segment is independent and different from each other.

[0090] In the embodiments of the present application, by selecting a suitable scattering strategy as the target scattering strategy from the preset scattering strategies according to the business logic of the data table in the database, and then storing and partitioning and scattering the data in the data table according to the target scattering strategy, the distributed storage of the data in the database is realized. Compared with the single-point deployment in the prior art, the cost of data storage is reduced.

[0091] In Figure 2 Based on the embodiments shown, in a possible implementation manner, Figure 3 FIG. is a schematic flowchart of a method for scattering database data provided by another embodiment of the present application. This method can be executed by a device for scattering database data, and this device can be implemented in software and / or hardware. Hereinafter, the method for scattering database data will be described with a server as the execution subject. As Figure 3 shown, after scattering the data of the data table according to the storage partition in step S103 of the method for scattering database data provided by the embodiments of the present application, it may further include:

[0092] Step S201: Analyze the distribution of the data in the data table in each area.

[0093] Step S202: If the difference between the amounts of data stored in any two areas exceeds a preset quantity or a preset ratio, then update the target scattering strategy by using the data of the data table and the determination method of the target scattering strategy.

[0094] The data in the data table is partitioned and the data in the data table is scattered according to the number of indexes and the index range in each area. By obtaining and analyzing the distribution of the data in the data table in each area, the embodiments of the present application can determine whether the data in the data table is evenly distributed in the area. For example, the number of areas is 3, namely area A, area B, and area C, and the preset number is 2000. After scattering the data in the data table according to the storage partition, there are 10,000 pieces of data in area A, 9,957 pieces of data in area B, and 2,000 pieces of data in area C. Then the data volume between area A and area B is within the preset number range, the data volume between area A and area C exceeds the preset number, and the data volume between area B and area C also exceeds the preset number. Then the data table is not evenly distributed in each area. Another example, the number of areas is 3, namely area A, area B, and area C, and the preset ratio is 20%. After scattering the data in the data table according to the storage partition, there are 10,000 pieces of data in area A, 9,957 pieces of data in area B, and 2,000 pieces of data in area C. Then the difference in the data volume between area A and area B is within the preset ratio range, the difference in the data volume between area A and area C exceeds the preset ratio, and the difference in the data volume between area B and area C also exceeds the preset ratio. Then the data table is not evenly distributed in each area.

[0095] If the distribution gap of the data in the data table in each area is large, it may cause excessive data processing pressure in some areas and less data processing in some areas, resulting in unreasonable resource allocation. Therefore, when the distribution does not meet the uniform expectation, that is, if the difference between the data volumes stored in any two areas exceeds the preset number or preset ratio, the target scattering strategy can be updated by using the data in the data table and the determination method of the target scattering strategy.

[0096] In a possible implementation manner, the preset scattering strategies include a scattering strategy based on the distribution of stock data, a scattering strategy based on the distribution of upstream data, and a scattering strategy based on the distribution of yesterday's data. In the method provided by the embodiments of the present application, the target scattering strategy is updated by using the data in the data table and the determination method of the target scattering strategy. Taking the target scattering strategy as the scattering strategy based on the distribution of stock data as an example, the scattering strategy based on the distribution of stock data is updated by using the data in the data table and the determination method of the scattering strategy based on the distribution of stock data. Among them, the scattering strategy based on the distribution of stock data is generated by analyzing the stock data and through a certain determination method. When updating the scattering strategy based on the distribution of stock data, it is to analyze the data in the data table and regenerate the scattering strategy based on the distribution of stock data through the above determination method. Among them, the specific scheme of this determination method can refer to the present application Figure 4For the embodiments shown, no specific details are provided here.

[0097] In the embodiments of the present application, after the data in the data table is shuffled, the data distribution in the storage partition is analyzed. When the distribution does not meet the uniform expectation, the target shuffling strategy is updated using the data in the data table, ensuring the reliability of the target shuffling strategy.

[0098] In Figure 2 or Figure 3 Based on the embodiments shown, in a possible implementation manner, for the method of shuffling database data provided by the embodiments of the present application, the preset shuffling strategies include a shuffling strategy based on the distribution of stock data, a shuffling strategy based on the distribution of upstream data, and a shuffling strategy based on the distribution of yesterday's data. The specific implementation manners of each preset shuffling strategy are introduced below.

[0099] In a possible implementation manner, for the method of shuffling database data provided by the embodiments of the present application, Figure 4 is a flowchart of the shuffling strategy based on the distribution of stock data provided by an embodiment of the present application. As Figure 4 shown, the shuffling strategy based on the distribution of stock data in the embodiments of the present application includes:

[0100] Step S301: Dedup the stock data according to the index to obtain the deduplicated index quantity of the index.

[0101] In the embodiments of the present application, taking the index in the stock data as the IOU number as an example, the stock data is deduplicated according to the IOU number. There may be one or more pieces of data for each IOU number. By deduplicating the IOU numbers, the quantity of IOU numbers in the stock data can be obtained, that is, the deduplicated index quantity of the index is obtained.

[0102] Step S302: Determine the index quantity and index range that can be stored in each region according to the deduplicated index quantity and the preset region quantity.

[0103] After obtaining the deduplicated index quantity of the index in the stock data, the index quantity and index range that can be stored in each region are determined according to the deduplicated index quantity and the preset region quantity. The embodiments of the present application do not limit the specific implementation manner of how to determine the index quantity and index range that can be stored in each region according to the deduplicated index quantity and the preset region quantity. Among them, the preset region quantity can be determined by estimating the data volume in the data table. In a possible implementation manner, the preset region quantity is 1000. The embodiments of the present application only take this as an example and are not limited thereto. The preset region quantity can be specifically determined according to the actual business requirements.

[0104] In a possible implementation, the embodiment of the present application uses the ratio of the deduplicated index data volume to the number of preset regions as the number of indexes that can be stored in each region. Then, by numbering the indexes, the index range stored in each region can be determined. Exemplarily, the number of preset regions is 1000, and the number of indexes that can be stored in each region is 10000. Then the index range stored in the first region can be 0 - 10000, and so on. The embodiment of the present application will not elaborate further.

[0105] In another possible implementation, the method for database data scattering provided by the embodiment of the present application determines the number of indexes and the index range that can be stored in each region according to the deduplicated index number and the number of preset regions, including: determining the ratio of the number of rows of the stock data to the deduplicated index number as the factor multiple of the index; generating a new array using the deduplicated index number and the factor multiple; sorting the new array according to the index, and determining the number of indexes and the index range that can be stored in each region using the number of rows of the new array and the number of preset regions.

[0106] For the sake of easy introduction, exemplarily, after deduplicating the stock data according to the index, the deduplicated array [1, 2, 3] is obtained. The deduplicated index number is 3, and the number of rows of the stock data is 9. Then the factor multiple of the index is 9 / 3 = 3. The deduplicated array [1, 2, 3] is expanded by the factor multiple and sorted to obtain the new array [1, 1, 1, 2, 2, 2, 3, 3, 3]. Then, the ratio of the number of rows 9 of the new array to the number of preset regions is used as the index data that can be stored in each region. For example, if the number of preset regions is 3, then the number of indexes that can be stored in each region is 1, and the index data that can be stored is 3. For example: the first region can store 3 index data of index 1, the second region can store 3 index data of index 2, and the third region can store 3 index data of index 3.

[0107] After determining the number of indexes and the index range that can be stored in each region, the data table is stored and partitioned according to the number of preset regions, and then the data table is scattered according to the number of indexes and the index range in each region.

[0108] Based on this, in a possible implementation, step S102 of the method for database data scattering provided by the embodiment of the present application, determining the target scattering strategy of the data table according to the business logic of the data table in the preset scattering strategy, may include:

[0109] If the similarity between the data distribution of the data table and the data distribution of the stock data reaches the first preset similarity, then determine the scattering strategy based on the stock data distribution in the preset scattering strategy as the target scattering strategy.

[0110] The data distribution of the data table is basically the same as that of the stock data. For example, the similarity reaches the first preset similarity, and the specific value of the first preset similarity is not limited in the embodiments of the present application. The shuffling strategy based on the stock data distribution is used as the target shuffling strategy for the data table.

[0111] If the target shuffling strategy is the shuffling strategy based on the stock data distribution, then the data in the data table is stored and partitioned according to the target shuffling strategy. Using the shuffling strategy based on the stock data distribution, the number of regions and the number of indexes and the index range that can be stored in each region are determined, and the data table is stored and partitioned according to the preset number of regions, and the data in the data table is shuffled according to the number of indexes and the index range in each region.

[0112] In another possible implementation, the method for shuffling database data provided by the embodiments of the present application Figure 5 is a schematic flowchart of the shuffling strategy based on the upstream data distribution provided by an embodiment of the present application. As Figure 5 shown, the shuffling strategy based on the upstream data distribution in the embodiments of the present application includes:

[0113] Step S401: Determine the amount of index data that can be stored in each region according to the amount of index data and the number of index regions in the upstream data.

[0114] Step S402: Determine the number of regions of the data table according to the number of rows of the data table and the amount of index data that can be stored in each region.

[0115] Step S403: Perform duplicate removal processing on the upstream data according to the index to generate the number of deduplicated indexes of the index.

[0116] Step S404: Determine the number of indexes and the index range stored in each region according to the number of deduplicated indexes and the number of regions.

[0117] In a possible implementation, the ratio of the amount of index data and the number of index regions in the upstream data is used as the amount of index data that can be stored in each region, and then the ratio of the number of rows of the data table to the amount of index data that can be stored in each region is used as the number of regions of the data table.

[0118] Then, by performing duplicate removal processing on the upstream data according to the index, the number of deduplicated indexes of the index is generated. Furthermore, the ratio of the number of deduplicated indexes to the number of regions is used as the number of indexes stored in each region. According to the order of the indexes and the order of each region, the index range of each region is determined. Finally, the data table is stored and partitioned according to the number of regions, and the data in the data table is shuffled according to the number of indexes and the index range in each region.

[0119] Based on this, in another possible implementation, step S102 of the method for database data scattering provided by the embodiments of the present application, determining the target scattering strategy of the data table according to the business logic of the data table in the preset scattering strategy, may include:

[0120] If there is upstream data for the data table and there is a proportional relationship between the upstream data and the data volume of the data table, then determine the scattering strategy based on the upstream data distribution in the preset scattering strategy as the target scattering strategy.

[0121] When there is a proportional relationship between the upstream data of the data table and the data volume of the data table, take the scattering strategy based on the upstream data distribution as the target scattering strategy of the data table. The data table stores and partitions the data table according to the number of regions determined by the scattering strategy based on the upstream data distribution, and scatters the data of the data table according to the number of indexes and the index range determined by the scattering strategy based on the upstream data distribution.

[0122] In yet another possible implementation, the method for database data scattering provided by the embodiments of the present application Figure 6 is a schematic flowchart of the scattering strategy based on the data distribution of yesterday provided by an embodiment of the present application. As Figure 6 shown, the scattering strategy based on the data distribution of yesterday provided by the embodiments of the present application includes:

[0123] Step S501: De-duplicate the data of yesterday according to the index to obtain the de-duplicated index quantity of the index.

[0124] Step S502: Determine the number of indexes and the index range that can be stored in each region according to the de-duplicated index quantity and the preset number of regions.

[0125] The specific manner of the scattering strategy based on the data distribution of yesterday is similar to that of the scattering strategy based on the inventory data distribution. The difference is that the scattering strategy based on the data distribution of yesterday uses the data of yesterday of the data table as a reference, while the scattering strategy based on the inventory data distribution uses the inventory data of the data table as a reference. The specific manner can refer to the Figure 4 embodiment shown, and the embodiments of the present application will not be elaborated herein.

[0126] In yet another possible implementation, step S102 of the method for database data scattering provided by the embodiments of the present application, determining the target scattering strategy of the data table according to the business logic of the data table in the preset scattering strategy, may include:

[0127] If there is data of yesterday for the data table and the similarity between the data of yesterday and the data in the data table reaches the second preset similarity, then determine the scattering strategy based on the data distribution of yesterday in the preset scattering strategy as the target scattering strategy.

[0128] In the embodiment of the present application, multiple preset scattering strategies are set, and through the business logic of the data table, a suitable scattering strategy is selected from the multiple preset scattering strategies to scatter the data in the data table, thereby improving the reliability of data scattering.

[0129] The embodiment of the present application has been stress-tested by using the method of breaking up the database data provided by the embodiment of the present application. The stress-test results show that, compared with the prior art, for a database with the same amount of data, the storage time is greatly reduced, and its efficiency is greatly improved. Moreover, when the amount of data increases and the server resources remain unchanged, the time consumption conforms to linear growth. When the amount of data increases and the server resources increase year-on-year, the time consumption remains basically unchanged. Therefore, the method of breaking up the database data provided by the embodiment of the present application is reliable. Moreover, when there is hot data in the database, the server can still work stably.

[0130] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0131] Figure 7 is a schematic diagram of the structure of a device for breaking up database data provided by an embodiment of the present application. The device can be implemented by software and / or hardware, for example, by a server, such as Figure 7 As shown, the device for breaking up database data provided in the embodiment of the present application may include: a determination module 61 and a processing module 62.

[0132] The determination module 61 is used to determine the business logic of the data table in the database.

[0133] The determination module 61 is further used to determine a target scattering strategy of the data table in the preset scattering strategies according to the business logic of the data table.

[0134] The processing module 62 is used to perform storage partitioning on the data in the data table according to the target scattering strategy, and scatter the data in the data table according to the storage partitions.

[0135] In a possible implementation manner, the apparatus for breaking up database data provided in the embodiment of the present application, the determination module 61 is further used to:

[0136] Determine the primary key type of the data table; if the primary key type is a numeric type, divide the data table into multiple different primary key interval segments according to the numeric auto-increment method.

[0137] The device of this embodiment can perform the above Figure 2 The technical principle and technical effect of the method embodiment shown are similar to those of the above-mentioned embodiment and will not be described in detail here.

[0138] exist Figure 7On the basis of the embodiments shown, further, Figure 8 FIG. Figure 8 is a schematic structural diagram of a device for scattering database data provided by another embodiment of the present application. The device can be implemented in software and / or hardware. For example, it can be implemented by a server, as Figure 8 shown, the device for scattering database data provided by the embodiment of the present application may further include an update module 63.

[0139] The update module 63 is configured to: analyze the distribution of data in the data table within the storage partition; if the difference between the amounts of data stored in any two regions exceeds a preset quantity or a preset ratio, update the target scattering strategy by using the data in the data table and the determination method of the target scattering strategy.

[0140] The device of this embodiment can execute the Figure 3 method embodiment shown above. Its technical principle and technical effect are similar to those of the above embodiment and will not be elaborated here.

[0141] On the basis of the above embodiments, in a possible implementation manner, for the device for scattering database data provided by the embodiment of the present application, the preset scattering strategies include a scattering strategy based on the distribution of stock data, a scattering strategy based on the distribution of upstream data, and a scattering strategy based on the distribution of yesterday's data.

[0142] The determination module 61 is specifically configured to:

[0143] If the similarity between the data distribution of the data table and the data distribution of the stock data reaches a first preset similarity, determine the scattering strategy based on the distribution of the stock data as the target scattering strategy among the preset scattering strategies.

[0144] If there is upstream data in the data table and there is a proportional relationship between the upstream data and the amount of data in the data table, determine the scattering strategy based on the distribution of the upstream data as the target scattering strategy among the preset scattering strategies.

[0145] If there is yesterday's data in the data table and the similarity between yesterday's data and the data in the data table reaches a second preset similarity, determine the scattering strategy based on the distribution of yesterday's data as the target scattering strategy among the preset scattering strategies.

[0146] In a possible implementation manner, for the device for scattering database data provided by the embodiment of the present application, the scattering strategy based on the distribution of stock data includes:

[0147] Perform duplicate removal processing on the stock data according to the index to obtain the number of de-duplicated indexes of the index; determine the number of indexes and the index range that can be stored in each region according to the number of de-duplicated indexes and the preset number of regions.

[0148] In a possible implementation, the apparatus for database data scattering provided by the embodiments of the present application determines the number of indexes and the index range that can be stored in each area according to the number of deduplicated indexes and the number of preset areas, including:

[0149] Determine the ratio of the number of rows of the stock data to the number of deduplicated indexes as the factor multiple of the indexes; generate a new array using the number of deduplicated indexes and the factor multiple; sort the new array according to the indexes, and determine the number of indexes and the index range that can be stored in each area by using the number of rows of the data table and the number of preset areas.

[0150] In a possible implementation, the apparatus for database data scattering provided by the embodiments of the present application, based on the scattering strategy of upstream data distribution, includes:

[0151] Determine the amount of index data that can be stored in each area according to the amount of index data in the upstream data and the number of areas of the indexes; determine the number of areas of the data table according to the number of rows of the data table and the amount of index data that can be stored in each area; perform deduplication processing on the upstream data according to the indexes to generate the number of deduplicated indexes of the indexes; determine the number of indexes and the index range stored in each area according to the number of deduplicated indexes and the number of areas.

[0152] In a possible implementation, the apparatus for database data scattering provided by the embodiments of the present application, based on the scattering strategy of yesterday's data distribution, includes:

[0153] Perform deduplication processing on yesterday's data according to the indexes to obtain the number of deduplicated indexes of the indexes; determine the number of indexes and the index range that can be stored in each area according to the number of deduplicated indexes and the number of preset areas.

[0154] The device embodiments provided by the present application are merely illustrative. Figure 7 and Figure 8 The module division in is merely a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system. The coupling between each module can be realized through some interfaces, and these interfaces are usually electrical communication interfaces, but it does not exclude the possibility of being mechanical interfaces or other forms of interfaces. Therefore, the modules described as separate components may or may not be physically separated, and can be located in one place or distributed to different positions of the same or different devices.

[0155] Figure 9 is a schematic structural diagram of an electronic device provided by the embodiments of the present application. The electronic device can be a server. As Figure 9 shown, the electronic device includes:

[0156] a receiver 70, a transmitter 71, a processor 72, a memory 73, and a computer program; wherein the receiver 70 and the transmitter 71 are configured to implement data transmission with other devices, the computer program is stored in the memory 73 and is configured to be executed by the processor 72, and the computer program includes instructions for performing the method for scattering the database data as described above. For the content and effects thereof, please refer to the method embodiments.

[0157] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When at least one processor of a user device executes the computer-executable instructions, the user device performs the various possible methods described above.

[0158] Among them, the computer-readable medium includes a computer storage medium and a communication medium, and the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also exist as discrete components in a communication device.

[0159] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0160] An embodiment of the present application further provides a computer program product, including computer instructions, which when executed by a processor, implement each step in the method for scattering database data in the above embodiments.

[0161] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0162] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for scrambling database data, characterized in that, Including: Determine the business logic of the data table in the database; According to the business logic of the data table, determine the target fragmentation strategy of the data table in the preset fragmentation strategies; According to the target fragmentation strategy, perform storage partitioning on the data in the data table, and fragment the data of the data table according to the storage partitioning; The preset fragmentation strategies include a fragmentation strategy based on the distribution of stock data, a fragmentation strategy based on the distribution of upstream data, and a fragmentation strategy based on the distribution of yesterday's data; The determining the target fragmentation strategy of the data table in the preset fragmentation strategies according to the business logic of the data table includes: If the similarity between the data distribution of the data table and the data distribution of the stock data reaches a first preset similarity, determine the fragmentation strategy based on the distribution of stock data as the target fragmentation strategy in the preset fragmentation strategies; If there is upstream data in the data table and there is a proportional relationship between the upstream data and the data volume of the data table, determine the fragmentation strategy based on the distribution of upstream data as the target fragmentation strategy in the preset fragmentation strategies; If there is yesterday's data in the data table and the similarity between the yesterday's data and the data in the data table reaches a second preset similarity, determine the fragmentation strategy based on the distribution of yesterday's data as the target fragmentation strategy in the preset fragmentation strategies.

2. The method according to claim 1, characterized in that, After fragmenting the data of the data table according to the storage partitioning, it further includes: Analyze the data distribution of the data table in each area; If the difference between the data volumes stored in any two areas exceeds a preset quantity or a preset ratio, update the target fragmentation strategy by using the data of the data table and the determination method of the target fragmentation strategy.

3. The method according to claim 1, wherein The fragmentation strategy based on the distribution of stock data includes: Perform deduplication processing on the stock data according to the index to obtain the deduplicated index quantity of the index; Determine the ratio of the number of rows of the stock data to the deduplicated index quantity as the factor multiple of the index; Generate a new array by using the deduplicated index quantity and the factor multiple; Sort the new array according to the index, and determine the index quantity and index range that can be stored in each area by using the number of rows of the new array and the preset number of areas.

4. The method according to claim 1, characterized in that, The fragmentation strategy based on the distribution of upstream data includes: According to the index data volume and the number of index areas in the upstream data, determine the index data volume that can be stored in each area; According to the number of rows of the data table and the index data volume that can be stored in each area, determine the number of areas of the data table; Perform deduplication processing on the upstream data according to the index to generate the deduplicated index quantity of the index; According to the deduplicated index quantity and the number of areas, determine the index quantity and index range stored in each area.

5. The method according to claim 1, characterized in that, The fragmentation strategy based on the distribution of yesterday's data includes: Perform deduplication processing on the yesterday's data according to the index to obtain the deduplicated index quantity of the index; According to the deduplicated index quantity and the preset number of areas, determine the index quantity and index range that can be stored in each area.

6. The method according to claim 1 or 2, characterized in that Before selecting the target fragmentation strategy of the data table in the business logic according to the data table in the preset fragmentation strategy, the following steps are also included: Determine the primary key type of the data table; If the primary key type is a numeric type, divide the data table into multiple different primary key range segments in the order of numeric auto-increment.

7. An apparatus for scrambling database data, characterized in that, Including: A determination module for determining the business logic of the data table in the database; The determination module is further configured to determine the target fragmentation strategy of the data table in the preset fragmentation strategy according to the business logic of the data table; A processing module for storing and partitioning the data in the data table according to the target fragmentation strategy, and fragmenting the data of the data table according to the storage partition; The preset fragmentation strategy includes a fragmentation strategy based on the distribution of stock data, a fragmentation strategy based on the distribution of upstream data, and a fragmentation strategy based on the distribution of yesterday's data; The determination module is specifically configured to, if the similarity between the data distribution of the data table and the data distribution of the stock data reaches a first preset similarity, determine the fragmentation strategy based on the distribution of the stock data as the target fragmentation strategy in the preset fragmentation strategy; If there is upstream data in the data table and there is a proportional relationship between the upstream data and the data volume of the data table, determine the fragmentation strategy based on the distribution of the upstream data as the target fragmentation strategy in the preset fragmentation strategy; If there is yesterday's data in the data table and the similarity between the yesterday's data and the data in the data table reaches a second preset similarity, determine the fragmentation strategy based on the distribution of yesterday's data as the target fragmentation strategy in the preset fragmentation strategy.

8. An electronic device, comprising: A processor and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by the processor, they are used to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Data storage method and device in distributed storage system

    CN112527492A