Data storage method and device, electronic equipment and storage medium
By determining the target data table and data distribution characteristics in the distributed database system, selecting appropriate shard keys and performing data sharding processing, the data tilt problem is solved, data uniform distribution and node load balancing are achieved, and system performance is improved.
Patent Information
- Application Number
- CN202510240883.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
Distributed database systems are prone to facing data tilt problems when the data volume increases sharply, resulting in excessive burden on some nodes, low processing efficiency, resource bottlenecks, and affect system performance and stability.
By determining the target data table and target data distribution characteristics, the target shard key is determined based on these characteristics, and using pre-configured data sharding rules, the target data table is sharded to make the data evenly distributed on multiple shards.
It prevents data tilt during distributed database sharding, ensures even data distribution and load balancing of each node, thereby improving the overall performance of the system.
Smart Images

Figure CN120179647A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to a data storage method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of the Internet and the continuous expansion of the scale of Internet data, the demand for data storage is increasing, and thus distributed databases have emerged. A distributed database is a large database system formed by connecting multiple database units that are geographically dispersed but logically centralized through a computer network.
[0003] However, in practical applications, with the rapid growth of data volume, distributed database systems often face the challenge of data skew. Data skew is mainly manifested as the uneven distribution of data storage, that is, some nodes are overloaded due to carrying excessive data, which not only leads to low processing efficiency of these nodes, but may also cause resource bottlenecks, thus affecting the performance and stability of the entire system.
[0004] To effectively solve this problem, a data storage method is proposed to prevent data skew during the sharding process of a distributed database, so as to ensure that data can be evenly distributed and the loads of each node are balanced, thereby improving the overall performance of the system. Summary of the Invention
[0005] The present invention provides a data storage method, apparatus, electronic device, and storage medium to prevent data skew during the sharding process of a target distributed database, so as to ensure that data can be evenly distributed and the loads of each node are balanced, thereby improving the overall performance of the system.
[0006] According to one aspect of the present invention, a data storage method is provided, and the method includes:
[0007] Determine a target data table, where the target data table includes multiple target data to be stored in a target distributed database;
[0008] Determine the target data distribution characteristics of the multiple target data in the target distributed database, and determine the target sharding key of the target data table based on the target data distribution characteristics;
[0009] Based on the data sharding rules pre-configured according to the target sharding key, perform data sharding processing on the target data table, so that the multiple target data are stored on multiple shards of the target distributed database.
[0010] Optionally, determining the target data distribution characteristics of the multiple target data in the target distributed database includes: inputting the multiple target data into a preset target data distribution model to obtain the target data distribution characteristics of the multiple target data in the target distributed database.
[0011] Optionally, the method further includes: inputting the multiple target data into a pre-constructed initial data distribution model to obtain the initial data distribution characteristics of the multiple target data in the target distributed database; determining an initial sharding key corresponding to the initial data distribution characteristics, and performing data sharding processing on the target data table based on the data sharding rule of the initial sharding key; determining the sharding load information of the target distributed database according to the processing result of the data sharding processing; and when the sharding load information meets a preset sharding load condition, determining the initial data distribution model as the target data distribution model.
[0012] Optionally, the method further includes: when the sharding load information does not meet the preset sharding load condition, adjusting the model parameters of the initial data distribution model to obtain the target data distribution model, where the model parameters are used to select the data distribution characteristics of the target data in the target data table.
[0013] Optionally, the method further includes: determining the service information of the target data table, and constructing an initial data distribution model of the target data table according to the service information.
[0014] Optionally, the method further includes: determining the data processing information of the multiple shards and displaying the data processing information, where the data processing information includes at least one of the read and write request volume of each shard, the response time of each shard, the ratio of the data volume of a single shard to the total data volume, and the growth rate of the data volume of each shard.
[0015] Optionally, the target sharding key includes a single sharding key or a combination of multiple sharding keys, and the data sharding rule includes at least one of a range sharding rule, a hash sharding rule, and a list sharding rule.
[0016] According to another aspect of the present invention, a data storage device is provided. The device includes:
[0017] A target data determination module, configured to determine a target data table, where the target data table includes multiple target data to be stored in a target distributed database;
[0018] A sharding key determination module, configured to determine the target data distribution characteristics of the multiple target data in the target distributed database, and determine the target sharding key of the target data table based on the target data distribution characteristics;
[0019] A target data storage module, configured to perform data sharding processing on the target data table based on a pre-configured data sharding rule of the target sharding key, so that multiple pieces of the target data are stored on multiple shards of the target distributed database.
[0020] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0021] At least one processor; and
[0022] A memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the data storage method according to any embodiment of the present invention.
[0024] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the data storage method according to any embodiment of the present invention when executed.
[0025] The technical solution of the embodiment of the present invention determines a target data table, where the target data table includes multiple pieces of target data to be stored in a target distributed database; further determines the target data distribution characteristics of the multiple pieces of target data in the target distributed database, and determines the target sharding key of the target data table based on the target data distribution characteristics; thereby performing data sharding processing on the target data table based on a pre-configured data sharding rule of the target sharding key, so that multiple pieces of the target data are stored on multiple shards of the target distributed database. The technical solution of the embodiment of the present invention realizes preventing data skew during the sharding process of the target distributed database, ensuring that the data can be evenly distributed and the loads of each node are balanced, thereby improving the overall performance of the system.
[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0028] Figure 1 Schematic flowchart of a data storage method provided by an embodiment of the present invention;
[0029] Figure 2 Schematic flowchart of a data storage method provided by an embodiment of the present invention;
[0030] Figure 3 Schematic structural diagram of a data storage device provided by an embodiment of the present invention;
[0031] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0035] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0036] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way to send a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0037] It can be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation of the present disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0038] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0039] Figure 1 It is a schematic flowchart of a data storage method provided by an embodiment of the present invention. This embodiment is applicable to the situation of storing data in a data table. This method can be executed by a data storage device, and the data storage device can be implemented in the form of hardware and / or software. The data storage device can be configured in an electronic device such as a computer or a server. As Figure 1 shown, the method of this embodiment includes:
[0040] S110. Determine a target data table, where the target data table includes a plurality of target data to be stored in a target distributed database.
[0041] Among them, the target data table can be understood as a data table for storing a plurality of target data. The number of target data tables can be one, two or more. The target data can be understood as data that already exists in the target data table but has not been stored in the target distributed database. The target distributed database can be understood as a distributed database for storing a plurality of the target data. In the embodiment of the present invention, the target distributed database can be set according to actual needs, and is not specifically limited herein. For example, MySQL database, MongoDB database, Hadoop database, etc.
[0042] In an embodiment of the present invention, the manner of obtaining the target data table may include: in response to a target trigger operation of at least one target object, obtaining at least one row of target information. After obtaining at least one row of the target information, a target data table may be obtained based on at least one row of the target information. It can be understood that a single target data in the target data table may be a row of target information in the data table, and each row of target information may be used as a row of the data table. The relationship between the target object and the target information may be one-to-one or one-to-many. That is, one target object corresponds to one row of target information or multiple target information.
[0043] Exemplarily, the target object may be a user. When the target trigger operation is a user registration operation, the target data table may be a user information table. Each row of target information in the target data table may be the user information input by the user after receiving each user registration operation. The user information includes a user identifier and user attribute information.
[0044] In an embodiment of the present invention, there are various ways to obtain the target data table based on at least one row of the target information. As an optional implementation manner in the embodiment of the present invention, obtaining the target data table based on at least one row of the target information includes: presetting a data table, and storing at least one row of the target information in the preset data table, so that a target data table can be obtained. As another optional implementation manner in the embodiment of the present invention, obtaining the target data table based on at least one row of the target information includes: generating a data table including at least one row of the target information based on at least one row of the target information, that is, the target data table.
[0045] S120. Determine the target data distribution characteristics of multiple target data in the target distributed database, and determine the target sharding key of the target data table based on the target data distribution characteristics.
[0046] Among them, the target data distribution characteristics can be understood as the target sharding key for determining the target data table. The target data distribution characteristics may include at least one of the central tendency, dispersion degree, skewness, and kurtosis of multiple target data in the target data table. Among them, the central tendency of multiple target data in the target data table describes the characteristics of the position where multiple target data in the target data table gather, reflecting the central position or average level of multiple target data. Optionally, the central tendency of multiple target data can be reflected by calculating at least one of the mean, median, and mode of multiple target data. The dispersion degree of multiple target data in the target data table describes the degree of deviation of data points from the central position or the characteristics of the dispersion degree of multiple target data, reflecting the difference degree between multiple target data in the target data table. Optionally, the difference degree between multiple target data in the target data table can be reflected by calculating at least one of the variance, standard deviation, and range of multiple target data. The skewness of multiple target data in the target data table describes the characteristics of the asymmetry of the data distribution of multiple target data, which can reflect whether the data distribution is biased towards one side. The data bias can be reflected by positive skewness and negative skewness. The kurtosis of multiple target data in the target data table can be used to determine the characteristics of the sharpness of the data distribution of multiple target data, which can reflect the flatness of the data distribution. The flatness of the data distribution can be reflected by leptokurtosis and platykurtosis.
[0047] In an embodiment of the present invention, the target sharding key may include a single sharding key or a combination of multiple sharding keys. Among them, the target sharding key is configured with a data sharding rule. In an embodiment of the present invention, the data sharding rule sets how to disperse multiple target data in the target data table to different shards. Specifically, when the target sharding key is a single sharding key, that is, the data sharding rule of this sharding key is the data sharding rule of the target sharding key. When the target sharding key is a combination of multiple sharding keys, a corresponding data sharding rule can be set for each assigned key. Then, the data sharding rule of the target sharding key is the data sharding rule obtained by combining the data sharding rules of multiple sharding keys.
[0048] In an embodiment of the present invention, the data sharding rule may include at least one of a range sharding rule, a hash sharding rule, and a list sharding rule. Among them, the range sharding rule can be used to divide data based on the range of the values of the sharding key. For example, store the data with the target object ID in the range of 1 - 10,000 on node A, and store the data with the target object ID in the range of 10,001 - 20,000 on node B. The hash sharding rule can be used to perform a hash operation on the value of the sharding key, and then disperse the data to different shards according to the hash value. The list sharding rule can be used to divide data according to a predefined list. For example, data can be assigned to different shards according to geographical locations (such as countries, cities).
[0049] In an embodiment of the present invention, determining the target data distribution characteristics of multiple pieces of the target data in the target distributed database may include: when the number of target data tables is multiple, determining the target data distribution characteristics of all the data in all the target data tables in the target distributed database. When the number of target data tables is one, determining the target data distribution characteristics of all the data in the target data table in the target distributed database.
[0050] In an embodiment of the present invention, determining the target sharding key of the target data table based on the target data distribution characteristics may include: determining at least one candidate sharding key of the target data table based on the target data distribution characteristics, and determining the target sharding key based on at least one candidate sharding key. Among them, determining the target sharding key based on at least one candidate sharding key includes: when the number of candidate sharding keys is one, the candidate sharding key can be determined as the target sharding key. When there are multiple candidate sharding keys, the target sharding key is obtained based on the multiple candidate sharding keys. Among them, there are multiple ways to obtain the target sharding key based on multiple candidate sharding keys. For example, at least some of the candidate sharding keys can be combined, and the combination of at least some of the candidate sharding keys can be used as the target sharding key. Or, one sharding key can be selected from the multiple candidate sharding keys as the target sharding key. In an embodiment of the present invention, specifically, selecting one sharding key from the multiple candidate sharding keys as the target sharding key can be to select the sharding key with a high cardinality from the multiple candidate sharding keys as the target sharding key, which can better ensure the uniform distribution of multiple target data on multiple shards of the target distributed database.
[0051] Based on the above embodiments, to select a suitable target sharding key, obtaining the target sharding key based on multiple candidate sharding keys may include: further evaluating each candidate sharding key. For example, the candidate sharding key can be used as the target sharding key to simulate the distribution of multiple pieces of the target data across multiple shards in the target distributed database. Thus, the actual evaluation result of the candidate sharding key can be obtained. When the actual evaluation result of the candidate sharding key reaches the expected evaluation result, the candidate sharding key can be used as the target sharding key. It can be understood that when the number of candidate sharding keys that reach the expected evaluation result is one, the candidate sharding key can be determined as the target sharding key. When the number of candidate sharding keys that reach the preset evaluation result is multiple, the combination of multiple candidate sharding keys can be determined as the target sharding key.
[0052] In an embodiment of the present invention, each candidate sharding key is configured with data sharding rules. Evaluating each candidate sharding key may include: when the number of data sharding rules configured for the candidate sharding key is multiple, each of the data sharding rules can be used to simulate the data distribution of multiple target data in the target distributed database, and a data sharding simulation result corresponding to each data sharding rule can be obtained. Furthermore, based on the data sharding simulation result, one data sharding rule corresponding to the candidate sharding key can be determined. It can be understood that when a candidate sharding key is configured with multiple data sharding rules, the types of the multiple data sharding rules can be the same or different.
[0053] S130. Perform data sharding processing on the target data table based on the data sharding rules pre-configured for the target sharding key, so that multiple pieces of the target data are stored on multiple shards of the target distributed database.
[0054] In an embodiment of the present invention, the target distributed database includes multiple shards to store multiple pieces of target data dispersedly.
[0055] As an optional implementation manner in an embodiment of the present invention, when the target sharding key is a single sharding key, the data sharding rules pre-configured for the target sharding key can be used to perform data sharding processing on the target data table, so that multiple pieces of the target data are stored on multiple shards of the target distributed database.
[0056] As another alternative implementation in the embodiments of the present invention, when the target sharding key is a combination of multiple sharding keys, the priority of each sharding key in the multiple sharding key combinations can be determined; and data sharding processing is performed on multiple target data in the target data table according to the data sharding rules configured with the corresponding sharding keys in sequence according to the priority, so as to disperse and store the multiple target data on multiple shards of the target distributed database.
[0057] Based on the above embodiments, the embodiments of the present invention further include: determining the data processing information of multiple shards, and displaying the data processing information. Wherein, the data processing information may include at least one of the read / write request volume of each shard, the response time of each shard, the ratio of the data volume of a single shard to the total data volume, and the growth rate of the data volume of each shard. In the embodiments of the present invention, by visually displaying the data processing information of each shard in the target distributed database, it is convenient to timely understand the data processing information of each shard in the target distributed database, so as to timely adjust the data storage strategy.
[0058] It should be noted that by dispersing and storing multiple target data in the target data table on multiple shards of the target database, the stability of data distribution can be ensured, unnecessary data migration can be prevented, this performance balance can be maintained, the situation that some shards are overloaded or underloaded due to changes in data distribution can be avoided, the access performance of hot data caused by data migration can be prevented from decreasing, the round-trip time and storage overhead of data in the network can be reduced, the relative stability of the storage structure can be maintained, the effective utilization rate of storage resources can be improved, the consistency of data on each shard can be ensured, and the integrity of transactions during execution can be guaranteed.
[0059] The technical solution of the embodiments of the present invention determines a target data table, where the target data table includes multiple target data to be stored in a target distributed database; then determines the target data distribution characteristics of the multiple target data in the target distributed database, and determines the target sharding key of the target data table based on the target data distribution characteristics; and then performs data sharding processing on the target data table according to the data sharding rules pre-configured based on the target sharding key, so that the multiple target data are stored on multiple shards of the target distributed database. The technical solution of the embodiments of the present invention realizes the prevention of data skew phenomenon during the sharding process of the target distributed database, so as to ensure that the data can be evenly distributed and the loads of each node are balanced, thereby improving the overall performance of the system.
[0060] Figure 2The flowchart of a data storage method provided by an embodiment of the present invention. Optionally, on the basis of the foregoing embodiment, determining the target data distribution characteristics of a plurality of the target data in the target distributed database includes: inputting the plurality of target data into a preset target data distribution model to obtain the target data distribution characteristics of the plurality of the target data in the target distributed database. Technical features that are the same as or similar to those in the above embodiment will not be described in detail herein. As Figure 2 shown, the method of this embodiment specifically includes:
[0061] S210. Determine a target data table, where the target data table includes a plurality of target data to be stored in the target distributed database.
[0062] S220. Input the plurality of target data into a preset target data distribution model to obtain the target data distribution characteristics of the plurality of the target data in the target distributed database, and determine the target sharding key of the target data table based on the target data distribution characteristics.
[0063] Among them, the target data distribution model can be understood as a data distribution model for obtaining the target data distribution characteristics of a plurality of the target data in the target distributed database. In the embodiment of the present invention, the method for obtaining the target data distribution module includes: inputting the plurality of target data into a pre-constructed initial data distribution model to obtain the initial data distribution characteristics of the plurality of the target data in the target distributed database; determining an initial sharding key corresponding to the initial data distribution characteristics, and performing data sharding processing on the target data table based on the data sharding rule of the initial sharding key; determining the sharding load information of the target distributed database according to the processing result of the data sharding processing; and when the sharding load information meets a preset sharding load condition, determining the initial data distribution model as the target data distribution model.
[0064] Among them, the initial data distribution model can be understood as a pre-constructed data distribution model for determining the data distribution characteristics of the target data. Optionally, the initial data distribution model can be a data model for data feature extraction, that is, a data feature extraction model. The data feature extraction model can be obtained by training in advance using business data and the corresponding data features of the business data as training samples. The initial data distribution characteristics can be the data distribution characteristics of the target data obtained through the initial data distribution model. The initial sharding key can be understood as a sharding key determined based on the initial distribution characteristics for sharding multiple target data in the target data table. The data sharding rule of the initial sharding key can be configured according to experience. The sharding load information can be used to determine the data carrying capacity and resource allocation amount of each shard in the target distributed database. The preset sharding load condition can be set according to actual needs. For example, the upper limit and / or lower limit of the data carrying capacity of each shard.
[0065] Specifically, multiple pieces of the target data can be input into a pre-constructed initial data distribution model. Thus, the initial data distribution characteristics of the multiple pieces of target data in the target distributed database can be obtained. Furthermore, a sharding key in the target data table can be determined based on the initial data distribution characteristics, that is, the initial sharding key corresponding to the initial data distribution characteristics can be obtained. Furthermore, based on the data sharding rule of the initial sharding key, data sharding processing can be performed on the target data table. Thus, the sharding load information of the target distributed database can be determined according to the processing result of the data sharding processing. When the sharding load information reaches the preset sharding load condition, the initial data distribution model can be determined as the target data distribution model.
[0066] Based on the above example, the method may further include: when the sharding load information does not reach the preset sharding load condition, the model parameters of the initial data distribution model can be adjusted to obtain the target data distribution model. Among them, the model parameters can be used to select the data distribution characteristics of the target data in the target data table. In the embodiments of the present invention, the number of model parameters can be one, two or more than two, and can be set according to actual needs.
[0067] In an embodiment of the present invention, the method for constructing an initial data distribution model includes: determining the business information of the target data table, and constructing the initial data distribution model of the target data table according to the business information. Herein, the business information of the target data table can be understood as the business information of the business to which the target data in the target data table belongs. The business information of the target data table includes business requirements, data sources, and data attributes. Among them, business requirements can be used to clarify what business requirements the target data table needs to meet. For example, what types of data need to be stored. Data sources can include data formats, data update frequencies, etc. Data attributes can identify the target attributes in the target data table, such as primary keys. It can be understood that the initial data distribution models corresponding to different business types can be the same or different. Constructing the initial data distribution model of the target data table according to the business information includes: constructing the initial data distribution model of the target data table according to the business type and business processing flow corresponding to the data in the target data table.
[0068] S230. Perform data sharding processing on the target data table based on the pre-configured data sharding rules for the target sharding key, so that multiple pieces of the target data are stored on multiple shards of the target distributed database.
[0069] The technical solution of the embodiment of the present invention realizes the function of accurately and conveniently determining the target data distribution characteristics of the target data by inputting multiple pieces of target data into a preset target data distribution model to obtain the target data distribution characteristics of the multiple pieces of target data in the target distributed database.
[0070] Figure 3 It is a schematic structural diagram of a data storage device provided by an embodiment of the present invention. As Figure 3 shown, the device includes: a target data determination module 310, a sharding key determination module 320, and a target data storage module 320.
[0071] Among them, the target data determination module 310 is used to determine a target data table, where the target data table includes multiple pieces of target data to be stored in the target distributed database; the sharding key determination module 320 is used to determine the target data distribution characteristics of the multiple pieces of target data in the target distributed database, and determine the target sharding key of the target data table based on the target data distribution characteristics; the target data storage module 320 is used to perform data sharding processing on the target data table based on the pre-configured data sharding rules for the target sharding key, so that multiple pieces of the target data are stored on multiple shards of the target distributed database.
[0072] In the technical solution of the embodiment of the present invention, a target data table is determined by a target data determination module, where the target data table includes a plurality of target data to be stored in a target distributed database; a target data distribution feature of the plurality of target data in the target distributed database is determined by a shard key determination module 320, and a target shard key of the target data table is determined based on the target data distribution feature; the target data storage module 320 performs data sharding processing on the target data table based on a pre-configured data sharding rule of the target shard key, so that the plurality of target data are stored on a plurality of shards of the target distributed database. The technical solution of the embodiment of the present invention realizes preventing data skew during the sharding process of the target distributed database, ensuring that data can be evenly distributed and the load of each node is balanced, thereby improving the overall performance of the system.
[0073] Optionally, the shard key determination module 320 includes a data distribution feature determination unit; wherein, the data distribution feature determination unit is configured to input a plurality of target data into a preset target data distribution model to obtain a target data distribution feature of the plurality of target data in the target distributed database.
[0074] Optionally, the device further includes a target model obtaining module, where the target model obtaining module is configured to input a plurality of the target data into a pre-constructed initial data distribution model to obtain an initial data distribution feature of the plurality of target data in the target distributed database; determine an initial shard key corresponding to the initial data distribution feature, and perform data sharding processing on the target data table based on the data sharding rule of the initial shard key; determine shard load information of the target distributed database according to the processing result of the data sharding processing; and determine the initial data distribution model as the target data distribution model when the shard load information reaches a preset shard load condition.
[0075] Optionally, the device further includes a module parameter adjustment module; wherein, the module parameter adjustment module is configured to adjust model parameters of the initial data distribution model to obtain the target data distribution model when the shard load information does not reach a preset shard load condition, where the model parameters are used to select a data distribution feature of target data in the target data table.
[0076] Optionally, the device further includes an initial model construction module; wherein, the initial model construction module is configured to determine service information of the target data table and construct an initial data distribution model of the target data table according to the service information.
[0077] Optionally, the device further includes a data processing information display module; wherein, the data processing information display module is configured to determine the data processing information of the multiple shards and display the data processing information, where the data processing information includes at least one of the read / write request volume of each shard, the response time of each shard, the ratio of the data volume of a single shard to the total data volume, and the growth rate of the data volume of each shard.
[0078] Optionally, the target shard key includes a single shard key or a combination of multiple shard keys, and the data sharding rule includes at least one of a range sharding rule, a hash sharding rule, and a list sharding rule.
[0079] The data storage device provided by the embodiments of the present invention can execute the data storage method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0080] It should be noted that the various units and modules included in the above data storage device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present invention.
[0081] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0082] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0083] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0084] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data storage method.
[0085] In some embodiments, the data storage method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data storage method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data storage method in any other appropriate way (e.g., by means of firmware).
[0086] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0087] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0088] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0090] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0091] A computing system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0092] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0093] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data storage method, characterized in that: include: Determine a target data table, wherein the target data table includes a plurality of target data to be stored in a target distributed database; Determine target data distribution characteristics of the plurality of target data in the target distributed database, and determine a target shard key of the target data table based on the target data distribution characteristics; Based on the data sharding rule pre-configured by the target sharding key, the target data table is subjected to data sharding processing, so that a plurality of the target data are stored in a plurality of shards of the target distributed database.
2. The method according to claim 1, characterized in that The determining target data distribution characteristics of the plurality of target data in the target distributed database comprises: A plurality of target data are input into a preset target data distribution model to obtain target data distribution characteristics of the plurality of target data in the target distributed database.
3. The method according to claim 2, characterized in that The method further comprises: Inputting the plurality of target data into a pre-built initial data distribution model to obtain initial data distribution characteristics of the plurality of target data in the target distributed database; Determine an initial sharding key corresponding to the initial data distribution feature, and perform data sharding processing on the target data table based on a data sharding rule of the initial sharding key; Determining shard load information of the target distributed database according to the processing result of the data shard processing; When the shard load information reaches a preset shard load condition, the initial data distribution model is determined as the target data distribution model.
4. The method according to claim 3, characterized in that The method further comprises: When the shard load information does not reach a preset shard load condition, the model parameters of the initial data distribution model are adjusted to obtain the target data distribution model, wherein the model parameters are used to select data distribution characteristics of the target data in the target data table.
5. The method according to claim 3, characterized in that: The method further comprises: The business information of the target data table is determined, and an initial data distribution model of the target data table is constructed according to the business information.
6. The method according to claim 1, characterized in that The method further comprises: Determine data processing information of the multiple shards and display the data processing information, wherein the data processing information includes at least one of the amount of read and write requests for each shard, the response time of each shard, the ratio of the data volume of a single shard to the total data volume, and the growth rate of the data volume of each shard.
7. The method according to claim 1, characterized in that The target sharding key includes a single sharding key or a combination of multiple sharding keys, and the data sharding rule includes at least one of a range sharding rule, a hash sharding rule, and a list sharding rule.
8. A data storage device, characterized in that: include: A target data determination module, used to determine a target data table, wherein the target data table includes a plurality of target data to be stored in a target distributed database; A shard key determination module, used to determine target data distribution characteristics of the plurality of target data in the target distributed database, and determine a target shard key of the target data table based on the target data distribution characteristics; The target data storage module is used to perform data sharding processing on the target data table based on the data sharding rules pre-configured by the target sharding key, so that multiple target data are stored in multiple shards of the target distributed database.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the data storage method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data storage method according to any one of claims 1 to 7 when executed.