A data processing method, system, electronic device and storage medium
By calculating the element distribution value of the target data table in the distributed database and determining the fields corresponding to the minimum value as the distribution key, the problem of query performance degradation caused by data skew in the distributed database is solved, and a more uniform data distribution and performance improvement is achieved.
Patent Information
- Application Number
- CN202210455078.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-24
AI Technical Summary
Data skew problems caused by the primary and backup mode and multiple backup copies in distributed databases have led to degradation in database query performance.
By obtaining the field and element information of the target data table, calculate the element distribution value of each field, and determine the field with the smallest element distribution value as the distribution key of the target distribution table to optimize the data distribution.
It effectively solves the data skew problem, makes the file distribution of the target data table more evenly, and improves the database query performance.
Smart Images

Figure CN114791912B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of Internet technologies, and particularly relates to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of information technology and the continuous advancement of the informatization and digitalization processes, distributed databases have greatly solved the problem of single-node data capacity, and have won favor with advantages such as good scalability and high service performance.
[0003] Since most distributed databases use the master-slave mode for data disaster recovery and generally have multiple backup copies, the total physical space occupied by single-table data at the bottom layer grows rapidly. As the amount of table data increases, the amount of data on some nodes will be much higher than that on other nodes. That is, the phenomenon of data skew occurs, which in turn leads to a decline in database query performance. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, device, and storage medium, which can solve the problem of data skew.
[0005] In a first aspect, embodiments of this application provide a data processing method, which includes:
[0006] Obtain the fields in the target data table and the element information of each field; wherein, the skew value of the target data table is greater than a preset threshold, and the skew value is used to describe the degree of data skew of the data table;
[0007] According to the element information of each field, determine the element distribution value of each field, and the element distribution value is used to describe the uniformity of the elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, where N is a positive integer;
[0008] Determine the field corresponding to the minimum value in the element distribution values as the distribution key of the target distribution table.
[0009] In a second aspect, embodiments of this application provide a data processing apparatus, and the data processing apparatus includes:
[0010] An obtaining module, configured to obtain the fields in the target data table and the element information of each field; wherein, the skew value of the target data table is greater than a preset threshold, and the skew value is used to describe the degree of data skew of the data table;
[0011] A determining module, configured to determine the element distribution value of each field according to the element information of each field, and the element distribution value is used to describe the uniformity of the elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, where N is a positive integer;
[0012] The determining module is further configured to determine the field corresponding to the minimum value in the element distribution values as the distribution key of the target distribution table.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0015] In the embodiments of the present application, by obtaining the fields in the target data table with tilt values greater than a preset threshold and the element information of each field; wherein, the tilt value is used to describe the degree of data tilt of the data table, so the data tilt degree of the target data table is relatively serious. Then, the element distribution value of each field is determined according to the element information of each field, and the element information at least includes the quantity value of the Nth element in each field. The element distribution value calculated in this way is used to characterize the uniformity of the elements included in each field. The smaller the element distribution value, the fewer the same elements in this column. Since the distribution of the distributed data table on each machine node is mainly allocated according to the hash of the content of the table distribution key, the same hash value will be distributed on the same machine, so the fewer the same contents of the distribution key, the better. Therefore, finally, the field corresponding to the minimum value in the element distribution values is determined as the distribution key of the target distribution table, which can effectively solve the problem of data tilt and make the file distribution of the target data table more uniform. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0017] Figure 1 is a schematic diagram of data tilt provided by an embodiment of the present application;
[0018] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0020] Figure 4 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0021] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only provided to provide a better understanding of the present application by showing examples of the present application.
[0022] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the elements.
[0023] The technical terms related to the present application will be briefly introduced below.
[0024] A distributed database generally uses smaller computer systems. Each computer can be placed separately in a location, and each computer may have a complete copy or a partial copy, and has its own local database. Many computers located in different locations are connected to each other through a network to jointly form a complete, globally logically centralized and physically distributed large database.
[0025] Data skew. For a cluster system, generally the cache is distributed, that is, different nodes are responsible for a certain range of cached data. The phenomenon that the dispersion degree of the cached data is insufficient, resulting in a large amount of cached data being concentrated on one or several service nodes, is generally referred to as data skew.
[0026] Data disaster tolerance refers to the process of establishing a data system in a different location. To protect data security and improve the continuous availability of data, an enterprise needs to consider multiple aspects such as redundant structure, data backup, and fault warning, and copy the necessary files of the database to storage devices.
[0027] A distribution key is a column (or a group of columns) used to determine the database partition where a specific data row is stored.
[0028] A hash function transforms an input of any length into an output of a fixed length through a hashing algorithm. This output is the hash value. Such a transformation is a compression mapping, that is, the space of hash values is generally much smaller than the space of the input.
[0029] The data processing method provided by the embodiments of this application can be applied to at least the following application scenarios, which will be described below.
[0030] With the continuous advancement of the informatization and digitalization processes, enterprise data not only carries the internal data of the unit itself but also includes a vast amount of user usage and business data. Due to the poor scalability of traditional single-database media such as database appliances, as the digitalization process of enterprise data accelerates, when the data volume reaches the storage bottleneck, it will limit the growth of enterprise data. Distributed databases have greatly solved the problem of single-node data capacity and won the favor of many large enterprises with advantages such as good scalability and high business performance.
[0031] Since most distributed databases use the primary and standby mode for data disaster recovery and generally have multiple backup copies, the size of the underlying total physical space occupied by single-table data grows rapidly. In a distributed database, data distribution mainly uses the hash value (HASH) of the distribution key value. When the redundancy of some values of the distribution key is too high, the data on the nodes where these values are located will be much higher than that of other nodes.
[0032] This makes it that as the amount of table data increases, the quality of the distribution key selection determines whether the data is reasonably distributed among the machines. In general development operations, developers usually specify the distribution key in an empirical way. In actual data development operations, there is a widespread phenomenon of large amounts of file data skew in the tables developed by developers.
[0033] As Figure 1 shown, the file corresponding to the data table includes: the first part, the second part, the third part,..., the Wth part, and the data volume of each part varies greatly, showing a phenomenon of data skew.
[0034] Therefore, since there are many nodes in a distributed database, the performance degradation of a single node will greatly affect the overall performance. In actual data operations, the skew of data table files is one of the main reasons for the degradation of database query performance.
[0035] Figure 2 is a flowchart of a data processing method provided by the embodiments of this application.
[0036] As Figure 2As shown, the data processing method may include step 210-step 230. This method is applied to a data processing device, specifically as follows:
[0037] Step 210, obtain the fields in the target data table and the element information of each field; wherein, the skew value of the target data table is greater than a preset threshold, and the skew value is used to describe the degree of data skew of the data table.
[0038] Step 220, determine the element distribution value of each field according to the element information of each field. The element distribution value is used to describe the uniformity of the elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, where N is a positive integer.
[0039] Step 230, determine the distribution key of the target distribution table as the field corresponding to the minimum value in the element distribution value.
[0040] In the embodiment of the present application, by obtaining the fields in the target data table with a skew value greater than the preset threshold and the element information of each field; wherein, the skew value is used to describe the degree of data skew of the data table, so the data skew degree of the target data table is relatively serious. Then, determine the element distribution value of each field according to the element information of each field. The element information at least includes the quantity value of the Nth element in each field. The element distribution value calculated in this way is used to characterize the uniformity of the elements included in each field. The smaller the element distribution value, the fewer the same elements in this column. Since the distribution of the distributed data table on each machine node is mainly allocated according to the HASH of the table distribution key content, the same HASH value will be distributed on the same machine. Therefore, the fewer the same contents of the distribution key, the better. Thus, finally, determine the distribution key of the target distribution table as the field corresponding to the minimum value in the element distribution value, which can effectively solve the problem of data skew and make the file distribution of the target data table more uniform.
[0041] Next, the content of step 210-step 230 will be described separately:
[0042] Regarding step 210.
[0043] Step 210, obtain the fields in the target data table and the element information of each field; wherein, the skew value of the target data table is greater than a preset threshold, and the skew value is used to describe the degree of data skew of the data table.
[0044] The element information at least includes: the quantity value of the Nth element in each field, where N is a positive integer.
[0045] Among them, the preset threshold may be 2.6.
[0046] Among them, the target data table includes multiple candidate fields. Step 210 may specifically include the following steps:
[0047] Determine the frequency value of each candidate field appearing in the target data table;
[0048] Select fields from the target data table according to the frequency value.
[0049] Here, according to the number of occurrences of each candidate field, determine the frequency value of each candidate field appearing in the target data table; then, according to the frequency value, select fields for the frequently used fields in the skewed table, so as to calculate the element distribution value of each field separately later, and select the field with a smaller element distribution value as the distribution key.
[0050] In a possible embodiment, before step 210, the following steps may further be included:
[0051] Obtain multiple data tables, where the data tables include file information, each data table corresponds to multiple files, and the file information includes: file identification information, total file size, and file size of each file;
[0052] Calculate the skew value of each data table according to the file information;
[0053] Determine the target data table from the multiple data tables according to the skew value.
[0054] Among them, the file information may include: file identification information, total file size, and file size of each file (including: maximum file size and minimum file size).
[0055] For example: The file information of Table A may include: table schema name (such as bss), file identification information (table name), total file size 8839246872 Byte, maximum file 3131375904 Byte, minimum file 78999176 Byte.
[0056] Then, calculate the skew value of each data table according to the file information and sort, and filter out the tables with higher skew values for processing, that is, the target data table. So as to calculate the element distribution value of the table field elements later for the severely skewed target data table, and filter out a new suitable distribution key.
[0057] Among them, in the step of determining the target data table from the multiple data tables according to the skew value, it includes: determining the data table with a skew value greater than the preset threshold as the target data table.
[0058] Here, by obtaining multiple data tables, calculating the skew value of each data table according to the file information of the data tables, and according to the skew value, the severely skewed target data table can be determined from the multiple data tables.
[0059] Among them, in the step of calculating the tilt value of each data table according to the file information mentioned above, the following steps can be specifically included:
[0060] For each data table, the following steps are performed separately:
[0061] Calculate the ratio of the size of the first file to the size of the second file, where the first file is the largest file among multiple files, and the second file is the smallest file among multiple files;
[0062] According to the preset level relationship, determine the file level corresponding to the total file size; the preset level relationship includes multiple groups of corresponding total file sizes and file levels;
[0063] Determine the tilt value according to the ratio and the file level.
[0064] Calculating the ratio of the size of the first file to the size of the second file is to collect the sizes of all data files of the table in the cluster-related instances and calculate the ratio Max / Min of the first file to the second file size.
[0065] Determining the file level corresponding to the total file size according to the preset level relationship is to collect the sizes of all data files of the table in the cluster-related instances, calculate the total file size of the single table, and divide it into several levels according to the actual situation.
[0066] For example, define three levels: Level 1 (below 1G), Level 2 (1 - 10G), Level 3 (above 10G).
[0067] For example: the total file size is 8839246872 Byte, and the file level corresponding to the total file size is 2.
[0068] Determine the tilt value and sort according to the ratio and the file level. For example, the tilt value of Table A is Rate = 0.8 * 39.6 + 0.2 * 2 = 32.08. Then calculate the tilt index for other tables in turn and sort them.
[0069] Thus, through the ratio of the largest file to the smallest file in each data table and the file level of the total file, the tilt value of the data table can be determined quickly and accurately.
[0070] Among them, in the step of determining the tilt value according to the ratio and the file level mentioned above, the following steps can be specifically included:
[0071] Obtain the first weight value and the second weight value, where the first weight value is greater than the second weight value;
[0072] According to the first weight value and the second weight value, perform weighted calculation on the ratio and the file level to obtain the tilt value.
[0073] Based on the first weight value and the second weight value, a weighted calculation is performed on the comparison value and the file level to obtain the tilt value, which can be specifically determined by the following formula:
[0074] Rate = 0.8 * Max / Min + 0.2 * Level, and other statistical algorithms can also be used.
[0075] Among them, Rate is the tilt value; Level is the file level corresponding to the total file size
[0076] 0.8 is the first weight value, and 0.2 is the second weight value;
[0077] Max / Min is the ratio of the size of the first file to the size of the second file.
[0078] Here, by assigning the first weight value to the ratio and the second weight value to the file level, and the first weight value is greater than the second weight value; the importance of the ratio of the largest file to the smallest file in the data table to the tilt value can be emphasized. Based on the first weight value and the second weight value, a weighted calculation of the comparison value and the file level can obtain an accurate tilt value.
[0079] It involves step 220.
[0080] According to the element information of each field, determine the element distribution value of each field. The element distribution value is used to describe the uniformity of the elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, and N is a positive integer.
[0081] The main principle of the element distribution value is that the distribution of the distributed database table on each machine node is mainly allocated according to the HASH of the table distribution key content. The same HASH value will be distributed on the same machine. Therefore, the less the same content of the distribution key, the better.
[0082] Calculate the element distribution values of the common fields in the target table. For example, the common fields included in Table A are: Field 1, Field 2, and Field F. According to the element information of each field, the determined element distribution values of each field are: the element distribution value of Field 1 is 1.0004, the element distribution value of Field 2 is 1.3156, and the other fields are calculated in turn.
[0083] In a possible embodiment, step 220 may specifically include the following steps:
[0084] According to the element information, determine the quantity value of all elements in each field;
[0085] According to the quantity value of all elements and the element information, determine the element distribution value of each field respectively.
[0086] Among them, the element distribution value mainly describes the uniformity of elements in the same field, and the calculation formula can be customized or calculated based on statistical algorithms.
[0087] For example, it can be through the following formula (1):
[0088] Where xt is the number of the same element in the field, n is the number of elements in the field, which can be valued according to actual needs, and total is the total number of records in the field.
[0089] Among them, what formula (1) calculates is the sum of the proportions of the total number of the first n same elements after grouping and sorting in the total number of records under the current field. Correspondingly, the smaller this value is, the better, indicating that there are fewer same elements in this column.
[0090] For example, the results of the element distribution values of each field can be: the element distribution value f1(x) of the first field, the element distribution value f2(x) of the second field,....
[0091] Thus, according to the element information, the quantity value of all elements in each field is determined. Based on the quantity value of all elements and the quantity value of the Nth element in each field, the uniformity of the elements included in each field can be effectively described to accurately determine the element distribution value of each field respectively.
[0092] In a possible embodiment, step 220 may specifically include the following steps:
[0093] According to the element information, determine the average value of the number of elements;
[0094] According to the average value of the number and the element information, determine the element distribution value of each field respectively.
[0095] The variance of the number of the same kind of elements can be directly calculated, that is, it can be determined through the following formula (2):
[0096]
[0097] Where xi is the number of the same element in the field, is the average value of the number of different elements.
[0098] According to the average value of the number and the element information xi, determine the element distribution value of each field respectively.
[0099] Thus, according to the average value of the number of elements and the quantity value of the Nth element in each field, the uniformity of the elements included in each field can be effectively described to accurately determine the element distribution value of each field respectively.
[0100] It involves step 230.
[0101] Determine the field corresponding to the minimum value in the element distribution value as the distribution key of the target distribution table.
[0102] For example: The element distribution values of each field are as follows: the element distribution value of field 1 is 1.0004, the element distribution value of field 2 is 1.3156,..., and the element distribution value of field F is 1.0512. The element distribution value is the smallest for field 1. Therefore, the best distribution key obtained should be field 1. Finally, perform the distribution key switch, and other target data tables are processed similarly.
[0103] Here, by calculating the element distribution values of the target data table and performing an automated distribution key switch, the file distribution corresponding to the target data table is optimized, while significantly reducing the manual and time costs. Thus, by reducing the data table skew rate of the distributed database, the availability of the distributed database is effectively enhanced. Moreover, without manual intervention, the work pressure of data developers is effectively reduced.
[0104] In the data processing method provided in this application, by obtaining the fields in the target data table with a skew value greater than the preset threshold and the element information of each field; where the skew value is used to describe the degree of data skew of the data table, so the data skew degree of the target data table is relatively serious. Then, determine the element distribution value of each field according to the element information of each field, and the element information includes at least the quantity value of the Nth element in each field. The element distribution value calculated in this way is used to characterize the uniformity of the elements included in each field. The smaller the element distribution value, the fewer the same elements in this column. Since the distribution of the distributed data table on each machine node is mainly allocated according to the HASH of the table distribution key content, the same HASH value will be distributed on the same machine. Therefore, the less the same content of the distribution key, the better. Thus, finally, determine the field corresponding to the minimum value in the element distribution value as the distribution key of the target distribution table, which can effectively solve the problem of data skew and make the file distribution of the target data table more uniform.
[0105] Based on the above Figure 2 shown data processing method, an embodiment of this application further provides a data processing device, as Figure 3 shown, the data processing device 300 may include:
[0106] An obtaining module 310, configured to obtain the fields in the target data table and the element information of each field; where the skew value of the target data table is greater than the preset threshold, and the skew value is used to describe the degree of data skew of the data table.
[0107] A determining module 320, configured to determine the element distribution value of each field according to the element information of each field, and the element distribution value is used to describe the uniformity of the elements included in each field; the element information includes at least: the quantity value of the Nth element in each field, and N is a positive integer.
[0108] The determining module 320 is further configured to determine the field corresponding to the minimum value in the element distribution values as the distribution key of the target distribution table.
[0109] In a possible embodiment, the determining module 320 is specifically configured to:
[0110] Determine the quantity value of all elements in each field according to the element information;
[0111] Determine the element distribution value of each field respectively according to the quantity value of all elements and the element information.
[0112] In a possible embodiment, the determining module 320 is specifically configured to:
[0113] Determine the average quantity value of the elements according to the element information;
[0114] Determine the element distribution value of each field respectively according to the average quantity value and the element information.
[0115] In a possible embodiment, the obtaining module 310 is specifically configured to:
[0116] Determine the frequency value of the occurrence of each candidate field in the target data table;
[0117] Select fields from the target data table according to the frequency value.
[0118] In a possible embodiment, the obtaining module 310 is further configured to:
[0119] Obtain a plurality of data tables, the data tables include file information, each data table corresponds to a plurality of files, and the file information includes: file identification information, total file size, and file size of each file;
[0120] The data processing device 300 may further include:
[0121] A calculation module, configured to calculate the skewness value of each data table according to the file information.
[0122] The determining module 320 is further configured to: determine the target data table from the plurality of data tables according to the skewness value.
[0123] In a possible embodiment, the calculation module is specifically configured to:
[0124] For each data table, respectively perform the following steps:
[0125] Calculate the ratio of the size of the first file to the size of the second file, where the first file is the largest file among the plurality of files, and the second file is the smallest file among the plurality of files;
[0126] Determine the file level corresponding to the total file size according to the preset level relationship; the preset level relationship includes multiple groups of corresponding total file sizes and file levels;
[0127] Determine the tilt value according to the ratio and the file level.
[0128] In a possible embodiment, the calculation module is specifically configured to:
[0129] Obtain a first weight value and a second weight value, where the first weight value is greater than the second weight value;
[0130] Perform weighted calculation on the ratio and the file level according to the first weight value and the second weight value to obtain the tilt value.
[0131] In the embodiment of the present application, by obtaining the fields in the target data table whose tilt value is greater than the preset threshold, and the element information of each field; wherein, the tilt value is used to describe the degree of data tilt of the data table, so the data tilt degree of the target data table is relatively serious. Then, determine the element distribution value of each field according to the element information of each field, and the element information at least includes the quantity value of the Nth element in each field. The element distribution value calculated in this way is used to characterize the uniformity of the elements included in each field. The smaller the element distribution value, the fewer the same elements in this column. Since the distribution of the distributed data table on each machine node is mainly allocated according to the HASH of the table distribution key content, the same HASH value will be distributed on the same machine, so the fewer the same contents of the distribution key, the better. Therefore, finally, the field corresponding to the minimum value in the element distribution value is determined as the distribution key of the target distribution table, which can effectively solve the problem of data tilt and make the file distribution of the target data table more uniform.
[0132] Figure 4 Shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application.
[0133] The electronic device may include a processor 401 and a memory 402 storing computer program instructions.
[0134] Specifically, the above-mentioned processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0135] The memory 402 may include a mass storage for data or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 402 may include removable or non-removable (or fixed) media. Where appropriate, the memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, the memory 402 is a non-volatile solid-state memory. In a particular embodiment, the memory 402 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0136] The processor 401 reads and executes the computer program instructions stored in the memory 402 to implement any of the data processing methods in the illustrated embodiments.
[0137] In one example, the electronic device may further include a communication interface 403 and a bus 410. As shown, Figure 4 the processor 401, the memory 402, and the communication interface 403 are connected via the bus 410 to complete communication with each other.
[0138] The communication interface 403 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application.
[0139] The bus 410 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a hypertransport (HT) interconnect, an industry standard architecture (ISA) bus, an infinite bandwidth interconnect, a low pin count (LPC) bus, a memory bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus, or other suitable bus, or a combination of two or more of these. Where appropriate, the bus 410 may include one or more buses. Although the embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.
[0140] The electronic device can execute the data processing method in the embodiments of the present application, so as to implement the combination with Figure 2 the described data processing method.
[0141] In addition, in combination with the data processing method in the above embodiments, the embodiments of the present application can provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, the Figure 2 data processing method is implemented.
[0142] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0143] The functional blocks shown in the above structure block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0144] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0145] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. A data processing method, characterized in that, the method includes: obtaining fields in a target data table and element information of each field; wherein, the skew value of the target data table is greater than a preset threshold, and the skew value is used to describe the degree of data skew of the data table; determining an element distribution value of each field according to the element information of each field, and the element distribution value is used to describe the uniformity of elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, and N is a positive integer; determining the field corresponding to the minimum value in the element distribution values as the distribution key of the target distribution table; wherein, before obtaining the fields in the target data table, the method further includes: obtaining a plurality of data tables, the data tables include file information, each data table corresponds to a plurality of files, and the file information includes: file identification information, total file size, and file size of each file; calculating the skew value of each data table according to the file information; determining the target data table from the plurality of data tables according to the skew value; wherein, calculating the skew value of each data table according to the file information includes: for each data table, respectively performing the following steps: calculating the ratio of the size of the first file to the size of the second file, where the first file is the largest file among the plurality of files, and the second file is the smallest file among the plurality of files; determining the file level corresponding to the total file size according to a preset level relationship; the preset level relationship includes multiple groups of corresponding total file sizes and file levels; determining the skew value according to the ratio and the file level; wherein, determining the skew value according to the ratio and the file level includes: obtaining a first weight value and a second weight value, and the first weight value is greater than the second weight value; performing weighted calculation on the ratio and the file level according to the first weight value and the second weight value to obtain the skew value.
2. The method according to claim 1, characterized in that, determining the element distribution value of each field according to the element information of each field includes: determining the quantity value of all elements in each field according to the element information; determining the element distribution value of each field respectively according to the quantity value of all elements and the element information.
3. The method according to claim 1, characterized in that, determining the element distribution value of each field according to the element information of each field includes: determining the average quantity value of the elements according to the element information; determining the element distribution value of each field respectively according to the average quantity value and the element information.
4. The method according to any one of claims 1-3, characterized in that, the target data table includes a plurality of candidate fields, and obtaining the fields in the target data table includes: determining the frequency value of each candidate field appearing in the target data table; selecting the fields from the target data table according to the frequency value.
5. A data processing device, characterized in that, the data processing device includes: an acquisition module, configured to acquire fields in a target data table and element information of each field; wherein, a tilt value of the target data table is greater than a preset threshold, and the tilt value is used to describe the degree of data tilt of the data table; a determination module, configured to determine an element distribution value of each field according to the element information of each field, and the element distribution value is used to describe the uniformity of elements included in each field; the element information at least includes: the quantity value of the Nth element in each field, where N is a positive integer; the determination module is further configured to determine the field corresponding to the minimum value in the element distribution values as the distribution key of the target distribution table; wherein, the acquisition module is further configured to: acquire a plurality of data tables, the data tables include file information, each data table corresponds to a plurality of files, and the file information includes: file identification information, total file size, and file size of each file; wherein, the data processing device further includes: a calculation module, configured to calculate the tilt value of each data table according to the file information; wherein, the determination module is further configured to: determine the target data table from the plurality of data tables according to the tilt value; wherein, the calculation module is specifically configured to: for each data table, respectively perform the following steps: calculate the ratio of the size of the first file to the size of the second file, the first file is the largest file among the plurality of files, and the second file is the smallest file among the plurality of files; determine the file level corresponding to the total file size according to a preset level relationship; the preset level relationship includes multiple groups of corresponding total file sizes and file levels; determine the tilt value according to the ratio and the file level; acquire a first weight value and a second weight value, the first weight value is greater than the second weight value; perform weighted calculation on the ratio and the file level according to the first weight value and the second weight value to obtain the tilt value.
6. An electronic device, characterized in that, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the data processing method described in any one of claims 1-4 is implemented.
7. A computer-readable storage medium, characterized in that, computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the data processing method described in any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Partitioning method and partitioning device based on database middleware and readable storage medium
CN111274028A
Electronic equipment and data processing method
CN114020747A