Data set repartitioning method and device, electronic equipment and storage medium
By sampling and estimating the data set to be processed, the total number of bytes occupied by them is determined, and the data set is repartitioned based on this value, which solves the problems of partitioning in the prior art that fails to consider the data itself, has high additional overhead and is highly tuning difficult, and realizes an efficient and low-cost data partitioning solution.
Patent Information
- Application Number
- CN202311658222.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
In the adaptive resource adjustment solution of existing cloud computing products, partitioning fails to consider the data itself, resulting in large additional overhead and high tuning difficulty.
By sampling and estimating the data set to be processed, the total number of bytes occupied by them is determined, and the data set is repartitioned based on this value.
Partitioning is realized based on the data size of the data set, reducing system complexity and cost, and improving time efficiency and resource utilization.
Smart Images

Figure CN120104296A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of distributed data processing technology, and in particular to a data set repartitioning method and device, an electronic device, and a storage medium. Background Art
[0002] At present, with the development of cloud computing, the adaptive resource adjustment function can automatically recommend the appropriate number of resources according to the user's current task requirements and cluster configuration, and is becoming more and more popular among service providers and users. The adaptive resource adjustment function allows the system to evaluate the type, scale and required resources of the task when the user submits a task, and recommend a suitable number of resources to the user. No similar solution for obtaining the number of partitions by sampling estimation has been found.
[0003] In cloud computing products in related technologies, adaptive resource adjustment solutions are generally implemented based on real-time monitoring of resources in the YARN cluster. This method has the following technical problems:
[0004] 1. Partitioning fails to take into account the data's own situation: The existing solution is based only on the resources and configuration of the cluster environment, and does not partition the data. Regardless of the size of the data, it is estimated according to the configuration of the environment.
[0005] 2. Large additional overhead: Implementing automatic resource adjustment solutions requires certain overhead, which may involve the transformation and upgrade of clusters and resource monitoring systems, as well as the design and development of related algorithms and strategies. These overheads may increase system complexity and cost.
[0006] 3. High tuning difficulty: Automatic resource adjustment involves the design and tuning of algorithms and strategies. How to accurately assess the resource requirements of tasks, reasonably allocate resources, and determine adjustment strategies are challenging issues. Sufficient algorithm optimization and testing are required to ensure the stability and performance of the system.
[0007] Therefore, there is at least one of the above technical problems in the related art. Summary of the invention
[0008] The present application provides a data set repartitioning method and device, an electronic device and a storage medium to at least solve at least one technical problem existing in the related art.
[0009] According to one aspect of an embodiment of the present application, a method for repartitioning a data set is provided, comprising:
[0010] Get the data set to be processed;
[0011] Determine the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed;
[0012] The data set to be processed is repartitioned according to the total number of occupied bytes.
[0013] Optionally, as in the aforementioned method, determining the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed includes:
[0014] Sampling the data set to be processed according to a preset sampling method to obtain a plurality of sampled data;
[0015] Determine the number of bytes of a single data piece occupied by each of the plurality of sampled data pieces;
[0016] Based on the number of bytes occupied by each piece of sampled data, the total number of bytes occupied by the data set to be processed is determined.
[0017] Optionally, as in the aforementioned method, determining the number of bytes of a single data piece occupied by each of the multiple pieces of sampled data comprises:
[0018] Determine all fields included in the sampled data;
[0019] Based on the field type of each field in all the fields, determine the number of single field bytes occupied by each field;
[0020] The number of bytes of a single data item occupied by the sampled data is determined based on the number of bytes of a single field occupied by each field.
[0021] Optionally, as in the aforementioned method, sampling the to-be-processed data set according to a preset sampling method to obtain a plurality of sampled data includes:
[0022] In the data set to be processed, a preset number of selections are randomly performed to obtain the plurality of sampling data.
[0023] Optionally, as in the aforementioned method, repartitioning the to-be-processed data set according to the total number of occupied bytes includes:
[0024] Determining a target data source of the data set to be processed;
[0025] Determine the average number of bytes of each piece of data in the to-be-processed data set based on the number of bytes of a single piece of data occupied by each piece of sampled data in the plurality of pieces of sampled data;
[0026] In the case where the target data source is a file type data source, determining a target partition size corresponding to the target data source;
[0027] Determine the target partition number corresponding to the data set to be processed based on the total number of occupied bytes, the average number of bytes, and the target partition size;
[0028] The data set to be processed is repartitioned according to the target number of partitions.
[0029] Optionally, as in the aforementioned method, repartitioning the to-be-processed data set according to the total number of occupied bytes includes:
[0030] Determining a target data source of the data set to be processed;
[0031] Determine the average number of bytes of each piece of data in the to-be-processed data set based on the number of bytes of a single piece of data occupied by each piece of sampled data in the plurality of pieces of sampled data;
[0032] When the target data source is a relational data source, configuring to obtain a specified partition size of the data set to be processed;
[0033] Determine the designated partition number corresponding to the data set to be processed based on the total number of occupied bytes, the average number of bytes, and the designated partition size;
[0034] The data set to be processed is repartitioned according to the specified number of partitions.
[0035] Optionally, as in the aforementioned method, after obtaining the data set to be processed, the method further includes:
[0036] Determining available computing resources corresponding to the data set to be processed;
[0037] Dividing the available computing resources to obtain single-thread computing resources corresponding to each thread and a target number of all single-thread computing resources, wherein the sum of all single-thread computing resources is the available computing resources;
[0038] The data set to be processed is repartitioned according to the target number.
[0039] According to another aspect of an embodiment of the present application, a data set repartitioning device is also provided, including:
[0040] An acquisition module is used to obtain the data set to be processed;
[0041] A determination module, configured to determine the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed;
[0042] The repartitioning module is used to repartition the data set to be processed according to the total number of occupied bytes.
[0043] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; wherein the memory is used to store a computer program; and the processor is used to execute the method steps in any of the above embodiments by running the computer program stored in the memory.
[0044] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the method steps in any of the above embodiments when executed.
[0045] In the embodiment of the present application, a sampling estimation is performed on the data set to be processed to determine the total number of occupied bytes of the data set to be processed; the data set to be processed is repartitioned according to the total number of occupied bytes, thereby realizing the repartitioning of the data set to be processed according to the size of the data set, and since the sampling estimation can calculate the total number of occupied bytes and repartition with a lower time complexity, there is no need to traverse the entire data set to be processed, which can improve time efficiency; in addition, compared with completely processing the entire data set, the sampling estimation only needs to process the sampled data, which can save resource consumption, such as computing resources and storage space; and the sampling estimation can be estimated and repartitioned when the data set has not been fully generated or changes rapidly; finally, only the sampled data itself is focused on, and the acquired data is analyzed and calculated, so as not to rely on external services; through the method of the above embodiment, since only the data itself is focused on, rather than partitioning according to resources, the technical problems existing in the related art that the partitioning fails to consider the data itself, the additional overhead is large, and the tuning difficulty is high can be effectively overcome. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0048] Figure 1 is a schematic diagram of a hardware environment of an optional data set repartitioning method according to an embodiment of the present application;
[0049] Figure 2is a flowchart of an optional data set repartitioning method according to an embodiment of the present application;
[0050] Figure 3 is a flowchart of an optional data set repartitioning method according to another embodiment of the present application;
[0051] Figure 4 is a flowchart of an optional data set repartitioning method according to another embodiment of the present application;
[0052] Figure 5 is a flowchart of an optional data set repartitioning method according to another embodiment of the present application;
[0053] Figure 6 is a structural block diagram of an optional data set repartitioning device according to an embodiment of the present application;
[0054] Figure 7 It is a structural block diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0056] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0057] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretation:
[0058] sparkRDD: It is the core data structure of spark and is used to represent distributed collections. It is a mutable, partitioned, fault-tolerant, parallel computing-oriented data collection.
[0059] According to one aspect of an embodiment of the present application, a method for repartitioning a data set is provided. Optionally, in this embodiment, the method for repartitioning a data set can be applied to Figure 1 In the hardware environment composed of terminal 1402 and server 1404 shown in FIG. Figure 1 As shown, server 1404 is connected to terminal 1402 via a network, and can be used to provide services (such as game services, application services, etc.) for the terminal or a client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 1404.
[0060] The above network may include but is not limited to at least one of the following: wired network, wireless network. The above wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the above wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal may not be limited to a PC, a mobile phone, a tablet computer, etc.
[0061] The data set repartitioning method of the embodiment of the present application can be executed by a server, or by a terminal, or by both a server and a terminal. The terminal can also execute the data set repartitioning method of the embodiment of the present application by a client installed thereon.
[0062] Taking the data set repartitioning method in this embodiment executed by the server as an example, Figure 2 A data set repartitioning method provided in an embodiment of the present application includes the following steps:
[0063] Step S101, obtaining a data set to be processed.
[0064] The data set repartitioning method in this embodiment can be applied to a scenario where a partitioning operation needs to be performed on a data set to be processed, so that data of each partition obtained by the partitioning operation can be processed by multiple threads at the same time later.
[0065] Specifically, all or part of the data to be processed can be determined in the sparkRDD and used as the data set to be processed. In addition, the data set to be processed can be a data set that has not been completed or is changing.
[0066] Step S102, determining the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed.
[0067] Specifically, after the data set to be processed is determined, the size of the data set to be processed must be determined so as to facilitate processing after the data set to be processed is stored on a disk.
[0068] In this embodiment, the total number of bytes occupied by the data set to be processed is determined by performing an abstract estimation on the data set to be processed.
[0069] Optionally, multiple records may be selected from the data set to be processed to estimate the average number of bytes of each record in the data set to be processed, and then the total number of bytes occupied by the data set to be processed may be determined according to the total number of records in the data set to be processed.
[0070] Step S103, repartitioning the data set to be processed according to the total number of occupied bytes.
[0071] Specifically, after determining the total number of occupied bytes, the data set to be processed can be repartitioned according to the total number of occupied bytes. Generally, the sizes of the partitions obtained by repartitioning are similar, that is, the error in the number of bytes of each partition is within a preset range (for example, 1MB, 5MB, etc.).
[0072] Furthermore, after determining the total number of occupied bytes, the size of each partition can be determined, and then the number of partitions can be determined, so that the data set to be processed can be repartitioned based on the number of partitions.
[0073] Repartitioning can be done by re-storing the data in the original partition of the dataset to be processed, and storing it in the new disk determined by the repartitioning. Optionally, repartitioning can be achieved through shuffle, which refers to the process of redistributing and recombining data between different partitions of a dataset. It usually occurs in operations that require data rearrangement or data aggregation across nodes or tasks. The process of redistributing and reorganizing data can help achieve parallel computing and aggregation operations of data.
[0074] According to the method of this embodiment, the total number of bytes occupied by the data set to be processed is determined by sampling and estimating the data set to be processed; and the data set to be processed is repartitioned according to the total number of bytes occupied, thereby realizing the repartitioning of the data set to be processed according to the size of the data set. Moreover, since the sampling estimation can calculate the total number of bytes occupied and repartition with a lower time complexity, it is not necessary to traverse the entire data set to be processed, which can improve the time efficiency. In addition, compared with completely processing the entire data set, the sampling estimation only needs to process the sampled data, which can save resource consumption, such as computing resources and storage space. Moreover, the sampling estimation can be estimated and repartitioned when the data set has not been completely generated or changes rapidly. Finally, only the sampled data itself is focused on, and the acquired data is analyzed and calculated, so as not to rely on external services. According to the method of the above embodiment, only the data itself is focused on, rather than partitioning according to resources, so as to effectively overcome the technical problems existing in the related art that the partitioning fails to consider the data itself, the additional overhead is large, and the tuning difficulty is high.
[0075] like Figure 3 As shown, as an optional embodiment, as in the above method, the step S102 determines the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed, including the following steps:
[0076] Step S201: sampling a plurality of sampled data in a data set to be processed according to a preset sampling method.
[0077] As an optional embodiment, as in the aforementioned method, sampling a plurality of sampled data in the data set to be processed according to a preset sampling method includes:
[0078] In the data set to be processed, a preset number of selections are randomly performed to obtain multiple sampling data.
[0079] That is to say, a preset number of selections can be performed in the data set to be processed by random selection to obtain multiple sampling data.
[0080] Furthermore, random selection can be achieved through random data sampling algorithms, which construct a representative sample set by randomly selecting a certain number of samples in the data set. The probability of each sample being selected is the same, so that each sample has an equal chance of being selected.
[0081] The preset number of times can be determined according to the total number of data items in the data set to be processed. Generally, the more data sets to be processed, the greater the preset number of times.
[0082] For example, the preset number of times may be 1%, 2%, etc. of the data set to be processed. Furthermore, the number of preset times may be selected according to actual applications and is not limited here.
[0083] Step S202, determining the number of bytes of a single data piece occupied by each of the multiple sampling data pieces.
[0084] Specifically, after a plurality of sampling data are determined, the number of bytes occupied by each sampling data may be counted, so as to determine the number of bytes of a single data piece occupied by each sampling data.
[0085] Step S203: determining the total number of bytes occupied by the data set to be processed based on the number of bytes occupied by each piece of sampled data.
[0086] After determining the number of bytes occupied by each piece of sampled data, the average number of bytes of multiple pieces of sampled data can be determined, and the average number of bytes can be used as the average number of bytes of the data set to be processed. The total number of bytes occupied by the data set to be processed can be determined by multiplying it by the number of data included in the data set to be processed.
[0087] With the method of this embodiment, the total occupancy times of the data set to be processed can be quickly determined through sampling, and the statistical efficiency can be improved.
[0088] like Figure 4 As shown, as an optional embodiment, as in the above method, the step S202 determines the number of bytes of a single data piece occupied by each of the multiple sampling data pieces, including the following steps:
[0089] Step S301, determining all fields included in the sampled data.
[0090] Specifically, after a plurality of sampled data are obtained by random selection, the process corresponding to steps S301 to S303 may be executed for each sampled data to determine the number of bytes of a single data piece occupied by each sampled data piece.
[0091] That is to say, after determining each piece of sampled data, all fields corresponding to any one of the sampled data can be determined; among them, the basic data types of Java may include: byte, bool, short, int, float, long, double, char.
[0092] Step S302: Based on the field type of each field in all the fields, determine the number of single field bytes occupied by each field.
[0093] After the field type of each field is determined, generally, the byte length occupied by each field type is fixed, so the number of single field bytes occupied by each field can be determined based on the field type of each field.
[0094] For example, the type with byte size of 1: byte, bool;
[0095] Types with a byte size of 2: short, char;
[0096] Types with byte size of 4: int, float;
[0097] Types with byte size 8: long, double.
[0098] Step S303: Based on the number of bytes of a single field occupied by each field, determine the number of bytes of a single data record occupied by the sampled data.
[0099] Specifically, after determining the number of bytes of a single field occupied by each field, the number of bytes of a single data record occupied by the sampled data can be calculated by adding up the number of bytes of all single fields.
[0100] Through the method of this embodiment, the number of bytes of a single data piece occupied by each sampling data piece can be quickly calculated, thereby effectively improving the calculation efficiency.
[0101] As an optional embodiment, as in the above method, step S103 repartitions the data set to be processed according to the total number of occupied bytes, including the following steps:
[0102] Step 401, determining the target data source of the data set to be processed.
[0103] Specifically, when the data set to be processed is obtained, the target data source of the data set to be processed can be determined.
[0104] Optionally, the target data source may include: File type data source: hive / / hdfs / cos / ks3 / obs, etc., based on the size of the overflow file buffer. Relational data source: mysql / oracle / sqlserver / pgsql, etc., based on the data file size configuration in the underlying table space.
[0105] Step 402: based on the number of bytes occupied by each piece of sampled data in the plurality of pieces of sampled data, determine the average number of bytes of each piece of data in the data set to be processed.
[0106] Specifically, after determining the number of bytes occupied by each piece of sampled data, the average number of bytes of multiple pieces of data may be calculated to determine the average number of bytes of the multiple pieces of sampled data, and the average number of bytes may be used as the average number of bytes of the data set to be processed.
[0107] Step 403, when the target data source is a file type data source, determine the target partition size corresponding to the target data source.
[0108] When the target data source is a file type data source, the file type data source generally has a default data block (i.e., block) size, for example, the block size on HDFS is 128 MB by default. Then the target partition size corresponding to the target data source can be directly determined as the default corresponding data block size.
[0109] Step 404: Determine the target partition number corresponding to the data set to be processed based on the total occupied bytes, the average bytes, and the target partition size.
[0110] Specifically, after determining the total number of occupied bytes, the average number of bytes, and the target partition size, the number of data items that can be stored in a partition can be determined; and then the target number of partitions corresponding to the data set to be processed can be determined.
[0111] Step 405: repartition the data set to be processed according to the target number of partitions.
[0112] Specifically, after the target number of partitions is determined, the data set to be processed can be repartitioned according to the target number of partitions.
[0113] Through the method of this embodiment, when the target data source is a file type data source, the target partition number corresponding to the data set to be processed can be quickly determined.
[0114] As an optional embodiment, as in the above method, step S103 repartitions the data set to be processed according to the total number of occupied bytes, including the following steps:
[0115] Step 501, determining the target data source of the data set to be processed.
[0116] Specifically, when the data set to be processed is obtained, the target data source of the data set to be processed can be determined.
[0117] Optionally, the target data source may include: File type data source: hive / / hdfs / cos / ks3 / obs, etc., based on the size of the overflow file buffer. Relational data source: mysql / oracle / sqlserver / pgsql, etc., based on the data file size configuration in the underlying table space.
[0118] Step 502: Based on the number of bytes of a single data piece occupied by each sampled data piece in the plurality of sampled data pieces, determine the average number of bytes of each data piece in the to-be-processed data set.
[0119] Specifically, after determining the number of bytes occupied by each piece of sampled data, the average number of bytes of multiple pieces of data may be calculated to determine the average number of bytes of the multiple pieces of sampled data, and the average number of bytes may be used as the average number of bytes of the data set to be processed.
[0120] Step 503: When the target data source is a relational data source, a specified partition size of the data set to be processed is configured.
[0121] When the target data source is a relational data source, the corresponding partition size can be configured through the preset configuration, for example, the disk size configuration of the data files in the underlying tablespace such as mysql / oracle / sqlserver / pgsql.
[0122] Configuration instructions can be, for example:
[0123] mysql-innodb_data_file_path, oracle-path_to_datafile, sqlserver-query sys.database_files, Pgsql-wal_buffers.
[0124] Step 504: Determine the designated partition number corresponding to the data set to be processed based on the total number of occupied bytes, the average number of bytes, and the designated partition size.
[0125] Specifically, after determining the total number of occupied bytes, the average number of bytes, and the specified partition size, the number of data items that can be stored in a partition can be determined; and further, the number of specified partitions corresponding to the data set to be processed can be determined.
[0126] Step 505: repartition the data set to be processed according to the specified number of partitions.
[0127] Specifically, after the specified number of partitions is determined, the data set to be processed can be repartitioned according to the specified number of partitions.
[0128] Through the method of this embodiment, when the target data source is a relational data source, the target number of partitions corresponding to the data set to be processed can be quickly determined.
[0129] like Figure 5 As shown, as an optional embodiment, as in the above method, after obtaining the data set to be processed, the method further includes the following steps:
[0130] Step S601: Determine available computing resources corresponding to the data set to be processed.
[0131] Specifically, the remaining computing resources may be analyzed to determine the computing resources allocated to the data set to be processed, that is, to determine the available computing resources corresponding to the data set to be processed.
[0132] Step S602, divide the available computing resources to obtain the single-thread computing resources corresponding to each thread and the target number of all single-thread computing resources, wherein the sum of all single-thread computing resources is the available computing resources.
[0133] Specifically, after determining the available computing resources, in order to implement multi-threaded processing of the data set to be processed, the available computing resources can be divided to determine the single-thread computing resources corresponding to each thread and the target number of single-thread computing resources.
[0134] Generally, the computing resource amounts corresponding to the various single-thread computing resources are the same or similar (ie, the difference in resource amounts is within a preset range).
[0135] For example, when the available computing resources include 5 CPUs, the available computing resources can be divided to obtain 5 single-threaded computing resources, each of which includes 1 CPU, and thus 5 threads can be started simultaneously.
[0136] Step S603: repartition the data set to be processed according to the target number.
[0137] Specifically, after the target number is determined, the data set to be processed can be repartitioned according to the target number to obtain partitions of the target number.
[0138] Through the method of this embodiment, repartitioning can be performed based on default resources (ie, available computing resources), thereby enabling multi-threaded processing of the data set to be processed, thereby effectively improving data processing efficiency.
[0139] Furthermore, after the data set to be processed is determined, the number of partitions manually set by the requester may be accepted, and then the data set to be processed may be repartitioned according to the manually set number of partitions.
[0140] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM (Read-Only Memory) / RAM (Random Access Memory), a disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0142] According to another aspect of the embodiments of the present application, a data set repartitioning device for implementing the above-mentioned data set repartitioning method is also provided. Figure 6 is a structural block diagram of an optional data set repartitioning device according to an embodiment of the present application, such as Figure 6 As shown, the device may include:
[0143] Acquisition module 1, used to acquire the data set to be processed;
[0144] Determination module 2, used for determining the total number of bytes occupied by the data set to be processed by sampling and estimating the data set to be processed;
[0145] The repartitioning module 3 is used to repartition the data set to be processed according to the total number of occupied bytes.
[0146] It should be noted that the acquisition module 1 in this embodiment can be used to execute the above step S101, the determination module 2 in this embodiment can be used to execute the above step S102, and the repartitioning module 3 in this embodiment can be used to execute the above step S103.
[0147] The device in this embodiment, in addition to the above modules, may also include a module for executing any method in any of the embodiments of the aforementioned data set repartitioning method.
[0148] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. It should be noted that the above modules as part of the device can be run in Figure 1 In the hardware environment shown, it can be implemented by software or by hardware, wherein the hardware environment includes a network environment.
[0149] According to another aspect of an embodiment of the present application, an electronic device for implementing the above-mentioned data set repartitioning method is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0150] According to another embodiment of the present application, there is also provided an electronic device, including: Figure 7 As shown, the electronic device may include: a processor 1501 , a communication interface 1502 , a memory 1503 and a communication bus 1504 , wherein the processor 1501 , the communication interface 1502 , and the memory 1503 communicate with each other via the communication bus 1504 .
[0151] Memory 1503, used for storing computer programs;
[0152] The processor 1501 is used to implement the following steps when executing the program stored in the memory 1503:
[0153] Step S101, obtaining a data set to be processed.
[0154] Step S102, determining the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed.
[0155] Step S103, repartitioning the data set to be processed according to the total number of occupied bytes.
[0156] Optionally, in this embodiment, the above-mentioned communication bus can be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0157] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0158] As an example, the memory 1503 may include, but is not limited to, the acquisition module 1, determination module 2, and repartitioning module 3 in the data set repartitioning device. In addition, it may also include, but is not limited to, other module units in the data set repartitioning device, which will not be described in detail in this example.
[0159] The above-mentioned processor can be a general-purpose processor, which can include but not be limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0160] An embodiment of the present application further provides a computer-readable storage medium, the storage medium including a stored program, wherein the method steps of the above method embodiment are executed when the program is run.
[0161] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0162] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0163] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0164] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0165] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0166] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.
[0167] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0168] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A dataset repartitioning method, It is characterized in that include: Get the data set to be processed; Determine the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed; The data set to be processed is repartitioned according to the total number of occupied bytes.
2. The method according to claim 1, It is characterized in that The step of determining the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed includes: Sampling the data set to be processed according to a preset sampling method to obtain a plurality of sampled data; Determine the number of bytes of a single data piece occupied by each of the plurality of sampled data pieces; Based on the number of bytes occupied by each piece of sampled data, the total number of bytes occupied by the data set to be processed is determined.
3. The method according to claim 2, It is characterized in that The determining of the number of bytes of a single data piece occupied by each of the plurality of sampled data pieces comprises: Determine all fields included in the sampled data; Based on the field type of each field in all the fields, determine the number of single field bytes occupied by each field; The number of bytes of a single data item occupied by the sampled data is determined based on the number of bytes of a single field occupied by each field.
4. The method according to claim 2, It is characterized in that The step of sampling the data set to be processed according to a preset sampling method to obtain a plurality of sampled data includes: In the data set to be processed, a preset number of selections are randomly performed to obtain the plurality of sampling data.
5. The method according to claim 2, It is characterized in that The repartitioning of the to-be-processed data set according to the total number of occupied bytes includes: Determining a target data source of the data set to be processed; Determine the average number of bytes of each piece of data in the to-be-processed data set based on the number of bytes of a single piece of data occupied by each piece of sampled data in the plurality of pieces of sampled data; In the case where the target data source is a file type data source, determining a target partition size corresponding to the target data source; Determine the target partition number corresponding to the data set to be processed based on the total number of occupied bytes, the average number of bytes, and the target partition size; The data set to be processed is repartitioned according to the target number of partitions.
6. The method according to claim 2, It is characterized in that The repartitioning of the to-be-processed data set according to the total number of occupied bytes includes: Determining a target data source of the data set to be processed; Determine the average number of bytes of each piece of data in the to-be-processed data set based on the number of bytes of a single piece of data occupied by each piece of sampled data in the plurality of pieces of sampled data; When the target data source is a relational data source, configuring to obtain a specified partition size of the data set to be processed; Determine the designated partition number corresponding to the data set to be processed based on the total number of occupied bytes, the average number of bytes, and the designated partition size; The data set to be processed is repartitioned according to the specified number of partitions.
7. The method according to any one of claims 1 to 6, It is characterized in that After obtaining the data set to be processed, the method further includes: Determining available computing resources corresponding to the data set to be processed; Dividing the available computing resources to obtain single-thread computing resources corresponding to each thread and a target number of all single-thread computing resources, wherein the sum of all single-thread computing resources is the available computing resources; The data set to be processed is repartitioned according to the target number.
8. A data set repartitioning device, It is characterized in that include: An acquisition module is used to obtain the data set to be processed; A determination module, configured to determine the total number of bytes occupied by the data set to be processed by performing sampling estimation on the data set to be processed; The repartitioning module is used to repartition the data set to be processed according to the total number of occupied bytes.
9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, in, The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method steps of any one of claims 1 to 7 by running the computer program stored in the memory.
10. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, wherein the computer program is configured to execute the method steps described in any one of claims 1 to 7 when run.