Data sample partitioning method and system

By processing the original sample dataset in parallel in a distributed cluster and generating data subsets using a bootstrap sampling method, the time and resource bottlenecks caused by serial processing are solved, achieving efficient data sample partitioning and expansion, and supporting subsequent ensemble learning and statistical analysis.

CN114911626BActive Publication Date: 2025-12-12SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210623161.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-12-12
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The serial processing method in the existing technology has the problems of long data processing time and easy resource bottleneck. Especially when the data size and complexity of big data samples increase, the demand for computing resources is high, and it is difficult to efficiently divide and expand the samples.

Method used

A distributed cluster approach is adopted. The master node obtains the original sample dataset and determines the target execution unit and child nodes. The target child nodes control the target execution unit to divide the original sample dataset according to the target partitioning method. Parallel processing is performed using a bootstrap sampling method to generate multiple data subsets and store them in a distributed file system.

Benefits of technology

It achieves efficient and rational use of computing resources, shortens data processing time, improves processing efficiency, breaks through the bottleneck of computing resources, and the generated data subset supports subsequent ensemble learning and statistical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114911626B_ABST
    Figure CN114911626B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data sample division method and system, comprising: main node obtains the original sample data set to be divided, determines the target execution unit of original sample data set, and determines the target subnode in at least one subnode of distributed cluster, and original sample data set is sent to target subnode;Target subnode receives original sample data set, controls corresponding target execution unit and carries out division operation to original sample data set according to target division mode, and each data subset obtained by division operation constitutes target data set;Wherein, target execution unit is the execution unit for dividing processing original sample data set, the number of target execution unit is two or more, and target subnode is the subnode to which target execution unit belongs.The technical scheme of the embodiment of the application realizes the purpose of saving data processing time and improving processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed data processing, and particularly relates to a data sample division method and system. BACKGROUND

[0002] With the continuous updating of Internet technology and the popularization of information network technology, the current era has entered the big data era; through expansion and division of small data samples to form big data samples, which has become an important way to build big data samples.

[0003] At present, when the original sample is a small data sample and the data volume of the original sample cannot meet the data volume requirement of the required sample, the original sample is usually divided and expanded in a serial processing manner. However, due to the increasing data scale and increasing data complexity of big data samples, the division and expansion in the serial processing manner leads to a long data processing time and a high requirement for computing resources, and the problem of unable to continue processing easily occurs. SUMMARY

[0004] The present application provides a data sample division method to solve the problem of long data processing time and resource bottleneck caused by the serial processing manner in the prior art, and achieves the purposes of saving data processing time and improving processing efficiency.

[0005] According to an aspect of the present application, a data sample division method is provided, comprising:

[0006] The master node obtains an original sample data set to be divided, determines a target execution unit of the original sample data set, and determines a target sub-node in at least one sub-node of a distributed cluster, and sends the original sample data set to the target sub-node;

[0007] The target sub-node receives the original sample data set, controls the corresponding target execution unit to perform a division operation on the original sample data set according to a target division manner, and each data subset obtained by the division operation forms a target data set;

[0008] The target execution unit is an execution unit for dividing and processing the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution unit belongs.

[0009] According to another aspect of the present application, a data sample division system is provided, comprising a master node and at least one sub-node, wherein,

[0010] The master node is configured to acquire an original sample data set to be divided, determine a target execution unit of the original sample data set, determine a target sub-node in at least one sub-node of a distributed cluster, and send the original sample data set to the target sub-node.

[0011] The target sub-node is configured to receive the original sample data set, control the corresponding target execution unit to perform a division operation on the original sample data set according to a target division mode, and obtain a target data set by each data subset obtained by the division operation.

[0012] The target execution unit is an execution unit configured to perform a division operation on the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution unit belongs.

[0013] The technical scheme of the embodiment of the present application acquires an original sample data set to be divided by the master node, determines a target execution unit of the original sample data set, determines a target sub-node in at least one sub-node of a distributed cluster, and sends the original sample data set to the target sub-node. The target sub-node controls the corresponding target execution unit to perform a division operation on the original sample data set according to a target division mode, and obtains a target data set by each data subset obtained by the division operation. The target execution unit is an execution unit configured to perform a division operation on the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution unit belongs. The embodiment of the present application solves the problem of long data processing time and resource bottleneck caused by the serial processing mode in the prior art, saves data processing time, improves processing efficiency, and efficiently and reasonably utilizes computing resources.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 A flowchart of a data sample division method provided by the embodiment of the present application is shown in FIG. 1.

[0017] Figure 2A schematic diagram of a data sample division process provided by an embodiment of the present application is shown in the figure;

[0018] Figure 3 A comparison schematic diagram of a data sample division result provided by an embodiment of the present application is shown in the figure;

[0019] Figure 4 A structure diagram of a data sample division system provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0020] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiment of the present application will be described clearly and completely below by combining the figures in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, not all. Based on the embodiment in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the protection scope of the present application.

[0021] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned figures are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0022] Embodiment one

[0023] Figure 1 A flowchart of a data sample division method provided by an embodiment of the present application is shown in the figure. The embodiment can be applicable to sample division of original sample data set to realize sample expansion. The method can be executed by a data sample division system which can be realized in the form of hardware and / or software. As shown in the figure, the method comprises: Figure 1

[0024] S110, the master node acquires the original sample data set to be divided, determines the target execution unit of the original sample data set, and determines the target sub-node in at least one sub-node of the distributed cluster, and sends the original sample data set to the target sub-node.

[0025] ​In this system, the target execution unit is the execution unit used to partition the original sample dataset. There are two or more target execution units, and the target child nodes are the child nodes to which the target execution units belong. The original sample dataset is the dataset to be partitioned and expanded into a large data sample dataset. There can be one master node and one or more child nodes. The master node is connected to each child node and is used to send control commands to each child node to control the operation of the child nodes.

[0026] For example, a computing framework consisting of a master node and child nodes can be the Hadoop framework or the Apache Spark framework. Apache Spark is a memory-based distributed computing framework suitable for handling iterative computations and performing parallel operations across multiple task nodes and data partitions. It utilizes the Resilient Distributed Datasets (RDD) data structure to achieve high scalability, low latency, and high flexibility, greatly improving the computational performance of complex machine learning and data mining analysis.

[0027] In this embodiment, the master node obtains the original sample dataset to be partitioned, which can be achieved by reading the original sample dataset into memory. After obtaining the original sample dataset, the target execution unit for processing the original sample dataset needs to be determined.

[0028] Specifically, the implementation method for determining the target execution unit includes: the master node determines the number of original data rows in the original sample dataset and the preset target number of data rows in the target dataset; the master node determines the target number of target execution units based on the target number of data rows and the original number of data rows; and the master node determines the execution units that meet the target number in the distributed cluster as the target execution units of the original sample dataset.

[0029] The target dataset is the dataset generated after the original sample dataset is partitioned and expanded.

[0030] In the specific implementation, the original sample dataset is set as D1, where D1 = {X1, X2, X3, ..., X...} N}, X1, X2, X3, ..., X N These represent the sample data in the original sample dataset, where N represents the number of original data entries. The number of target data entries in the target dataset can be set according to the actual application. Each execution unit included in each child node can be determined as the target execution unit, or one or more execution units can be randomly selected from each execution unit as the target execution unit; alternatively, the target quantity can be determined based on the number of target data entries and the number of original data entries, and the target execution units can be determined based on the target quantity.

[0031] Optionally, the master node determines the target quantity of the target execution units based on the target data quantity and the original data quantity, including: the master node divides the target data quantity by the original data quantity to obtain a data multiple value; and the master node performs an integer operation on the data multiple value, and determines a value obtained after the integer operation as the target quantity of the target execution units.

[0032] When the data multiple value is an integer, the data multiple value is determined as the target quantity; and when the data multiple value is a decimal, an integer operation is performed to determine two integers adjacent to the data multiple value, and a maximum value of the two integers is determined as the target quantity.

[0033] In a specific implementation, the target execution units of the target quantity can be randomly determined in each sub-node; or an equal quantity of execution units can be determined in each sub-node as the target execution units in a uniform division manner, and a sum of the target execution units determined in each sub-node is equal to the target quantity.

[0034] Optionally, the target execution units of the original sample data set are determined, including: the master node determines working states of each sub-node in the distributed cluster; the working states include an idle state or a busy state; and the master node determines execution units in the sub-nodes in the idle state as the target execution units for performing the division processing on the original sample data set.

[0035] In order to perform the division operation on the original sample data set in time and quickly, the target execution units can be determined based on the working states of the sub-nodes. Specifically, the working states of each sub-node are detected, the sub-nodes in the idle state are determined, and the target execution units are determined in the sub-nodes in the idle state, so as to improve the processing efficiency of the data division processing.

[0036] Further, the sub-nodes to which the target execution units belong are determined as target sub-nodes for controlling the target execution units to perform the division operation.

[0037] S120, the target sub-node receives the original sample data set, controls the corresponding target execution unit to perform a division operation on the original sample data set according to a target division manner, and each data subset obtained by the division operation constitutes a target data set.

[0038] In a specific implementation, the master node can simultaneously and in parallel send the original sample data set to each target sub-node, or determine a sending order of the original sample data set according to the working states of each target sub-node, and sequentially send the original sample data set to the corresponding target sub-node. For example, the working states of each target sub-node are detected, it is determined that there is a target sub-node in the idle state and not receiving the original sample data set, the original sample data set is sent to the target sub-node, and until all the target sub-nodes receive the original sample data set.

[0039] Further, it can be determined whether all target child nodes receive the original sample dataset within a preset time period, and if there is a target child node that has not received the original sample dataset, an error prompt is given, and the target execution unit and the target child node can be re-determined to ensure that the data sample division operation can be normally executed.

[0040] Optionally, the target division mode includes a bootstrap sampling mode, and the corresponding target execution unit is controlled to perform the division operation on the original sample dataset according to the target division mode, including: the target child node controls the corresponding target execution unit to perform the division operation on the original sample dataset according to the bootstrap sampling mode.

[0041] The bootstrap sampling mode is a popular and classical sampling analysis method in modern statistics. The core idea is to randomly and repeatedly sample a limited observation sample of an original dataset to obtain a data subset, that is, each time a sample is selected, it is equally likely to be selected again and added to the result sample set, so some samples may be selected multiple times, and some samples may be ignored. It is mainly used for calculating and constructing the confidence interval of various statistical quantities and other estimation problems.

[0042] In specific implementation, the division operation on the original sample dataset is that the target execution unit performs random and repeated sampling with replacement according to the bootstrap sampling mode to obtain a data subset. In this embodiment, each target execution unit performs the division operation, so each target execution unit corresponds to a data subset.

[0043] Optionally, the data subset includes a sampling block obtained by performing the division operation on the original sample dataset according to the bootstrap sampling mode; each data subset obtained by the division operation constitutes a target dataset, including: the target child node stores the sampling block in the pre-established distributed file system to constitute the target dataset.

[0044] The sampling block can be a BSP block, and the number of BSP blocks existing on each target child node is consistent with the number of target execution units existing on the target child node. In this embodiment, the sampling block is stored in the pre-established distributed file system to constitute the target dataset in the distributed file system. The distributed file system includes an HDFS (Hadoop Distributed File System).

[0045] In the embodiment, the specific implementation of the control of the corresponding target execution unit to perform the dividing operation on the original sample data set according to the target dividing manner can include: when two or more target execution units are included in the target sub-node, the target sub-node controls the target execution units to simultaneously perform the dividing operation on the original sample data set according to the target dividing manner.

[0046] Specifically, based on the independence between each sample in the original sample data set, random sampling with replacement can be performed on the original sample data set, and a plurality of Bootstrap sample sets can be generated in parallel on a distributed cluster, so as to realize the data sample dividing and expanding of the original sample data set. The computing resource bottleneck is broken, the time cost is greatly reduced, the processing efficiency is high, and according to the central limit theorem and the law of large numbers, more BSP blocks can be used to more reliably approach the statistical quantity of the true population. The generated BSP blocks are distributed stored on the HDFS, and only need to be generated once, can be reused afterwards, can be used at any time, and the flexibility is high.

[0047] In the embodiment of the application, before the target execution unit performs the data sample dividing operation on the original sample data set, the master node further determines the data amount of the data subset to be divided by each target execution unit and sends the data amount to the target sub-node corresponding to the target execution unit.

[0048] Specifically, the master node can set the data amount of the data subset to be divided by the target execution unit, so that the target execution unit divides the data subset with the data amount, thereby improving the flexibility of the dividing process. For example, the data amount can be the data amount corresponding to the original sample data set, or a preset data amount specified in advance, which can be set according to actual application by those skilled in the art, and the embodiment of the application is not limited in this regard.

[0049] Further, the target sub-node determines the dividing parameter of the original sample data set based on the received data amount corresponding to each target execution unit, to control the target execution unit to perform the dividing operation on the original sample data set according to the dividing parameter.

[0050] The dividing parameter can include a dividing rate. When the data amount is large, a larger dividing rate can be provided to control the target execution unit to perform the dividing operation on the original sample data set according to the provided dividing rate; when the data amount is small, a smaller dividing rate can be provided.

[0051] The technical scheme of the embodiment of the present application acquires the original sample data set to be divided by the master node, determines the target execution unit of the original sample data set, and determines the target sub-node in at least one sub-node of the distributed cluster, sends the original sample data set to the target sub-node, controls the corresponding target execution unit to perform the division operation on the original sample data set according to the target division mode by the target sub-node, and each data subset obtained by the division operation constitutes a target data set; wherein the target execution unit is an execution unit for performing the division processing on the original sample data set, the number of target execution units is two or more than two, and the target sub-node is a sub-node to which the target execution unit belongs. The embodiment of the present application solves the problem of long data processing time and resource bottleneck caused by the serial processing mode in the prior art, achieves the purpose of saving data processing time and improving processing efficiency, and can efficiently and reasonably utilize the computing resources.

[0052] Embodiment two

[0053] The above describes the embodiments corresponding to the data sample division method in detail. In order to make the technical scheme of the method more clear to those skilled in the art, the specific application scenarios are given below.

[0054] Figure 2 A schematic diagram of a data sample division process provided by the embodiment of the present application is shown in FIG. 1. Figure 2 As shown in the figure, the original data set in the figure can be understood as the original sample data set of the above-mentioned stored sample data, and can be object or attribute type data. The original sample data set is set as D1, D1={X1,X2,X3,……,X N}, X1, X2, X3, …, X N respectively represent the sample data in the original sample data set, N represents the number of original data, and the executor can be an execution unit located on each sub-node.

[0055] In a specific implementation, the master node can send the original data set to each sub-node to control each execution unit to perform the division operation on the original data set. Each sub-node adds a partition task for the executor in the task queue to perform the data division on the original data set D, and the number of partitions corresponds to the number of executors in the Spark cluster.

[0056] Further, each executor applies the Bootstrap method to the data in the data partition on each executor according to the Bootstrap principle to obtain the BSP block on each node. And each BSP block is saved to the distributed file system, indicating that a BSP block is stored on the distributed cluster.

[0057] Wherein, each node divides the respective BSP block, that is, the Bootstrap method is simultaneously and distributedly applied to each partition.

[0058] The bootstrap method can be: randomly repeated sampling with replacement to obtain a sample sample, that is, each time a sample is selected, it is equally likely to be selected again and added to the target data set again. Finally, the target data set obtained includes N*B samples, N represents the number of original data, B is the expansion multiple of the original sample data set, and the target data set is D2={X1*a,X2*b,X3*c,……,X N , a, b, c and z can be the number of each sample in the target data set respectively, so that (a+b+c+……+z)=N*B. In this embodiment, distributed clusters and Spark distributed computing engines are used, BSP blocks are divided on each node and stored locally, and all BSP blocks only need to be generated once and can be reused multiple times.

[0059] Figure 3 A data sample division result comparison schematic diagram is provided for the embodiment of the application, which includes time overhead of dividing 20 data blocks and dividing 50 data blocks. The experimental data is a real data set of 5GB, kaggle-500MB, with dimensions of 50 and 100, and class label numbers of two and five. The conversion efficiency, that is, the time overhead required for conversion, is recorded when multiple different methods are used to convert data blocks of multiple different sizes. If the time overhead exceeds 1800s, it is considered to be too long and is not recorded in the table, and "OT" is used to represent this case. The division method provided by the embodiment of the application can maintain the optimal conversion efficiency.

[0060] As Figure 3 shown, the serial mode, single machine parallel, broadcast distribution mode and BSP method toBSP() operator in the experiment can divide the data amount and time overhead. Among them, when the serial mode is used to divide more than 250MB, the overhead is more than 30mins. The single machine thread parallel mode can only divide 500MB and 250MB within 30mins when dividing 20 and 50 data blocks respectively. The conversion efficiency of the Spark broadcast mode is almost equal to that of the BSP method when the data is 50MB, but it cannot divide data greater than 250MB within 30mins. When the distributed BSP method is used, the optimal conversion efficiency can be maintained. The single machine mode is limited by the processor, memory and other computing resources of the single machine, and the distributed broadcast mode is limited by the running mechanism of the Spark broadcast operation itself. In the case of large data, only the broadcast step will become very time-consuming. Therefore, the distributed BSP method proposed in the embodiment of the application can divide the sample data quickly and efficiently.

[0061] In this embodiment, the data after the division operation is expanded, and the expanded data is expressed as a plurality of BSP blocks stored in the distributed file system. In order to make data preparation for subsequent distributed ensemble learning, the distributed stored BSP blocks support large-scale Bootstrap statistical analysis. The support of the ensemble learning scheme can read the BSP blocks on each computing node, then use the distributed computing framework and the corresponding evaluation method for different algorithms to select the quality of the corresponding blocks, and then perform ensemble learning on the selected BSP blocks to more accurately realize approximate calculation and further improve the prediction effect of ensemble learning.

[0062] Embodiment three

[0063] Figure 4 A structural diagram of a data sample division system provided by an embodiment of the present application is provided. The system is used to perform the data sample division method provided by any of the above embodiments. The system and the target detection method of each of the above embodiments belong to the same inventive concept. Details not described in the embodiment of the data sample division system can be referred to the embodiment of the data sample division method. As shown in the figure, the system includes a master node and at least one sub-node, wherein, Figure 4

[0064] The master node 10 is used to obtain an original sample data set to be divided, determine a target execution unit of the original sample data set, and determine a target sub-node in at least one sub-node of a distributed cluster, and send the original sample data set to the target sub-node.

[0065] The target sub-node 11 is used to receive the original sample data set, control the corresponding target execution unit to perform a division operation on the original sample data set according to a target division manner, and form a target data set by each data subset obtained by the division operation.

[0066] The target execution unit is an execution unit for dividing the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution unit belongs.

[0067] On the basis of any optional technical solution in the embodiments of the present application, optionally, it includes:

[0068] The master node is used to determine the number of original data in the original sample data set and the preset target data number of the target data set.

[0069] The master node is used to determine the target number of the target execution unit based on the target data number and the original data number.

[0070] ​The master node is configured to determine, in the distributed cluster, the execution units satisfying the target number as the target execution units of the original sample data set.

[0071] In any of the optional technical solutions in the embodiments of the present application, the method further comprises the following steps.

[0072] The master node is configured to divide the target data number by the original data number to obtain a data multiple value.

[0073] The master node is configured to perform an integer operation on the data multiple value, and determine a value obtained after the integer operation as the target number of the target execution units.

[0074] In any of the optional technical solutions in the embodiments of the present application, the method further comprises the following steps.

[0075] The master node is configured to determine the working state of each sub-node in the distributed cluster, wherein the working state comprises an idle state or a busy state.

[0076] The master node is configured to determine the execution units in the sub-nodes in the idle state as the target execution units for performing the division processing on the original sample data set.

[0077] In any of the optional technical solutions in the embodiments of the present application, the method further comprises the following steps.

[0078] The master node is further configured to determine the data amount of the data subset to be divided corresponding to each target execution unit, and send the data amount to the target sub-node corresponding to the target execution unit.

[0079] In any of the optional technical solutions in the embodiments of the present application, the method further comprises the following steps.

[0080] The target sub-node is further configured to determine the division parameter of the original sample data set based on the received data amount corresponding to each target execution unit, so as to control the target execution unit to perform the division operation on the original sample data set according to the division parameter.

[0081] In any of the optional technical solutions in the embodiments of the present application, the target division mode comprises a bootstrap sampling mode.

[0082] The target sub-node is configured to control the corresponding target execution unit to perform the division operation on the original sample data set according to the bootstrap sampling mode.

[0083] Optionally, based on any of the optional technical solutions of the embodiments of the present application, the data subset comprises a sampling block obtained by performing a division operation on the original sample data according to the self-service sampling manner.

[0084] The target sub-node is configured to store the sampling block to a pre-established distributed file system to form the target data set.

[0085] Optionally, based on any of the optional technical solutions of the embodiments of the present application, the data subset comprises a sampling block obtained by performing a division operation on the original sample data according to the self-service sampling manner.

[0086] The target sub-node is configured to control each of the target execution units to simultaneously perform a division operation on the original sample data set according to the target division manner when the target sub-node comprises two or more target execution units.

[0087] The data sample division system provided in the embodiments of the present application can execute the data sample division system method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0088] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0089] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of data sample partitioning, the method comprising: The method comprises the following steps: The master node obtains an original sample data set to be divided, determines a target execution unit of the original sample data set, and determines a target sub-node in at least one sub-node of a distributed cluster, and sends the original sample data set to the target sub-node; The target sub-node receives the original sample data set, controls the target execution unit to perform a division operation on the original sample data set according to a target division mode, and obtains each data subset from the division operation to form a target data set; wherein the target division mode is a bootstrap sampling mode, and the target sub-node is configured to control multiple target execution units to simultaneously perform repeated random sampling with replacement on the original sample data set to generate a distributed storage sampling block to form the target data set. The target execution unit is an execution unit for performing division processing on the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution unit belongs.

2. The method of claim 1, wherein, The method further comprises the following steps: The master node determines the number of original data in the original sample data set and a preset target data number of the target data set; The master node determines the target number of target execution units based on the target data number and the number of original data; The master node determines an execution unit satisfying the target number in the distributed cluster as the target execution unit of the original sample data set.

3. The method of claim 2, wherein, The method further comprises the following steps: The master node obtains a data multiple value by dividing the target data number by the number of original data; The master node performs an integer operation on the data multiple value, and determines the value obtained after the integer operation as the target number of target execution units.

4. The method of claim 1, wherein, The method further comprises the following steps: The master node determines the working state of each sub-node in the distributed cluster; wherein the working state comprises an idle state or a busy state; The master node determines an execution unit in a sub-node in the idle state as the target execution unit for performing division processing on the original sample data set.

5. The method of claim 1, wherein, The method further comprises the following steps: The master node determines the data amount of each data subset to be divided corresponding to each target execution unit, and sends the data amount to a target sub-node corresponding to the target execution unit.

6. The method of claim 5, wherein, The method further comprises the following steps: The target sub-node determines a division parameter of the original sample data set based on the data amount corresponding to each target execution unit, and controls the target execution unit to perform a division operation on the original sample data set according to the division parameter.

7. The method of claim 1, wherein, The target division mode comprises a bootstrap sampling mode, and the method further comprises the following steps: The target sub-node controls the target execution unit to perform a division operation on the original sample data set according to the bootstrap sampling mode.

8. The method of claim 7, wherein, The data subsets include sampling blocks obtained by performing a division operation on the original sample data according to the self-service sampling manner. The target data set is composed of the data subsets obtained by the division operation, and includes: The target sub-node stores the sampling blocks in the pre-established distributed file system to form the target data set.

9. The method of claim 1, wherein, The control of the corresponding target execution unit on the division operation of the original sample data set according to the target division manner includes: When the target sub-node includes two or more target execution units, the target sub-node controls each target execution unit to simultaneously perform the division operation on the original sample data set according to the target division manner.

10. A data sample partitioning system, characterized by, It includes: The master node and at least one sub-node, wherein The master node is configured to obtain an original sample data set to be divided, determine a target execution unit of the original sample data set, determine a target sub-node in at least one sub-node of a distributed cluster, and send the original sample data set to the target sub-node. The target sub-node is configured to receive the original sample data set, control a corresponding target execution unit to perform a division operation on the original sample data set according to a target division manner, and compose a target data set from data subsets obtained by the division operation; wherein the target division manner is a self-service sampling manner, and the target sub-node is configured to control multiple target execution units to simultaneously perform random repeated sampling with replacement on the original sample data set to generate distributed storage sampling blocks to form the target data set. The target execution unit is an execution unit configured to perform a division operation on the original sample data set, the number of target execution units is two or more, and the target sub-node is a sub-node to which the target execution units belong.

Citation Information

Patent Citations

  • Data processing method and device and computer readable storage medium

    CN112800020A

  • Parallel task processing method and system for large-scale data

    CN113821329A