Data processing method, electronic device, storage medium and program product

By performing slice operations based on the amount of data in the target column in the file in a distributed computing system, the problem of long-tail processing time is solved, and more balanced data allocation and higher processing efficiency are achieved.

CN120029751APending Publication Date: 2025-05-23HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311573256.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When processing multiple files through distributed computing systems, it often takes a long time to process data fragmentation of certain files, resulting in long tail phenomenon and low overall processing efficiency.

Method used

By obtaining the pending request, the amount of data of the target column in the multiple files is determined, and the file is sliced ​​according to the amount of data is distributed to multiple computing nodes for processing.

Benefits of technology

The file is sliced ​​through the amount of data in the target column, so that the data can be distributed more evenly into each data slice, so that the amount of data actually processed by each computing node is more balanced, reducing the long-tail phenomenon, and improving the overall processing efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029751A_ABST
    Figure CN120029751A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining a to-be-processed request which is used for indicating to process data of a target column, for any file in a plurality of to-be-processed files, determining the data volume of the target column in the file, and according to the data volume of the target column corresponding to the plurality of files, performing slicing operation on each file, the slicing operation being used for segmenting the file into at least one data fragment, and distributing the data fragments corresponding to the plurality of files to a plurality of computing nodes for processing. According to the method, the file can be segmented according to the data volume of the target column, so that the data of the target column can be more uniformly distributed to each data fragment, the data volume actually processed by each computing node is more balanced, the long tail phenomenon can be reduced, and the overall processing efficiency of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a data processing method, electronic equipment, storage medium and program product. Background Art

[0002] Distributed computing systems can process the data contained in files in parallel through multiple computing nodes, which can effectively improve the effect of data processing and has been increasingly widely used in various fields.

[0003] When processing multiple files through a distributed computing system, you can first slice each file to divide one file into multiple data slices, and then distribute the multiple data slices obtained after the slicing operation to multiple computing nodes for separate processing.

[0004] Although this method can realize parallel processing of multiple files, it often takes a long time to process the data fragments of certain files, resulting in a long tail phenomenon and low overall processing efficiency. Summary of the invention

[0005] The present application provides a data processing method, electronic device, storage medium and program product to improve data processing efficiency.

[0006] In a first aspect, an embodiment of the present application provides a data processing method for processing a plurality of files, wherein the files include at least one column of data; the method includes:

[0007] Obtain a pending request, where the pending request is used to instruct processing of data in a target column;

[0008] For any file among the multiple files, determine the data amount of the target column in the file;

[0009] According to the data amount of the target columns corresponding to the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice;

[0010] The data slices corresponding to the multiple files are distributed to multiple computing nodes for processing.

[0011] Optionally, the method further includes:

[0012] In response to a request to write a file, determine statistical information corresponding to the file to be written, and store the file to be written and the corresponding statistical information;

[0013] The statistical information includes the amount of data in each column of the file;

[0014] Accordingly, for any file among the multiple files, determining the data amount of the target column in the file includes:

[0015] For any file, the data volume of the target column in the file is determined according to pre-stored statistical information corresponding to the file.

[0016] Optionally, for any file, determining the data volume of a target column in the file according to pre-stored statistical information corresponding to the file includes:

[0017] For any file, determine whether the file has pre-stored statistical information, and if so, determine the data volume of the target column in the file according to the statistical information;

[0018] The method further includes: for any file, if the file has no pre-stored statistical information, slicing the file according to the file size corresponding to the file.

[0019] Optionally, performing a slicing operation on each file according to the data amount of the target column corresponding to the multiple files includes:

[0020] Determining the number of slices corresponding to each file according to the data amount of the target columns corresponding to the multiple files;

[0021] For any file, the file is divided into at least one data slice according to the number of slices corresponding to the file.

[0022] Optionally, determining the number of slices corresponding to each file according to the data amount of the target columns corresponding to the multiple files includes:

[0023] For any file, the number of slices corresponding to the file is determined according to the data volume of the target column corresponding to the file and the preset slice size.

[0024] Optionally, the multiple files are compressed and stored in the memory; the data volume of the target column is the data volume before compression corresponding to the target column;

[0025] According to the number of slices corresponding to the file, the file is divided into at least one data slice, including:

[0026] Determine the size of the data slice corresponding to the file according to the compressed file size and the number of slices corresponding to the file;

[0027] According to the size of the data slice corresponding to the file, a slicing operation is performed on the compressed file stored in the memory to obtain at least one compressed data slice, so that the computing node decompresses the compressed data slice and processes the decompressed data.

[0028] Optionally, allocating data shards corresponding to the multiple files to multiple computing nodes for processing includes:

[0029] According to the principle of even distribution, the data slices corresponding to the multiple files are distributed to a preset number of computing nodes for processing; or,

[0030] According to the total number of data shards corresponding to the multiple files, a corresponding number of computing nodes are called, and the computing nodes correspond to the data shards one by one, so that each data shard is allocated to a corresponding computing node for processing.

[0031] Optionally, the method further includes:

[0032] Obtaining a slice size for a target column and / or a slice size for a file input by a user through an interactive interface;

[0033] The slice size for the target column is used to: if statistical information of a certain file is pre-stored, determine the number of slices corresponding to the file according to the data volume of the target column in the file and the slice size for the target column;

[0034] The slice size for a file is used to: if the statistical information corresponding to a certain file is not stored, then the number of slices corresponding to the file is determined according to the file size of the file and the slice size for the file.

[0035] In a second aspect, an embodiment of the present application further provides a data processing method for processing a plurality of files, wherein the data contained in the files is divided into at least one data block; the method comprises:

[0036] Obtaining a pending request, where the pending request is used to instruct processing of data in a target data block;

[0037] For any file among the multiple files, determining the data amount of the target data block in the file;

[0038] According to the data amount of the target data blocks corresponding to the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice;

[0039] The data slices corresponding to the multiple files are distributed to multiple computing nodes for processing.

[0040] In a third aspect, an embodiment of the present application further provides a data processing method for processing a plurality of files through a distributed computing system, wherein the files are used to store a plurality of attribute information of a commodity; the method comprises:

[0041] Obtaining a pending request, where the pending request is used to instruct processing of target attribute information;

[0042] For any file among the multiple files, determining the amount of data corresponding to the target attribute information in the file;

[0043] According to the data amount corresponding to the target attribute information in the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice;

[0044] The data slices corresponding to the multiple files are distributed to multiple computing nodes in the distributed computing system for processing.

[0045] In a fourth aspect, an embodiment of the present application provides an electronic device, including:

[0046] at least one processor; and

[0047] a memory communicatively coupled to the at least one processor;

[0048] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method described in any one of the above aspects.

[0049] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method described in any of the above aspects is implemented.

[0050] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method described in any of the above aspects.

[0051] The data processing method, electronic device, storage medium and program product provided in the embodiments of the present application can obtain a request to be processed, wherein the request to be processed is used to indicate processing of data in a target column, and for any file among multiple files to be processed, determine the data amount of the target column in the file, and perform a slicing operation on each file according to the data amount of the target column corresponding to the multiple files, wherein the slicing operation is used to slice the file into at least one data slice, and allocate the data slices corresponding to the multiple files to multiple computing nodes for processing. Since the data amount of the target column can reflect the actual data amount that the computing node needs to process, the file is sliced ​​according to the data amount of the target column, so that the data of the target column can be more evenly distributed to each data slice, thereby making the data amount actually processed by each computing node more balanced, which is conducive to reducing the long tail phenomenon and improving the overall processing efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0053] Figure 1 A schematic diagram of a row storage;

[0054] Figure 2 A schematic diagram of column storage;

[0055] Figure 3 A schematic diagram of an application scenario provided for an embodiment of the present application;

[0056] Figure 4 A schematic diagram of load imbalance caused by different compression rates provided in an embodiment of the present application;

[0057] Figure 5 A schematic diagram of load imbalance caused by different amounts of column data provided in an embodiment of the present application;

[0058] Figure 6 A schematic diagram of the principle of a data processing method provided in an embodiment of the present application;

[0059] Figure 7 A flowchart of a data processing method provided in an embodiment of the present application;

[0060] Figure 8 A schematic diagram of a slicing operation process provided in an embodiment of the present application;

[0061] Fig. 9 A schematic diagram of the principle of collecting file information provided in an embodiment of the present application;

[0062] Fig.10 A schematic diagram of a process of performing a slicing operation using statistical information provided in an embodiment of the present application;

[0063] Fig.11 A flowchart of another data processing method provided in an embodiment of the present application;

[0064] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0065] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0066] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.

[0067] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0068] First, the terms involved in this application are explained:

[0069] Slicing: Divide the file into several splits according to a certain strategy to determine the number of computing nodes. A computing node can be assigned to each split to process the corresponding split through the computing node.

[0070] Input Split: Input data sharding, also known as data sharding, is a logical concept. It does not split files into shards for storage on the disk. Input Split only records the metadata information of the shards, such as the starting position and length, for subsequent use.

[0071] Split Size: Slice size, used to determine the amount of data in the file slices.

[0072] The application scenario of this application is first described below.

[0073] The embodiments of the present application can be applied to process various types of files, and in particular can be used to process files stored in columnar format.

[0074] Figure 1 A schematic diagram of row storage. Figure 2 This is a schematic diagram of column storage. Figure 1 and Figure 2 As shown, the data to be stored includes multiple rows and columns, which can be embodied in the form of a table. The arrows in the table indicate the direction of storage. When stored in row format, the data is saved row by row. When stored in column format, a row of data is split into different columns and saved separately.

[0075] Row-based storage has great advantages over column-based storage in terms of data writing and data modification. However, in terms of data reading, when only part of the columns need to be queried, row-based storage can usually only read out a row of data completely and cannot read only part of the columns, which will increase unnecessary workload and increase processing time.

[0076] refer to Figure 1 For example, when you need to query the transaction volume greater than 100 in the "Transaction Volume" column, the actual data to be processed is only the transaction volume column. If row storage is used, the complete data of each row needs to be read out. In the case of large data volume, the number of rows may reach hundreds of millions, resulting in low overall processing efficiency. Therefore, column storage can be used to save data, and data can be read by column when querying to improve overall efficiency.

[0077] Optionally, when performing columnar storage, if there is a lot of data in the table, it can also be stored in segments according to a preset number of rows. For example, if the table has a total of 100,000 rows and the preset number of rows is 10,000, it can be stored in segments of 10,000 rows, first storing rows 1 to 10,000 of the first column, then storing rows 1 to 10,000 of the second column, until the first 10,000 rows of all columns are stored, and then storing rows 10,001 to 20,000 of the first column, and so on, until all the data in the table is stored.

[0078] In the case of column storage, the stored data can be processed by a distributed computing system. The distributed computing system may include multiple computing nodes, and the multiple computing nodes may process the data to be processed in parallel. In practical applications, the data to be processed may be sliced ​​first to obtain multiple data slices, and then processed by multiple computing nodes.

[0079] Figure 3 A schematic diagram of an application scenario provided by an embodiment of the present application. Figure 3 As shown, the input data stored in the memory is a plurality of files, and the files can be firstly segmented according to the specified slice size to obtain a plurality of data slices, and each data slice is assigned to a computing node for processing.

[0080] Specifically, the records in the data shards may be read one by one by a record reader, and the calculation may be performed by a computing node, wherein the record reader and the computing node may be separate devices or may be integrated together, for example, a computing node may include a record reader.

[0081] Normally, data shards of the same size are allocated to each computing node, so that the amount of data on each computing node is basically the same, in order to ensure load balancing on the computing nodes. However, in actual applications, due to the different conditions of each file, the time spent processing some data shards may not meet expectations, resulting in a long tail phenomenon and a waste of resources.

[0082] Specifically, there are two reasons that may cause load imbalance, which are explained below.

[0083] The first reason is that the slicing operation is for compressed files, and the compression rates of different files are different, resulting in different sizes of data slicing processed by each computing node, thus causing a long tail.

[0084] Figure 4 A schematic diagram of load imbalance caused by different compression rates provided in an embodiment of the present application. Figure 4 As shown in the figure, the compression ratios of file 1, file 2, and file 3 are 100, 50, and 10, respectively, where the compression ratio is used to represent the ratio of the size before compression to the size after compression. According to the compressed size of the file and the slice size of 128M, each file is segmented, and the data slices obtained after segmentation are decompressed and processed. The decompressed size of each data slice corresponding to file 1 is 128M*100=12800M, the decompressed size of each data slice corresponding to file 2 is 128M*50=6400M, and the decompressed size of each data slice corresponding to file 3 is 120M*10=1280M. Due to the different compression ratios of each file, the amount of data finally processed by each computing node varies greatly, resulting in a long tail of computing nodes.

[0085] The second reason is that the slicing operation is performed on the entire file, but the current task may only read some columns, or even a very small number of columns, which causes the actual amount of data processed by the computing node to be far less than expected, resulting in a long tail or waste of resources.

[0086] Figure 5A schematic diagram of load imbalance caused by different column data amounts provided for an embodiment of the present application. As shown in Figure 5, the current task is to select data with c1>5 from column c1 (the corresponding input statement is: select c1 from srcwhere c1>5). In file 1, file 2, and file 3, the data amount of column c1 is 100M, 20M, and 1M, respectively, where the data amount may refer to the size of the data actually to be processed by the computing system. In the process of writing files, the data amount of each column can be determined and stored. A simple example is: a column is used to store transaction volume, the data type is integer (int), corresponding to 4 bytes, assuming that the column has a total of 10,000 rows, that is, it contains 10,000 transaction volume data, then the data amount of the column is 4*10000=40,000 bytes.

[0087] Different amounts of column data may also cause unbalanced load during processing, such as Figure 5 As shown in the figure, after the entire file is sliced ​​according to the slice size of 128M, only the data in column c1 in the obtained data slice will be processed. Therefore, the computing node used to process the data slice of file 1 needs to process 100M of data, and the computing nodes used to process the data slices of file 2 and file 3 need to process 20M and 1M of data respectively. In this way, when the slicing operation is performed according to the size of the entire file, since the sizes of the required columns in each file are different, the actual amount of data processed by the computing nodes will vary greatly, resulting in a long tail of computing nodes.

[0088] To address the above problems, this solution proposes a data processing method based on column statistical information. By using the pre-compression size of the column to be processed and the compressed size of the entire file, the file is sliced, which can effectively solve the load imbalance problem caused by the above two reasons.

[0089] Specifically, statistical information can be pre-stored for files, and the statistical information includes the pre-compression size of each column in the file and the compressed size of the entire file. When obtaining the task to be processed, the columns that need to be processed by the task can be determined first. For each file, the pre-compression size of the column that needs to be processed corresponding to the file is determined based on the statistical information of the file, as the actual data volume corresponding to the file. The number of slices of the file can be determined based on the actual data volume of the file. The larger the actual data volume, the larger the number of slices. Then, the compressed size of the entire file and the number of slices are used to determine the size of each data slice corresponding to the file, thereby completing the slicing operation.

[0090] Figure 6 The following is a schematic diagram of the principle of a data processing method provided in an embodiment of the present application. Figure 6As shown, the current task is to select data with c1>5 from column c1. The slice size of the column is 128M. For each file, first determine the size of column c1 before compression in the file based on the statistical information corresponding to the file, divide the size before compression by the slice size 128M to get the number of slices of the file, and divide the compressed size of the file by the corresponding number of slices to get the size of each data slice corresponding to the file. The size of the data slices obtained after different files are split may be different, but the actual amount of data contained in them is basically the same. Allocating a computing node to each data slice for processing can effectively achieve load balancing.

[0091] In summary, the data processing method based on column statistical information proposed in the embodiment of the present application can use the columns actually read by the task and their statistical information to divide the file more evenly, thereby effectively avoiding the long tail and improving the task execution efficiency.

[0092] Some embodiments of the present application are described in detail below in conjunction with the accompanying drawings. In the case where there is no conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0093] Figure 7 A flow chart of a data processing method provided in an embodiment of the present application. The method in this embodiment can be implemented on any device with data processing function, for example, on the cloud, locally deployed, client-side, IOT (Internet of Things) device implementation, etc.

[0094] The method in this embodiment can be used to process multiple files through a distributed computing system. For example, when implemented on the cloud, the distributed computing system on the cloud can be used to implement file processing. In a cloud-native scenario, the cloud-native big data computing service can be used to execute Figure 7 The steps shown are used to implement the file slicing operation and distribute the data slicing to the computing nodes in the distributed computing system for processing.

[0095] Optionally, the cloud native big data computing service can read data from a file system, especially a distributed file system, where each file in the file system can include at least one column of data. Of course, the embodiments of the present application can also be applied to other arbitrary distributed scenarios, such as reading files through other computing engines, which is not limited here. Alternatively, the embodiments of the present application can also be applied to non-distributed computing systems.

[0096] like Figure 7 As shown, the method may include:

[0097] Step 701: Obtain a pending request, where the pending request is used to instruct processing of data in a target column.

[0098] Optionally, the pending request may be initiated by a user, or may be automatically generated according to user requirements, or may be automatically generated through a pre-set task. The pending request may be expressed in a structured query language (SQL) or other forms.

[0099] The pending request can be used to process the data in the file, and can be used to process all the data or part of the data. In the case of column storage, the pending request can generally be used to process the data of one or more columns in the file, and the one or more columns can be used as target columns.

[0100] For example, if the pending request is to find the transaction volume greater than 100, the target column is "transaction volume". Alternatively, if the pending request is to find the product name with a price less than 10, the target columns are "price" and "product name".

[0101] Step 702: For any file among the multiple files, determine the data volume of the target column in the file.

[0102] The data volume of the target column may specifically refer to the data volume of the data to be processed in the target column. The data volume of the target column may be obtained in a variety of ways, for example, the data volume corresponding to each column may be pre-stored, or may be determined by information such as the number of rows and data type corresponding to the target column.

[0103] The multiple files may be specified by the user, for example, the user may select to query and process certain files. Alternatively, the multiple files to be processed may be automatically determined according to the target column, for example, the files containing the target column in the file system may be used as the multiple files to be processed.

[0104] The number and specific contents of the columns contained in the multiple files may be the same or different. In general, the multiple files may contain the same columns, for example, all contain three columns of "product name", "price", and "transaction volume". However, in some cases, some files may contain more or fewer columns, and this embodiment also supports processing these cases.

[0105] For each file, the data volume of the target column in the file can be determined. When there are multiple target columns, the data volume of the target column in the file can be specifically the total data volume of the multiple target columns.

[0106] The amount of data in the target column in different files can be different, which serves as the basis for subsequent slicing operations.

[0107] Step 703: Perform a slicing operation on each file according to the data amount of the target column corresponding to the multiple files, wherein the slicing operation is used to slice the file into at least one data slice.

[0108] Each file may correspond to one or at least two data shards. After performing slicing operations on multiple files, multiple data shards may be obtained. Optionally, the goal of the slicing operation may be to make the size of the data to be processed in each data shard obtained after slicing as consistent as possible. The data to be processed may be data belonging to the target column.

[0109] Figure 8 A schematic diagram of a slicing operation process provided in an embodiment of the present application. Figure 8 As shown, step 703 can be specifically implemented through steps 7031 and 7032.

[0110] Step 7031: Determine the number of slices corresponding to each file according to the data amount of the target columns corresponding to the multiple files.

[0111] The number of slices is used to indicate the number of data slices obtained after the file is sliced.

[0112] In an optional implementation, for any file, the number of slices corresponding to the file can be determined according to the data volume of the target column corresponding to the file and a preset slice size.

[0113] Among them, the preset slice size can be used to split the data volume of the target column. Splitting the data volume of the target column here is a logical concept. It does not directly split the target column, but calculates how many data slices will be obtained when the target column is split according to the slice size, thereby determining the number of slices corresponding to the file.

[0114] Optionally, you can divide the amount of data in the target column by the slice size to get the number of slices corresponding to the file.

[0115] Exemplarily, the preset slice size is 128M, and the data volume of the target column corresponding to a certain file is 256M, then it can be determined that the number of slices corresponding to the file is 256M÷128M=2.

[0116] Optionally, the preset slice size may be set by the user, or may adopt a default value, or may be determined according to the data volume of the target columns in the current multiple files, for example, by taking the greatest common divisor of the data volume of the target columns corresponding to the multiple files.

[0117] By presetting the slice size and the amount of data in the target column of each file, the number of slices for each file can be quickly determined, and the amount of data to be processed in the multiple data slices obtained after slicing is basically the same, which is conducive to achieving load balancing among multiple computing nodes.

[0118] In another optional implementation, the total number of slices may be determined first, and then the number of slices for each file may be determined based on the total number of slices and the data volume of the target column of each file.

[0119] The total number of slices can be preset and is also determined according to the number of multiple computing nodes.

[0120] Exemplarily, the total number of slices is 3, and the multiple files to be processed include file A and file B. The data volume of the target column in file A is 256M, and the data volume of the target column in file B is 512M. Then, file A can be divided into 1 data slice, and file B can be divided into 2 data slices, so that the size of the data to be processed contained in the obtained 3 data slices is as consistent as possible.

[0121] Step 7032: For any file, divide the file into at least one data slice according to the number of slices corresponding to the file.

[0122] Optionally, the size of each data slice corresponding to a file may be determined based on the file size and the number of slices, and the file may be segmented based on the size of the data slice to obtain at least one data slice, wherein the size of the data slice may be equal to the file size divided by the number of slices.

[0123] For example, the file size is 1024M, and the corresponding number of slices is 2, then the size of each data slice is 1024M÷2=512M, and the file is split according to the size of 512M to obtain two data slices.

[0124] Through the above method, the number of slices corresponding to each file can be determined according to the data volume of the target column. The data volume of the target column corresponding to the file and the number of slices can be positively correlated. Then, the file can be divided according to the number of slices, which can quickly and accurately implement the slicing operation and improve the slicing efficiency.

[0125] Optionally, in order to save storage space, the multiple files may be stored in a compressed manner in the memory, that is, compressed files are stored in the memory.

[0126] When calculating the number of slices, the data volume of the target column used may be the data volume before compression corresponding to the target column. The data volume before compression may reflect the data volume that the computing node actually needs to process. Therefore, using the data volume before compression of the target column may more accurately ensure load balancing and reduce resource waste.

[0127] Optionally, in step 7032, the file is divided into at least one data slice according to the number of slices corresponding to the file, which may include: determining the size of the data slice corresponding to the file according to the compressed file size and the number of slices corresponding to the file; and slicing the compressed file stored in the memory according to the size of the data slice corresponding to the file to obtain at least one compressed data slice, so that the computing node decompresses the compressed data slice and processes the decompressed data.

[0128] For example, the compressed file size of a certain file is 1024M, and the file is divided into two data slices, and the compressed size of each data slice is 512M. Assuming the compression ratio is 10, the decompressed size of each data slice is 5120M, and the computing node can process the target column in the 5120M data.

[0129] refer to Figure 6 Different files can be sliced ​​using different data slice sizes. Although the data slice size corresponding to each file is different, the actual amount of data that needs to be processed contained in each data slice is basically the same.

[0130] Through the above method, compressed and stored files can be processed, which can effectively avoid load imbalance caused by different compression rates of multiple files or different target column data amounts, and improve the overall processing efficiency of the system.

[0131] The above introduces a method for implementing slicing operations based on the number of slices. In other optional implementation methods, slicing operations are performed on each file according to the data volume of the target column corresponding to the multiple files. This can also be achieved through the following steps: based on the data volume of the target column corresponding to the file and the file size, the size of the data slice corresponding to the file is directly determined by calling a function or looking up a table, wherein the function or table used can be set in advance based on the principle of load balancing, thereby eliminating the intermediate operation of calculating the number of slices.

[0132] Step 704: Allocate data slices corresponding to the multiple files to multiple computing nodes for processing.

[0133] This embodiment does not limit the specific form of the computing node. For example, in some technologies, the computing node can be embodied in the form of a map task (source machine).

[0134] In an optional implementation, the data slices corresponding to the multiple files may be distributed to a preset number of computing nodes for processing according to the principle of even distribution.

[0135] Exemplarily, assuming that multiple files are split into N data shards, which are assigned to M computing nodes for processing, then the number of data shards processed by each computing node is N / M. Therefore, N / M data shards can be assigned to each computing node.

[0136] It should be noted that in this implementation, data shards are allocated based on the principle or goal of equal distribution, but it does not necessarily need to be completely equal. For example, when N cannot be divided by M, N / M can be rounded up to get the number of data shards corresponding to each computing node. Although it is not completely equal at this time, the principle of equal distribution is still used, which enables each computing node to process the data shards as evenly as possible.

[0137] In another optional implementation, a corresponding number of computing nodes can be called according to the total number of data shards corresponding to the multiple files, and the computing nodes correspond one-to-one to the data shards, so that each data shard can be assigned to a corresponding computing node for processing.

[0138] Exemplarily, assuming that multiple files are segmented into N data shards, N computing nodes can be called to process the N data shards. The N computing nodes correspond to the N data shards one by one, and each computing node is used to process the corresponding data shard.

[0139] Optionally, if the number of available computing nodes in the distributed computing system is less than the number of current data shards, data processing can be implemented through multiple rounds. In the first round, a data shard is allocated to each computing node. After the data shards in the first round are processed, the next round of processing is performed, and another data shard is allocated to each computing node until the multiple data shards are processed.

[0140] Through the above method, the number of data shards processed by each computing node can be made as consistent as possible, thereby providing a guarantee for the load balancing of the system.

[0141] In the case of column storage, when the pending request is to process the target column in the file, the computing node can only read and process the data in the target column without paying attention to the data in other columns. Therefore, after the computing node is assigned to each data shard, the computing node can read the data shard from the corresponding position and process it according to the metadata information of the data shard, such as the starting position, length, etc.

[0142] In the compressed storage scenario, what is read is the compressed data shards. You can first decompress the compressed data shards, read the data of the target column one by one from the decompressed data, and process them.

[0143] In other optional implementations, the file may not be compressed and directly stored in the memory. For uncompressed files, the size of the data slices may be obtained by dividing the file size by the number of slices, and the file may be segmented according to the size of the data slices to obtain data slices. The computing nodes may directly process the data slices without decompressing the data slices.

[0144] In summary, the data processing method provided in this embodiment can determine the data volume of the target column to be processed in each file when a request to be processed is obtained. The data volume of the target column can reflect the actual data volume that the computing node needs to process. The file is split according to the data volume of the target column, so that the data of the target column can be more evenly distributed to each data slice, thereby making the actual data volume processed by each computing node more balanced, which is conducive to reducing the long tail phenomenon and improving the overall processing efficiency of the system.

[0145] In one or more embodiments of the present application, optionally, statistical information of a file may be collected when the file is written, and the collected statistical information may be used to subsequently segment the file.

[0146] Specifically, in response to a request to write a file, statistical information corresponding to the file to be written can be determined, and the file to be written and the corresponding statistical information can be stored; wherein the statistical information includes the data volume of each column in the file.

[0147] Fig. 9 A schematic diagram of the principle of collecting file information provided by the embodiment of the present application. Fig. 9 As shown, the request to write a file can be expressed as an insert data query. After obtaining the request to write a file, statistical information of the file can be collected while writing the file, and the file and the statistical information can be written together to a storage device (such as a disk).

[0148] Specifically, the process of writing files can be completed by the control node and the computing node. The control node can include a compiler and an optimizer, etc. After obtaining a request to write a file, the request can be compiled and optimized, thereby converting the request into instructions that can be executed by the machine. The computing node can execute the instructions output by the control node, thereby completing the storage of files and statistical information.

[0149] The statistical information of any file may include the data volume of each column in the file. Accordingly, when determining the data volume of a target column in the file, the data volume of the target column in the file may be determined based on the statistical information corresponding to the file stored in advance.

[0150] In the scenario of compressed storage, the data volume of each column stored in the statistical information may include the data volume of each column before compression. In addition to the data volume of each column, the statistical information may also include various information that may be used in subsequent processing, such as the compressed file size of the file.

[0151] Exemplarily, the statistical information of the file may include: name, size (file size; in a compression scenario, size may be the compressed file size), row count, raw-size (size before compression), and table statistics (table statistical information).

[0152] Among them, table statistics can be used to represent the statistics of each column in the file. The statistics of each column can include: col-ID (column ID), min (minimum value), max (maximum value), raw-size (original size), compressed-size (compressed size), and null-count (number of null values).

[0153] Through the above method, statistical information can be collected and saved in the process of writing files, and the amount of data in each column can be efficiently and accurately counted during the writing process. Moreover, in the subsequent processing of the file, the pre-stored statistical information can be used to assist in the slicing operation, thereby improving the efficiency of file processing.

[0154] When the above-mentioned method of collecting statistical information is applied to a file system, statistical information can be collected for files newly written into the file system, while old files stored in the file system a long time ago may not have statistical information. The embodiment of the present application also provides a solution for mixed processing of new and old files.

[0155] Specifically, for any file, determining the data amount of the target column in the file based on the pre-stored statistical information corresponding to the file may include: for any file, determining whether the file has pre-stored statistical information, and if so, determining the data amount of the target column in the file based on the statistical information.

[0156] For any file, if there is no pre-stored statistical information for the file, a slicing operation is performed on the file according to the file size corresponding to the file.

[0157] Fig.10 A schematic diagram of a process of performing a slicing operation using statistical information provided in an embodiment of the present application. Fig.10 As shown, slicing operations using statistical information can include the following steps:

[0158] Step a: After the request to be processed is processed by the compiler and the optimizer, a target column is obtained, that is, the column actually required to be read.

[0159] Step b: When performing a file slicing operation, determine whether the target column has corresponding statistical information. If so, use a new slicing method; otherwise, use the old slicing method.

[0160] Among them, the input splitor (input data slicer) can be configured respectively for the new segmentation method and the old segmentation method. The new segmentation method can be implemented based on the new input data slicer, and the old segmentation method can be implemented based on the old input data slicer.

[0161] The new slicing method will perform slicing based on the target column and its statistical information. Specifically, the amount of data in the target column can be divided by the corresponding slice size to get the number of slices, and the file can be sliced ​​based on the number of slices. The old slicing method can directly perform slicing operations based on file size. Specifically, the file size can be divided by the corresponding slice size to get the number of slices.

[0162] Among them, in the new and old segmentation methods, the corresponding slice sizes can be the same or different.

[0163] Optionally, a slice size for a target column and / or a slice size for a file input by a user through an interactive interface may be obtained.

[0164] The slice size for the target column may be the slice size used in the new segmentation method, and is specifically used for: if the statistical information of a file is pre-stored, then the number of slices corresponding to the file is determined according to the data volume of the target column in the file and the slice size for the target column. The slice size for a file may be the slice size used in the old segmentation method, and is specifically used for: if the statistical information corresponding to a file is not stored, then the number of slices corresponding to the file is determined according to the file size of the file and the slice size for the file.

[0165] Optionally, the user can set the slice size in the new and old segmentation methods at one time, and the slice size set by the user can be used each time the data in the file is processed. Alternatively, the user can be supported to dynamically adjust the current slice size each time the data in the file is processed. Optionally, the information about the current file to be processed and the target column can be displayed on the interactive interface, so that the user can determine the slice size based on the current information.

[0166] In addition to user configuration, the slice size in the new and old segmentation methods can also be determined by other means, such as using default settings, or automatically adjusting according to the actual situation of the current file.

[0167] Through the above solution, new and old files in the file system can be mixed according to the slice size for the target column and the slice size for the file to meet the actual usage requirements. In addition, the user is allowed to configure the slice size of the target column and the slice size of the file to improve the user experience.

[0168] In one or more embodiments of the present application, optionally, when the statistical information of the file also includes: the number of rows in the file and the maximum value, minimum value, and number of null values ​​corresponding to each column in the file, the statistical information can be used to further optimize the processing flow.

[0169] Optionally, when there are multiple target columns and the request to be processed is specifically used to instruct to select data that meets preset requirements from multiple target columns, for any file, the data amount of the target column in the file may be determined by the following method, wherein the data amount of the target column in the file is the size of the data to be processed in the multiple target columns:

[0170] For any target column in the file, if it is judged that the target column does not contain data that meets the preset requirements based on at least one of the maximum value, the minimum value, the number of null values, and the number of rows of the target column, the size of the data to be processed in the target column is determined to be zero; if it is judged that the target column may contain data that meets the preset requirements based on at least one of the maximum value, the minimum value, the number of null values, and the number of rows of the target column, the size of the data to be processed in the target column is determined to be the data amount of the target column recorded in the statistical information;

[0171] The sum of the sizes of the to-be-processed data corresponding to the multiple target columns is calculated to obtain the sizes of the to-be-processed data in the multiple target columns.

[0172] Exemplarily, the target column is column c1, and the request to be processed is: select data with c1>5 from column c1. If the maximum value of column c1 is determined to be 3 based on the statistical information of the file, it means that there is no data in column c1 of the file that meets the preset requirements, then the size of the data to be processed corresponding to column c1 is 0.

[0173] Similarly, when the request to be processed is to select data with c1<100, if it is determined according to the statistical information that the minimum value of the c1 column is greater than 100, it is determined that there is no data in the c1 column that meets the preset requirements, or, if it is determined according to the statistical information that the number of rows in the c1 column is equal to the number of null values, it means that all the values ​​in the c1 column are null and there is no data that meets the preset requirements.

[0174] If the possibility that data that meets the preset requirements exists in column c1 cannot be ruled out based on information such as the maximum value, minimum value, number of null values, and number of rows, then it can be considered that there may be data that meets the preset requirements in column c1. At this time, the corresponding size of the data to be processed is the amount of data in column c1 recorded in the statistical information.

[0175] When there are multiple target columns, the sizes of the data to be processed in the multiple target columns can be added together to obtain the final data volume for calculating the number of slices. When there is only one target column, the number of slices can be directly calculated using the size of the data to be processed in the target column.

[0176] Through the above scheme, the size of the data to be processed in the target column can be determined according to the statistical information, so as to filter out the data that does not actually need to be processed, reduce the amount of calculation in the subsequent decompression and processing process, and further improve the efficiency of the system.

[0177] To summarize, the data processing method based on column statistical information provided in the embodiment of the present application records fine-grained file-level statistical information. During the task processing process, slicing operations can be performed according to the columns that actually need to be read in the task and the original size of the column in the file. Compared with slicing operations using the compressed file size, this method can make the data shards after slicing more uniform, so that the amount of data processed by each computing node is more uniform, effectively improving the overall execution efficiency of the computing nodes, thereby improving the execution efficiency of the tasks.

[0178] In addition to the above-mentioned data processing method applicable to column storage, the embodiment of the present application also provides a data processing method applicable to other storage methods, which is described below.

[0179] Fig.11 A flowchart of another data processing method provided in an embodiment of the present application. The method in this embodiment is used to process multiple files, and the data contained in the files is divided into at least one data block. The division method of the data blocks is not limited. For example, when dividing by rows, each row can be used as a data block, and when dividing by columns, each column can be used as a data block. Of course, other division methods can also be used, for example, a certain number of rows or columns can be used as a data block. Fig.11 As shown, the method includes:

[0180] Step 1101: Obtain a pending request, where the pending request is used to instruct processing of data in a target data block.

[0181] Step 1102: For any file among the multiple files, determine the data volume of the target data block in the file.

[0182] Step 1103: Perform a slicing operation on each file according to the data amount of the target data blocks corresponding to the multiple files, wherein the slicing operation is used to divide the file into at least one data slice.

[0183] Step 1104: Allocate data slices corresponding to the multiple files to multiple computing nodes for processing.

[0184] The specific implementation principle and process of this embodiment can refer to the above-mentioned embodiments, as long as the columns in the above-mentioned embodiments are replaced by data blocks.

[0185] Exemplarily, in the case of row-based storage, a data block may include a row of data. When recording the statistical information of the file, the corresponding pre-compression size of each row may be recorded. When executing a task, the pending request may be used to process the data of the target row. At this time, the data volume of the target row may be determined based on the statistical information, and the file may be sliced ​​based on the data volume of the target row, so that the data of the target row is distributed as evenly as possible in the data shards obtained after slicing, thereby improving the load balancing effect in the row-based storage scenario.

[0186] The data processing method provided in this embodiment can determine the data volume of the target data blocks to be processed in each file when a request to be processed is obtained. The data volume of the target data blocks can reflect the data volume that the computing node actually needs to process. The file is split according to the data volume of the target data blocks, so that the data of the target data blocks can be more evenly distributed to each data slice, thereby making the data volume actually processed by each computing node more balanced, which is conducive to reducing the long tail phenomenon and improving the overall processing efficiency of the system.

[0187] The embodiment of the present application further provides a data processing method, which is applied to an e-commerce scenario. The method is used to process multiple files through a distributed computing system, and the files are used to store multiple attribute information of commodities. The method includes:

[0188] Obtaining a pending request, where the pending request is used to instruct processing of target attribute information;

[0189] For any file among the multiple files, determining the amount of data corresponding to the target attribute information in the file;

[0190] According to the data amount corresponding to the target attribute information in the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice;

[0191] The data slices corresponding to the multiple files are distributed to multiple computing nodes in the distributed computing system for processing.

[0192] Exemplarily, the multiple attribute information may include: name, price, transaction volume, brand, origin, etc. Each attribute information may be stored as a column. According to the solution provided in this embodiment, the file can be sliced ​​according to the target attribute information.

[0193] The specific implementation principle and process of this embodiment can refer to the aforementioned embodiment, as long as the columns in the aforementioned embodiment are replaced with attribute information.

[0194] The data processing method provided in this embodiment can determine the data volume of the target attribute information to be processed in each file when a request to be processed is obtained. The data volume of the target attribute information can reflect the actual data volume that the computing node needs to process. The file is split according to the data volume of the target attribute information, so that the target attribute information can be more evenly distributed to each data slice, thereby making the actual data volume processed by each computing node more balanced, which is conducive to reducing the long tail phenomenon and improving the overall processing efficiency of the system.

[0195] In addition to e-commerce scenarios, the embodiments of the present application can also be applied to any other scenarios with data processing requirements. The specific implementation principles are similar and will not be repeated here.

[0196] Corresponding to the above data processing method, on one hand, an embodiment of the present application further provides a data processing device for processing a plurality of files, wherein the files include at least one column of data; the device includes:

[0197] An acquisition module, used for acquiring a pending request, wherein the pending request is used for instructing to process data of a target column;

[0198] A determination module, configured to determine, for any file among the plurality of files, the amount of data in a target column in the file;

[0199] A slicing module, used to perform a slicing operation on each file according to the data amount of the target column corresponding to the multiple files, wherein the slicing operation is used to slice the file into at least one data slice;

[0200] The allocation module is used to allocate the data slices corresponding to the multiple files to multiple computing nodes for processing.

[0201] On the other hand, an embodiment of the present application further provides a data processing device for processing a plurality of files, wherein the data contained in the files is divided into at least one data block; the device comprises:

[0202] An acquisition module, used for acquiring a pending request, wherein the pending request is used for instructing to process the data in the target data block;

[0203] A determination module, configured to determine, for any file among the plurality of files, the data volume of a target data block in the file;

[0204] A slicing module, configured to perform a slicing operation on each file according to the data amount of the target data blocks corresponding to the multiple files, wherein the slicing operation is used to slice the file into at least one data slice;

[0205] The allocation module is used to allocate the data slices corresponding to the multiple files to multiple computing nodes for processing.

[0206] In another aspect, an embodiment of the present application further provides a data processing device for processing a plurality of files through a distributed computing system, wherein the files are used to store a plurality of attribute information of a commodity; the device comprises:

[0207] An acquisition module, used for acquiring a pending request, wherein the pending request is used for instructing to process target attribute information;

[0208] A determination module, configured to determine, for any file among the plurality of files, the amount of data corresponding to the target attribute information in the file;

[0209] A slicing module, configured to perform a slicing operation on each file according to the amount of data corresponding to the target attribute information in the multiple files, wherein the slicing operation is used to slice the file into at least one data slice;

[0210] The allocation module is used to allocate the data slices corresponding to the multiple files to multiple computing nodes in the distributed computing system for processing.

[0211] Each device provided in the embodiments of the present application is used to execute the corresponding method embodiments described above. The specific implementation principles and beneficial effects can be found in the above embodiments and will not be repeated here.

[0212] Fig.12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.12 As shown, the electronic device of this embodiment may include:

[0213] At least one processor 1201; and a memory 1202 in communication with the at least one processor; wherein the memory 1202 stores instructions executable by the at least one processor 1201, and the instructions are executed by the at least one processor 1201 to enable the electronic device to perform the method as described in any of the above embodiments. Optionally, the memory 1202 can be independent or integrated with the processor 1201.

[0214] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.

[0215] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.

[0216] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the above embodiments when executed by a processor.

[0217] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0218] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.

[0219] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The memory may include a high-speed random access memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.

[0220] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0221] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0222] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0223] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0224] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0225] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A data processing method, It is characterized in that Used to process a plurality of files, wherein the files include at least one column of data; the method includes: Obtain a pending request, where the pending request is used to instruct processing of data in a target column; For any file among the multiple files, determine the data amount of the target column in the file; According to the data amount of the target columns corresponding to the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice; The data slices corresponding to the multiple files are distributed to multiple computing nodes for processing.

2. The method according to claim 1, It is characterized in that Also includes: In response to a request to write a file, determine statistical information corresponding to the file to be written, and store the file to be written and the corresponding statistical information; The statistical information includes the amount of data in each column of the file; Accordingly, for any file among the multiple files, determining the data amount of the target column in the file includes: For any file, the data volume of the target column in the file is determined according to pre-stored statistical information corresponding to the file.

3. The method according to claim 2, It is characterized in that For any file, the amount of data in the target column in the file is determined according to pre-stored statistical information corresponding to the file, including: For any file, determine whether the file has pre-stored statistical information, and if so, determine the data volume of the target column in the file according to the statistical information; The method further includes: for any file, if the file has no pre-stored statistical information, slicing the file according to the file size corresponding to the file.

4. The method according to claim 1, It is characterized in that According to the data amount of the target columns corresponding to the multiple files, a slicing operation is performed on each file, including: Determining the number of slices corresponding to each file according to the data amount of the target columns corresponding to the multiple files; For any file, the file is divided into at least one data slice according to the number of slices corresponding to the file.

5. The method according to claim 4, It is characterized in that Determining the number of slices corresponding to each file according to the data amount of the target columns corresponding to the multiple files includes: For any file, the number of slices corresponding to the file is determined according to the data volume of the target column corresponding to the file and the preset slice size.

6. The method according to claim 4, It is characterized in that The multiple files are compressed and stored in the memory; the data volume of the target column is the data volume before compression corresponding to the target column; According to the number of slices corresponding to the file, the file is divided into at least one data slice, including: Determine the size of the data slice corresponding to the file according to the compressed file size and the number of slices corresponding to the file; According to the size of the data slice corresponding to the file, a slicing operation is performed on the compressed file stored in the memory to obtain at least one compressed data slice, so that the computing node decompresses the compressed data slice and processes the decompressed data.

7. The method according to any one of claims 1 to 6, It is characterized in that Allocating data slices corresponding to the multiple files to multiple computing nodes for processing includes: According to the principle of even distribution, the data slices corresponding to the multiple files are distributed to a preset number of computing nodes for processing; or, According to the total number of data shards corresponding to the multiple files, a corresponding number of computing nodes are called, and the computing nodes correspond to the data shards one by one, so that each data shard is allocated to a corresponding computing node for processing.

8. The method according to claim 3, It is characterized in that Also includes: Obtaining a slice size for a target column and / or a slice size for a file input by a user through an interactive interface; The slice size for the target column is used to: if statistical information of a certain file is pre-stored, determine the number of slices corresponding to the file according to the data volume of the target column in the file and the slice size for the target column; The slice size for a file is used to: if the statistical information corresponding to a certain file is not stored, then the number of slices corresponding to the file is determined according to the file size of the file and the slice size for the file.

9. A data processing method, It is characterized in that Used to process multiple files, wherein the data contained in the files is divided into at least one data block; the method comprises: Obtaining a pending request, where the pending request is used to instruct processing of data in a target data block; For any file among the multiple files, determining the data amount of the target data block in the file; According to the data amount of the target data blocks corresponding to the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice; The data slices corresponding to the multiple files are distributed to multiple computing nodes for processing.

10. A data processing method, It is characterized in that The method is used to process a plurality of files through a distributed computing system, wherein the files are used to store a plurality of attribute information of commodities; the method comprises: Obtaining a pending request, where the pending request is used to instruct processing of target attribute information; For any file among the multiple files, determining the amount of data corresponding to the target attribute information in the file; According to the data amount corresponding to the target attribute information in the multiple files, a slicing operation is performed on each file, wherein the slicing operation is used to slice the file into at least one data slice; The data slices corresponding to the multiple files are distributed to multiple computing nodes in the distributed computing system for processing.

11. An electronic device, It is characterized in that include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method described in any one of claims 1 to 10.

12. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 10 is implemented.

13. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.