Large-scale data processing method, distributed system and device
By sorting the cached data locally in large-scale data processing and reading data in a specified database in a way that is equivalent to global sorting, the problem of inefficient large-scale data processing in the prior art is solved, and efficient data processing and query are achieved.
Patent Information
- Application Number
- CN202510158420.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
AI Technical Summary
In large-scale data processing, it is difficult for the prior art to efficiently process large-scale data, especially in the process of data screening and sorting, which has the problem of inefficiency.
Avoid global sorting of multiple database files by sorting multiple data in the cache locally and writing them to disk in the form of database files. Then, multiple database files in the disk are loaded by specifying a database (such as RocksDB), and data is read from multiple database files in a manner equivalent to global sorting.
This method improves the efficiency of large-scale data processing, reduces read and write operations to disk, saves performance losses, and implements fast data queries.
Smart Images

Figure CN120104572A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of computer technology, and in particular, to a large-scale data processing method, distributed system, and device. Background Art
[0002] In the context of big data and cloud computing, large-scale data processing is becoming more and more frequent. In order to process data tasks more efficiently, such as filtering data by conditions, data can usually be distributed to different node devices for task processing. This process involves the import and export of large amounts of data, and often requires global sorting of data to facilitate data task processing. The data processing process also requires privacy protection.
[0003] At present, it is hoped that there will be improved solutions that can improve the efficiency of processing large-scale data. Summary of the invention
[0004] One or more embodiments of this specification describe a large-scale data processing method, distributed system and device to improve the efficiency of processing large-scale data. The specific technical solution is as follows.
[0005] In a first aspect, an embodiment provides a large-scale data processing method, including:
[0006] Storing multiple data from the original dataset into the cache;
[0007] When the amount of data in the cache reaches a preset value, the multiple data in the cache are sorted, and the sorted multiple data are written to the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted;
[0008] By running the designated database, multiple database files stored in the disk are loaded; the designated database can read data from the multiple database files in a manner equivalent to global sorting.
[0009] In one implementation, after loading the multiple database files stored in the disk, the method further includes: when data needs to be queried, reading data from the multiple database files in a manner equivalent to global sorting by specifying a database.
[0010] In one implementation, the method is performed by any first device in a second group of devices in a distributed system, the original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the second group of devices, wherein a first partition corresponds to the first device;
[0011] The step of storing a plurality of data from the original data set into the cache comprises:
[0012] A plurality of data belonging to the first partition in the original data set is stored in the cache.
[0013] In one implementation, the distributed system further includes a first group of devices;
[0014] The step of storing a plurality of data belonging to the first partition from the original data set into the cache comprises:
[0015] Receive multiple data belonging to the first partition sent by the first group of devices respectively, and store the multiple data in a cache; wherein the multiple data belonging to the first partition are screened out by each device in the first group of devices from their respective partial data, and the partial data are partial data in the original data set.
[0016] In one implementation, the distributed system further includes a third group of devices;
[0017] The step of reading data from the plurality of database files in a manner equivalent to global sorting comprises:
[0018] When receiving a data acquisition request sent by a device in the third group of devices, reading data from the multiple database files through the designated database in a manner equivalent to global sorting;
[0019] The method further includes: sending the read data to the third group of devices, so that the devices in the third group of devices perform specified processing on the data and send the processed data to the target end.
[0020] In one implementation, the step of loading the plurality of database files stored in the disk by running the designated database includes: recording the data extreme value of each database file;
[0021] The step of reading data from the multiple database files through the designated database in a manner equivalent to global sorting includes: using the designated database to read data from the multiple database files by comparing with the data extreme values of each database file.
[0022] In one implementation, the designated database includes a RocksDB database.
[0023] In one implementation, the method is executed based on a MapReduce framework, which includes a Map phase and a Reduce phase; the step of loading the multiple database files stored in the disk by running a specified database is executed in the Reduce phase.
[0024] In a second aspect, an embodiment provides a distributed system suitable for large-scale data processing, including a second group of devices consisting of a plurality of devices, wherein the second group of devices includes any first device;
[0025] The first device is used to store multiple data belonging to the first partition from the original data set into a cache; when the amount of data in the cache reaches a preset value, sort the multiple data in the cache, and write the sorted multiple data into the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all data contained in the multiple database files written to the disk are not globally sorted; by running a specified database, the multiple database files stored in the disk are loaded; the specified database can read data from the multiple database files in a manner equivalent to global sorting;
[0026] The original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the second group of devices, and the plurality of partitions include the first partition.
[0027] In a third aspect, an embodiment provides a large-scale data processing device, including:
[0028] a data cache module configured to store a plurality of data from an original data set into a cache;
[0029] A local sorting module is configured to sort the multiple data in the cache when the amount of data in the cache reaches a preset value, and write the sorted multiple data to the disk in the form of a specified database file, so that the data in any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted;
[0030] The file loading module is configured to load multiple database files stored in the disk by running a specified database; the specified database can read data from the multiple database files in a manner equivalent to global sorting.
[0031] In a fourth aspect, an embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute any one of the methods described in the first aspect.
[0032] In a fifth aspect, an embodiment provides a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the first aspects is implemented.
[0033] In the method and device provided in the embodiments of this specification, when multiple data in the cache are stored in the disk, the data are locally sorted so that each database file is locally sorted; then, by running the specified database to load the multiple database files stored in the disk, the specified database can read the data in a manner equivalent to globally sorting the multiple database files, without the need to globally sort the multiple database files written to the disk, thereby saving at least one read and write operation on the disk, thereby improving the efficiency of processing large-scale data. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 A schematic diagram of an implementation scenario of an embodiment disclosed in this specification;
[0036] Figure 2 A flowchart of a large-scale data processing method provided in an embodiment;
[0037] Figure 3 A schematic diagram of a flow chart of a distributed processing method for large-scale data provided in an embodiment;
[0038] Figure 4 A state diagram of a partition state machine provided in an embodiment;
[0039] Figure 5 A schematic block diagram of a distributed system suitable for large-scale data processing provided by an embodiment;
[0040] Figure 6 A schematic block diagram of a large-scale data processing device provided in an embodiment. DETAILED DESCRIPTION
[0041] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0042] Figure 1This is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification, which includes several computing devices, which can be used as node devices. Large-scale data is transmitted to the computing device, which performs task processing on the large-scale data and sorts the received data locally, that is, sorts the data inside each database file without globally sorting multiple database files. After executing the task processing, the computing device sends the data processing results to the target end. Here, the number of computing devices can be one or more.
[0043] Large-scale data refers to data with a very large number of items, such as millions, tens of millions, or even hundreds of millions of items, represented by N data, which constitute the original data set. Data can be data about an object, and each data or each item of data can include multiple attributes and corresponding attribute values. The objects here can include users, merchants, commodities, events, or transactions. For example, when the data is about transactions, the data may include the following attributes: transaction time, transaction amount, transaction currency, transaction method, and transaction accounts of both parties. The attribute value is the specific value of the corresponding attribute.
[0044] Task processing can be to sort N data based on the attribute value of a specified attribute, or to filter out the attribute value of a specified attribute from N data to meet a certain preset condition, or to filter out the target data from N data and then perform some processing on the filtered target data, etc. In short, the object of task processing is large-scale N data.
[0045] Global sorting of large-scale data is usually the basis for various task processing. Global sorting refers to sorting the attribute values of specified attributes of N data.
[0046] When using computing devices to process large-scale data, it is necessary to import the original data set from the data source, and after the processing is completed, the data needs to be imported into the target end. This process involves data transmission between different devices, as well as disk reading and writing inside the device. In addition, when computing devices process large-scale data, the processing process may be divided into different stages, and there may also be import and export of large-scale data. In summary, regardless of whether the number of computing devices performing task processing is one or more, the performance loss of large-scale data task processing is mainly concentrated in two aspects: network transmission and disk reading and writing.
[0047] Each computing device includes hardware such as CPU, cache and disk. The disk is used to persist data, while the cache is used to assist the CPU in processing data. The global sorting of large-scale data mentioned above involves a large amount of disk reading and writing. Especially when the amount of data is large, it is necessary to frequently read data from the disk for sorting and write the sorting results back to the disk, which increases the disk I / O (input / output) burden.
[0048] In order to improve the performance of computing devices when processing large-scale data, an embodiment of this specification provides a large-scale data processing method. In this method, the data in the cache is locally sorted and written to the disk in the form of a database file. In this way, multiple data are written to the disk in the form of multiple database files, and a specified database is run to load multiple database files in the disk. The specified database can read data from multiple database files in a manner equivalent to global sorting. This processing method does not require global sorting of multiple database files in the disk, which at least reduces one read and write operation of multiple database files in the disk. This embodiment can reduce performance loss from the perspective of reducing disk read and write.
[0049] Combine the following Figure 2 This embodiment is described in detail.
[0050] Figure 2 A flow chart of a large-scale data processing method provided in an embodiment. The method is executed by a computing device and includes the following steps. The computing device includes a cache and a disk.
[0051] Step S210: storing a plurality of data from the original data set into a cache.
[0052] The computing device can obtain data in the original data set from the source. The source can send the data in the original data set to the computing device one by one, or send the data to the computing device in the form of a data stream. The computing device stores the received multiple data in a cache. The original data set contains large-scale data.
[0053] Step S220, when the amount of data in the cache reaches a preset value, sort the multiple data in the cache, and write the sorted multiple data to the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted.
[0054] The data volume can be the amount of data, that is, the number of items or pieces, or the size of the data, for example, the amount of data measured in bytes. The preset value is preset according to parameters such as cache capacity.
[0055] When the amount of multiple data in the cache reaches a preset value, the multiple data are sorted in the cache. Specifically, the sorting operation can be performed according to the attribute value of the specified attribute of the multiple data. For example, when the data is a transaction, the multiple data can be sorted according to the transaction amount. After the multiple data are sorted, the multiple data can be generated into a database file in the form of a specified database file. The above operations are all performed in the cache. After the multiple data are written to the disk, the cache will continue to receive data, and the operation of step S220 is repeated.
[0056] That is, a large amount of data is stored in the cache, and database files are periodically generated through the cache, and the database files are written from the cache to the disk. In this way, the data inside each database file is sorted based on the attribute values of the specified attributes, that is, the data is sorted locally. This sorting process is completed in the cache, not after writing to the disk and then reading it out from the disk and then sorting it globally.
[0057] For the N data in the original data set, these N data are stored as n database files. The data in each database file is sorted, and the n database files are written to the disk. This process performs a disk write operation for the N data. Table 1 is an example of two database files.
[0058] Table 1
[0059]
[0060]
[0061] As can be seen from Table 1, the transaction data in database files 1 and 2 are sorted according to the transaction amount. Each transaction ID represents a transaction data. Table 1 only shows the transaction amount data in one transaction data. In practice, a transaction data may also include attribute data such as transaction time, transaction method and transaction device.
[0062] The multiple database files stored in the disk are similar to those in Table 1. Only the data within each database file is sorted, and no major sorting is performed between different database files. In addition, each database file has been sorted when it is first written to the disk, and will not be read from the disk for sorting after being written to the disk, thus saving disk read and write operations.
[0063] The above-mentioned designated database file may be a preset database file and a database file that can be recognized by the designated database. The designated database is a database that can read data from multiple database files in a manner equivalent to global sorting.
[0064] That is, the specified database has a special ability when reading data, that is, it can read data from multiple database files in a manner equivalent to a global sort, without having to read all database files from disk and globally sort all database files. Because the specified database can read data from multiple database files in a manner equivalent to a global sort of all data in multiple database files.
[0065] The specified database may be, but is not limited to, a RocksDB database, and the database file format that can be recognized by the RocksDB database may be an SST (Sorted String Table) file. Multiple data may be stored in the database file in the form of key-value pairs. The RocksDB database is an open source database that can provide the above special capabilities. The RocksDB database is a high-performance key-value pair storage database that is extended and optimized based on the LevelDB database.
[0066] Taking the RocksDB database as an example, when generating an SST file, some tools or methods can be used to generate an SST file from multiple sorted data. For example, the SstFileWriter function can be used to generate an SST file from multiple data according to the SST file format.
[0067] Step S230, loading multiple database files stored in the disk by running the specified database.
[0068] Step S240, after loading the multiple database files stored in the disk, when data needs to be queried, data is read from the multiple database files in a manner equivalent to global sorting by specifying a database.
[0069] In this embodiment, a designated database is used to load multiple database files that have been locally sorted, and data query is performed through the designated database, so that target data can be quickly queried from multiple database files. The target data is the data required in the set task processing, and the query operation on the designated database is performed under the task processing request. This fast data query process is equivalent to the query performed in multiple database files that have been globally sorted, so the query efficiency is very high. However, the embodiment does not actually implement the operation of globally sorting all the data in the multiple database files, that is, the read and write operations corresponding to the global sorting are not performed on the disk, so the I / O operations for the disk can be saved.
[0070] Specifically, when the number of database files stored in the disk reaches a preset number, the specified database can be run to load multiple database files stored in the disk. After loading multiple database files in the specified database, data can be read from the multiple database files through the specified database, that is, the disk can be read.
[0071] When loading multiple database files stored in the disk, the data extreme value of each database file can be recorded, and the specified database can be used to read data from multiple database files by comparing with the data extreme value of each database file, so as to achieve the effect of searching after globally sorting multiple database files. This is only one implementation method, but not the only one.
[0072] The data extreme value of each database file can be the extreme value of the attribute value of the specified attribute, including the maximum value max and the minimum value min. When data needs to be queried, for example, when transaction data with a transaction amount in the range of 20 to 50 needs to be queried, the specified database can compare the amount max and min of each database file with the range of 20 to 50 respectively, determine whether the target data exists in the database file, and the location of the target data, and then read the target data from the database file.
[0073] By running the specified database, the operation of loading multiple database files stored in the disk can be a lightweight load, that is, only the key value data of the primary key in each database file needs to be read from the disk, for example, only the value of the amount needs to be read, and the min and max of the amount need to be found, without reading the full data of each data. In addition, the database file contains the primary key data and non-primary key data of multiple data, and the primary key data is the data that is locally sorted according to its value.
[0074] Lightweight loading also requires reading data from the disk, but the amount of data read is very small, and only the primary key data is read. Lightweight loading is different from full data reading from the disk. Full data reading requires reading all the data in the database file on the disk, that is, it needs to read the primary key data and non-primary key data of the database file.
[0075] When you need to query data, you need to read all the data from multiple database files on the disk through the specified database. In other words, when you query and read data through the specified database, you need to read all the data from the disk.
[0076] It can be seen from the execution process of steps S210 to S240 that N data can be written to the disk in the form of multiple database files through steps S210 and S220. After the loading process of multiple database files by the specified database in step S230, a lightweight disk read process is performed. When querying data in step S240, a disk read process of N data on the disk is realized by specifying the database. The entire processing process saves reading multiple database files from the disk, then globally sorting them, and storing the globally sorted N data on the disk, such a full disk write and disk read process. The amount of data in the lightweight disk read process is much smaller than the amount of data in the full disk write and disk read process.
[0077] In actual applications, when a specified database loads multiple database files from a disk, it can use the open function to access the data file, determine the handle of the data file, and record the metadata information corresponding to the data file in memory. The metadata information includes the max and min values of the primary key data of the data file.
[0078] Taking RocksDB as an example, you can use its Bulkload capability to quickly load large batches of data into the database, thereby improving processing efficiency. Specifically, you can use the IngestExternalFile function to import multiple SST files into the RocksDB database. The RocksDB database uses its Bulkload capability to quickly add externally generated SST files to the database, thereby significantly improving the efficiency of data loading.
[0079] During the data loading process, you can turn off the compaction function of the RocksDB database and turn on the move_files configuration. Turning off the compaction function prevents the merge operation of multiple SST files. Turning on the move_files configuration allows you to configure the data file movement process, thereby reducing disk I / O operations.
[0080] When data needs to be queried, that is, when a query is received, the data corresponding to the query can be read from multiple database files in a manner equivalent to global sorting by specifying the database. The query can be to query data within a certain range of primary key data, or to query N data after global sorting, or to query the global max or min, etc.
[0081] The RocksDB database listed above is only one implementation. In actual applications, other databases can also read data from multiple database files in a manner equivalent to globally sorting multiple locally sorted database files when reading data. For example, when reading data, the database reads the extreme value (maximum value or minimum value) from each database file, and determines the location of the target data from each database file based on the extreme value. This data query method does not require the determination of the extreme value of each database file when loading multiple database files on the disk, but rather determines the extreme value of each database file when querying data from the database.
[0082] In the above embodiment, the computing device may be one or more. When the computing device is one, it is necessary to use the computing device to process N data. When there are multiple computing devices, each computing device can be used as a node device, and multiple node devices can jointly process N data to improve processing efficiency.
[0083] In another embodiment of the present specification, when there are multiple computing devices, that is, the above method can be executed by a distributed system, and the distributed system can include multiple devices, and the multiple devices form a second group of devices, including any first device. At the same time, the original data set is divided into a plurality of partitions, and the plurality of partitions correspond to the plurality of devices included in the second group of devices, and the first partition corresponds to the first device.
[0084] In step S210, the first device may store a plurality of data belonging to the first partition from the original data set into a cache. The first device may continue to perform the above steps S220-240.
[0085] When the original data set is divided into several partitions, the original data set can be divided according to the business logic, and multiple devices in the second group of devices can process the corresponding partition data respectively, so that the devices correspond to the partitions, thereby splitting the large-scale data into partial data, and completing the task processing through multiple devices. Among them, the correspondence between the device and the partition can be one-to-one or not, and one device can process one or more partitions.
[0086] For example, when the original data set contains N transaction data, different intervals can be divided according to the transaction amount: 0-20, 20-40, 40-60, 60-80 and 80-100. The corresponding partition data is processed by the corresponding device. For example, the first device can be used to process transaction data with a transaction amount in the interval of 20-40. Alternatively, different intervals can be divided according to the transaction ID: ID suffixes from 0 to 1, ID suffixes from 2 to 3, ID suffixes from 4 to 5, ID suffixes from 6 to 7, and ID suffixes from 8 to 9.
[0087] In step S230, when the first device loads the multiple database files belonging to the first partition in the original data set, the multiple database files belonging to the first partition may be taken as an example to load the multiple database files stored in the disk.
[0088] When the first device corresponds to multiple partition data, multiple database files of different partitions are loaded as different instances. For example, when the first device is responsible for processing data of two partitions, the database files of the two partitions can be loaded into the specified database as different instances. Each partition is used as an instance, so that the data of different partitions in the database can be kept independent and do not affect each other.
[0089] In one implementation, the distributed system can be built based on the MapReduce framework. The MapReduce framework is a programming model and computing framework that is mainly used to process and generate large-scale data sets. The MapReduce framework includes a Map phase and a Reduce phase. In a conventional MapReduce framework, the Map phase is used to map large-scale data into different partitioned data and locally sort the partitioned data. The Reduce phase is used to merge and sort the partitioned data.
[0090] In this embodiment, steps S210 to S220 can be executed in the Map phase, that is, the step of writing the locally sorted multiple data into the disk in the form of a specified database file, which is executed in the Map phase. Step S230, that is, the step of loading multiple database files stored in the disk by running the specified database, is executed in the Reduce phase. In this embodiment, the Reduce phase does not need to merge and sort the multiple database files, that is, it does not need to perform global sorting.
[0091] In another embodiment of the present specification, in order to improve the overall processing efficiency when processing large-scale data, complex processing steps can be assigned to different device groups for execution. Figure 3 This embodiment is described.
[0092] Figure 3A flow chart of a distributed processing method for large-scale data provided in an embodiment. The method is executed by a distributed system, and the distributed system can be built based on the MapReduce framework. The distributed system includes a first group of devices, worker1, a second group of devices, engine, and a third group of devices, worker3. The first group of devices, worker1, is connected to the source end through the network, and the source end device holds the original data set. The third group of devices, worker3, is connected to the target end device, and the target end device needs to obtain the data set after task processing. Each group of devices includes several devices. Each worker and engine is an independent node device, responsible for different execution steps.
[0093] Among them, any device in the first group of devices worker1 can obtain partial data from the original data set of the source end through network transmission, filter out data belonging to several partitions from the partial data, and send the data belonging to several partitions to the second group of devices engine corresponding to the partitions, including sending multiple data belonging to the first partition to the first device engine_1. Among them, any partial data contains data belonging to several partitions. For example, when the first group of devices includes m devices, the original data set can be split into m partial data from front to back, and each partial data is continuous data. Any device in the first device worker1 can filter out data belonging to several partitions from the partial data according to business logic, such as filtering out transaction data belonging to different transaction amount segments.
[0094] The first group of devices worker1 can also perform format conversion on the partition data, convert the partition data into a data format recognizable by the second group of devices engine, and send the format-converted partition data to the corresponding second group of devices engine. The format conversion provided by the first group of devices worker1 enables the embodiment to support multiple data sources, that is, the data in the original data set can include files, graph data, relational data, message queues and other types of data.
[0095] The first device engine_1, in step S210, receives a plurality of data belonging to the first partition respectively sent by the first group of devices worker1, and stores the plurality of data in a cache. The first group of devices worker1 and the second group of devices engine are transmitted via a network.
[0096] Next, the first device engine_1 executes steps S220 to S230. For details, see Figure 2 Each engine in the second group of engine devices executes steps similar to those of the first device engine_1 to process the partition data corresponding to it.
[0097] In step S240, the third group device worker3 may send a data acquisition request to each device engine in the second group device engine. When the first device engine_1 receives the data acquisition request sent by the device in the third group device worker3, it reads data from multiple database files in a manner equivalent to global sorting by specifying a database, and sends the read data to the corresponding third group device worker3.
[0098] The devices in the third group of devices worker3 perform specified processing on the received data and send the processed data to the target end. The specified processing may include format conversion, etc. The third group of devices worker3 may convert the received data into a data format that can be recognized by the target end.
[0099] The target end stores the data sent by the third group of devices worker3 into the source end data set.
[0100] The third group of devices, worker3, transmit data to the second group of devices, engine, through the network. The third group of devices, worker3, transmit data to the target end through the network.
[0101] The MapReduce framework includes a Map phase and a Reduce phase. In this embodiment, the steps performed in the Map phase include: the first group of device worker1 obtains part of the data from the original data set, filters out data belonging to several partitions from the part of the data, and sends the data belonging to several partitions to the corresponding second group of device engines, and the second group of device engines receive the partition data sent by the first group of device worker1, store it in the cache, and locally sort the partition data, and write the locally sorted multiple data to the disk in the form of a specified database file. The steps performed in the Reduce phase include: the second group of device engines loads multiple database files stored in the disk by running the specified database, and when the loading is completed, the specified database is in a state where data can be queried.
[0102] It can be seen that there are disk read and write operations in the second group of devices, engine, and there may be no disk read and write operations in the first group of devices, worker1, and the third group of devices, worker3. In this embodiment, the steps that require disk read and write operations are set in the second group of devices, engine, but not in the first group of devices, worker1, which can reduce the number of devices that need to perform disk read and write operations, making resource allocation more efficient and reasonable.
[0103] In one implementation, the second group of devices may also be directly connected to the target end without passing through the third group of devices. In other words, the second group of devices may directly read data from the specified database and send the data to the target end.
[0104] In one implementation, the above operations performed by the first group of devices worker1 may also be performed by the second group of devices, that is, the first group of devices may be combined with the second group of devices into one group of devices for implementation.
[0105] In this embodiment, by processing different simple steps respectively through different device groups, complex processing steps can be simplified and the efficiency of the overall processing process can be improved. In addition, generally speaking, the throughput of network transmission is much higher than the throughput of disk. Adding different device groups increases the network bandwidth much less than the saved disk bandwidth, so the overall running time can be effectively reduced, thereby reducing resource consumption. After actual testing, it was found that compared with the traditional MapReduce architecture, for the same data source, under the condition of the same data processing time, the CPU resource usage in this embodiment can be reduced by more than 40%, and the storage usage can be reduced by 30% to 50%.
[0106] The embodiments of this specification also provide a state machine for partition data. Partition data is data belonging to a partition in the original data set. At a certain processing node, the device can set different states of the partition data. Figure 4 Describe the partition data partition state machine.
[0107] Figure 4 A state diagram of a partition state machine provided for an embodiment. The following states are included: CREATED, SEALED, INGESTING, READY, DELETED, and ERROR. The first three states are process states, and the last three states are termination states. CREATED and SEALED are states of the map phase, and NGESTING and READY are states of the reduce phase. The first four states CREATED, SEALED, INGESTING, READY, and DELETED are sequential changes of the states of the partition partition under correct circumstances, and these four states constitute a state change cycle.
[0108] Among them, the CREATED state is set by the second group of devices engine, indicating that the correspondence between the partition partition and the engine device has just been established, and the partition data is written to the second group of devices engine, that is, the partition data is being distributed and the data cannot be consumed.
[0109] The SEALED state means that all data of the partition has been sent to the engine, and additional data writing to the partition is prohibited. All cached data will be forced to be flushed, that is, the sorted data in the cache will be written to the disk. The second group of devices, the engine, sets this state after determining that all data of the partition has been sent by the first group of devices, worker1.
[0110] The INGESTING state is set by the second group of devices, engines, and indicates that when all data of the partition has been flushed, you can start building a RocksDB instance, start the RocksDB database, and load all database files of the partition. In this state, flushing data is prohibited.
[0111] The READY state is set by the second group of equipment engines and is a state where data can be queried. This means that the RocksDB instance has been built, all data files for the partition have been loaded, and each data file is arranged in order. The RocksDB database can provide query services for all data in the partition.
[0112] The DELETED state is set by the second group of equipment engines, indicating that all data of the partition has been sent to the target end, that is, consumption is completed, and the storage space of the second group of equipment engines can be released. In all other states, when an exception occurs, the state of the partition can be converted to this state.
[0113] ERROR can be set for all devices, indicating an abnormal state, in which data cannot be read or written. In all other states, when an abnormality occurs, the state of the partition can be converted to this state.
[0114] In this embodiment, setting a partition state machine can manage the data processing process more effectively and efficiently.
[0115] In this specification, the word "first" in terms such as the first group of devices, the first device, the first partition, etc., and the corresponding "second" (if any) in the text, etc., are merely for the convenience of distinction and description and do not have any limiting meaning.
[0116] In this specification, the computing device, the first group of devices, the second group of devices and the third group of devices can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.
[0117] The foregoing describes certain embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in an order different from that in the embodiments, and the desired results may still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] Figure 5 A schematic block diagram of a distributed system suitable for large-scale data processing provided in an embodiment. Figure 3 The distributed system 500 includes a second group of devices 520 consisting of a plurality of devices, and the second group of devices 520 includes any first device 523 .
[0119] The first device 523 is used to store multiple data belonging to the first partition from the original data set into the cache; when the amount of data in the cache reaches a preset value, sort the multiple data in the cache, and write the sorted multiple data into the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted; by running the specified database, the multiple database files stored in the disk are loaded; the specified database can read data from the multiple database files in a manner equivalent to global sorting;
[0120] The original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the first device 523, and the plurality of partitions include the first partition.
[0121] In one implementation, after loading the multiple database files stored in the disk, the first device 523 is further used to: when data needs to be queried, read data from the multiple database files in a manner equivalent to global sorting by specifying a database.
[0122] In one implementation, the distributed system 500 also includes a first set of devices 510;
[0123] Any device in the first group of devices 510 is used to obtain part of the data from the original data set, filter out data belonging to a plurality of partitions from the part of the data, and send the data belonging to the plurality of partitions to the second group of devices corresponding to the partitions, including sending a plurality of data belonging to the first partition to the first device 523;
[0124] The first device 523 is specifically configured to receive a plurality of data belonging to the first partition respectively sent by the first group of devices 510 , and store the plurality of data in a cache.
[0125] In one implementation, the distributed system further includes a third set of devices 530;
[0126] The first device 523 is specifically configured to read data from multiple database files in a manner equivalent to global sorting by specifying a database when receiving a data acquisition request sent by a device in the third group of devices 530;
[0127] The first device 523 is further used to send the read data to the third group of devices 530;
[0128] The devices in the third group of devices 530 are used to perform specified processing on the data received from the first device 523 and send the processed data to the target end.
[0129] In one implementation, the first device 523, when loading a plurality of database files stored in a disk by running a specified database, includes: recording a data extreme value of each database file;
[0130] The first device 523 is specifically configured to use a designated database to read data from a plurality of database files by comparing with data extreme values of each database file.
[0131] The above system embodiment corresponds to the method embodiment, and the specific description can refer to the description of the method embodiment part, which will not be repeated here. The system embodiment has the same technical effect as the corresponding method embodiment, and the specific description can refer to the corresponding method embodiment.
[0132] Figure 6 A schematic block diagram of a large-scale data processing device provided in an embodiment. Figure 2 and Figure 3 The method embodiment shown corresponds to the embodiment shown in the figure. The apparatus 600 is deployed in a computing device, and includes:
[0133] A data cache module 610, configured to store a plurality of data from an original data set into a cache;
[0134] The local sorting module 620 is configured to sort the multiple data in the cache when the amount of data in the cache reaches a preset amount, and write the sorted multiple data to the disk in the form of a specified database file, so that the data in any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted;
[0135] The file loading module 630 is configured to load multiple database files stored in the disk by running a specified database; the specified database can read data from multiple database files in a manner equivalent to global sorting.
[0136] In one implementation, the apparatus 600 further includes:
[0137] The data reading module 640 is configured to read data from the multiple database files stored in the disk in a manner equivalent to global sorting by specifying a database when data needs to be queried after the multiple database files are loaded.
[0138] In one implementation, the method is executed by any first device in the second group of devices in the distributed system, the original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the second group of devices, wherein the first partition corresponds to the first device. The data cache module 610 is specifically configured to store a plurality of data belonging to the first partition from the original data set into the cache.
[0139] In one implementation, the distributed system further includes a first group of devices. The data cache module 610 is specifically configured to: receive multiple data belonging to the first partition respectively sent by the first group of devices, and store the multiple data in the cache. The multiple data belonging to the first partition are selected by each device in the first group of devices from their respective partial data, and the partial data are partial data in the original data set.
[0140] In one implementation, the distributed system further includes a third group of devices. The data reading module 640 is specifically configured to: when receiving a data acquisition request sent by a device in the third group of devices, read data from multiple database files in a manner equivalent to global sorting by specifying a database. The apparatus 600 further includes: a data sending module 650. The data sending module 650 is configured to send the read data to the third group of devices, so that the devices in the third group of devices perform specified processing on the data and send the processed data to the target end.
[0141] In one implementation, the file loading module 630 is specifically configured to record the data extreme value of each database file. The data reading module 640 is specifically configured to read data from multiple database files by using a designated database and comparing with the data extreme value of each database file.
[0142] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0143] The present specification also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute Figures 1 to 4 Any of the methods described above.
[0144] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, Figures 1 to 4 Any of the methods described above.
[0145] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the storage medium and computing device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0146] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0147] The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the embodiments of the present invention in detail. It should be understood that the above description is only a specific implementation method of the embodiments of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A large-scale data processing method, comprising: Storing multiple data from the original dataset into the cache; When the amount of data in the cache reaches a preset value, the multiple data in the cache are sorted, and the sorted multiple data are written to the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted; By running the designated database, multiple database files stored in the disk are loaded; the designated database can read data from the multiple database files in a manner equivalent to global sorting.
2. The method according to claim 1, after loading the plurality of database files stored in the disk, further comprises: When data needs to be queried, data is read from the multiple database files through the designated database in a manner equivalent to global sorting.
3. The method according to claim 2, wherein the method is performed by any first device in a second group of devices in a distributed system, the original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the second group of devices, wherein a first partition corresponds to the first device; The step of storing a plurality of data from the original data set into the cache comprises: A plurality of data belonging to the first partition in the original data set is stored in the cache.
4. The method according to claim 3, wherein the distributed system further comprises a first group of devices; The step of storing a plurality of data belonging to the first partition from the original data set into the cache comprises: receiving a plurality of data belonging to the first partition respectively sent by the first group of devices, and storing the plurality of data in a cache; The multiple data belonging to the first partition are selected by each device in the first group of devices from their respective partial data, and the partial data are partial data in the original data set.
5. The method according to claim 3, wherein the distributed system further comprises a third group of devices; The step of reading data from the plurality of database files in a manner equivalent to global sorting comprises: When receiving a data acquisition request sent by a device in the third group of devices, reading data from the multiple database files through the designated database in a manner equivalent to global sorting; The method further comprises: The read data is sent to the third group of devices, so that the devices in the third group of devices perform specified processing on the data and send the processed data to the target end.
6. The method according to claim 2, wherein the step of loading the plurality of database files stored in the disk by running the specified database comprises: Record the data extreme value of each database file; The step of reading data from the plurality of database files through the designated database in a manner equivalent to global sorting comprises: The designated database is used to read data from the plurality of database files by comparing with the data extreme value of each database file. The method according to claim 1 , wherein the designated database comprises a RocksDB database.
8. According to the method of claim 1, the method is executed based on the MapReduce framework, and the MapReduce framework includes a Map stage and a Reduce stage; the step of loading the multiple database files stored in the disk by running the specified database is executed in the Reduce stage.
9. A distributed system suitable for large-scale data processing, comprising a second group of devices consisting of a plurality of devices, wherein the second group of devices includes any first device; The first device is used to store a plurality of data belonging to the first partition from the original data set into a cache; When the amount of data in the cache reaches a preset value, the multiple data in the cache are sorted, and the sorted multiple data are written to the disk in the form of a specified database file, so that the data inside any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted; By running a specified database, multiple database files stored in the disk are loaded; the specified database can read data from the multiple database files in a manner equivalent to global sorting; The original data set is divided into a plurality of partitions, the plurality of partitions respectively correspond to a plurality of devices included in the second group of devices, and the plurality of partitions include the first partition.
10. A large-scale data processing device, comprising: a data cache module configured to store a plurality of data from an original data set into a cache; A local sorting module is configured to sort the multiple data in the cache when the amount of data in the cache reaches a preset value, and write the sorted multiple data to the disk in the form of a specified database file, so that the data in any database file obtained are sorted, and all the data contained in the multiple database files written to the disk are not globally sorted; The file loading module is configured to load multiple database files stored in the disk by running a specified database; the specified database can read data from the multiple database files in a manner equivalent to global sorting.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.
12. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.