Data processing method and device
By building a system that coexists batch processing and stream processing in one system, the complex operation and maintenance problem in the stream batch separation architecture model is solved, and the rapidity of data processing and the rationality of resource utilization is achieved.
Patent Information
- Application Number
- CN202510379863.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing stream batch separation architecture model requires maintenance of two sets of code logic, resulting in high operation and maintenance complexity.
Build a system where batch processing and stream processing coexist, and realize that data is compatible with batch processing and stream processing in the same system by converting the to-process data into a unified data format and determining the processing mode based on the data characteristics and system state characteristics.
In a system, the rapid processing of offline data and real-time data is realized, reducing operation and maintenance costs and rationality of resource utilization.
Smart Images

Figure CN120277113A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of big data, and in particular, to a data processing method and apparatus. Background Art
[0002] With the development and application of information technology, a vast amount of various data can be generated at any time and place. Analyzing the vast amount of data is inseparable from big data technology. For example, currently, the main methods for processing real-time data information in the financial field are architectures that adopt a separated mode of batch processing and stream processing. The core idea of this architecture is to design independent processing links for real-time stream data and offline batch data respectively. Among them, a stream-batch separation architecture mode adopts the Lambda architecture proposed by Nathan Marz in 2012, which includes a data source, a batch processing system, a stream processing system, and resource management. Among them, the batch processing system and the stream processing system operate independently. Batch processing is used to process historical data or a large amount of static data, which emphasizes accuracy and integrity, while stream processing is used to process data in real time, which emphasizes speed and low latency.
[0003] Since the above-mentioned stream-batch separation architecture mode uses different technology stacks and architectures when in use, two sets of code logics need to be maintained, resulting in high operation and maintenance complexity. Summary of the Invention
[0004] The present application provides a data processing method to overcome the disadvantages of the traditional architecture of "either stream or batch", and to achieve the data processing effects of compatible batch processing and stream processing in one system, thereby reducing the operation and maintenance difficulty.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, which is applicable to a system where batch processing and stream processing coexist; the method includes: obtaining first data to be processed; processing the first data into second data with a unified data format; the unified data format is applicable to the batch processing and the stream processing; determining a first processing mode for processing the second data according to the data characteristics of the first data and / or the instant state characteristics of the system; the first processing mode is one of the batch processing and the stream processing; processing the second data through the first processing mode.
[0006] In the above solution, by constructing a system where batch processing and stream processing coexist, and in this system, by converting the first data to be processed into second data in a unified data format that is suitable for both batch processing and stream processing, then after determining the first processing mode for processing the second data, the second data can be directly processed according to the first processing mode. Through this method, the effect of quickly processing offline data and real-time data in a system where batch processing and stream processing coexist is achieved. At the same time, since there is no need to maintain two sets of code logics, the operation and maintenance cost is relatively low.
[0007] In a possible implementation method, after processing the second data through the first processing mode, the method further includes: during the process of processing the second data through the first processing mode, based on the processing state of the second data and / or the immediate state characteristics of the system, determining to switch the first processing mode to a second processing mode, where the second processing mode is a processing mode different from the first processing mode; storing the intermediate processing result of the second data, and after obtaining the intermediate processing result through the second processing mode, continuing to process the second data in the second processing mode.
[0008] In the above solution, by evaluating some states during the processing of the second data according to the first processing mode, including the processing of the second data and / or the immediate state characteristics of the system, determining whether to change the processing mode for processing the second data, and when it is determined that a change is needed, by storing the intermediate processing result of the second data, then after switching to the new processing mode, the second data can be continuously processed based on the stored intermediate processing result. Through this method, it is also realized that in a system where batch processing and stream processing coexist, even for the same data to be processed, the present application can process the data in two different processing modes based on the actual situation, so as to achieve the maximization of resource utilization and take into account the accuracy and real-time nature of data processing.
[0009] In a possible implementation method, the processing of the second data in the second processing mode includes: after converting the calculation logic in the first processing mode into the calculation logic in the second processing mode, processing the second data in the second processing mode.
[0010] In the above solution, after determining that the processing mode for the second data needs to be switched, by converting the calculation logic in the previous processing mode, that is, the first processing mode, into the calculation logic in the subsequent processing mode, that is, the second processing mode, the effect of quickly processing the same data with the new processing mode can be achieved.
[0011] In a possible implementation method, the data characteristics of the first data include at least one of the data volume and data latency; the immediate state characteristics of the system include at least one of the system's resource status and the priorities of various services within the system; the processing status of the second data includes at least one of the processing duration and the amount of processed data.
[0012] In the above solution, by defining the data characteristics of the first data, the immediate state characteristics of the system, and the processing status of the second data, when deciding which processing mode to use to process the data to be processed, it is possible to implement a decision-making processing mode in a comprehensive evaluation manner. Therefore, the processing mode determined by this method will also be more reasonable and closer to the actual data processing effect.
[0013] In a possible implementation method, determining a first processing mode for processing the second data according to the data characteristics of the first data and / or the immediate state characteristics of the system includes: if the data volume of the first data is lower than a first value and the data latency is lower than a second value, then determine that the first processing mode is the stream processing; otherwise, determine that the first processing mode is the batch processing; or if the resource status of the system is higher than a third value, then determine that the first processing mode is any one of the batch processing and the stream processing; or if the resource status of the system is lower than the third value, then determine the first processing mode for processing the second data according to the priority of the service to which the first data belongs.
[0014] In the above solution, by using the stream processing mode for the data to be processed with a small amount and high latency requirements, and vice versa using the batch processing mode, the effect of reasonably processing data can be achieved; and by not restricting the processing mode of the data to be processed when the system resources are sufficient, and by determining the corresponding processing mode according to the priority of the service to which the data to be processed belongs when the system resources are scarce, it is also a manifestation of reasonably using system resources to process data, achieving the effect of improving the efficiency of resource utilization under the consideration of actual conditions.
[0015] In a possible implementation method, obtaining the first data to be processed includes: obtaining data from different data sources through a plug-in unified interface; wherein, obtaining the first data to be processed from the real-time data source according to a time window or a quantity window; or obtaining the first data to be processed from the offline data source according to the primary key time range sharding or file size chunking.
[0016] In the above solution, by defining a unified interface that supports obtaining data from both real-time data sources and offline data sources, this meets the prerequisite for implementing both batch processing and stream processing modes in the same system.
[0017] In a possible implementation method, third data from the same data source is respectively processed through stream processing to obtain a first processing result and through batch processing to obtain a second processing result; the first processing result is stored in the real-time partition of the transactional data lake, and the second processing result is stored in the offline partition of the transactional data lake; wherein, the second processing result in the offline partition is used to correct the first processing result in the real-time partition.
[0018] In the above solution, the first processing result and the second processing result generated by respectively performing stream processing and batch processing on the third data from the same data source are stored in the real-time partition and the offline partition of the transactional data lake respectively. Then, according to the property of the same storage of the transactional data lake, if necessary later, the result in the offline partition can be used to correct the result in the real-time partition, ensuring the ultimate consistency of data in a system where batch processing and stream processing coexist.
[0019] In a second aspect, an embodiment of the present application provides a data processing device, which is applicable to a system where batch processing and stream processing coexist; the device includes: a data acquisition unit for acquiring first data to be processed; a data format conversion unit for processing the first data into second data with a unified data format; the unified data format is applicable to the batch processing and the stream processing; a processing mode determination unit for determining a first processing mode for processing the second data according to the data characteristics of the first data and / or the instant state characteristics of the system; the first processing mode is one of the batch processing and the stream processing; a processing unit for processing the second data through the first processing mode.
[0020] In the above solution, by constructing a system where batch processing and stream processing coexist, and in this system, by converting the first data to be processed into second data with a unified data format that is applicable to both batch processing and stream processing, then after determining the first processing mode for processing the second data, the second data can be directly processed according to the first processing mode. Through this method, the effect of quickly processing offline data and real-time data in a system where batch processing and stream processing coexist is achieved. At the same time, since there is no need to maintain two sets of code logics, the operation and maintenance cost is relatively low.
[0021] In a third aspect, an embodiment of the present application provides a computing device, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory and executing, according to the obtained program, any implementation method in the first aspect.
[0022] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute any implementation method as in the first aspect.
[0023] Fifthly, an embodiment of the present application provides a computer program product including computer-executable instructions for causing a computer to execute any implementation method as in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for description in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic diagram of a data processing method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a data processing device provided by an embodiment of the present application; Figure 3 It is a schematic diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0027] Regarding the problems of complex structure and high maintenance cost of two sets of codes existing in the current solution for processing massive data using a stream-batch separation architecture mode, the present application proposes a data processing method applicable to a system where batch processing and stream processing coexist. As Figure 1 shown, it is a schematic diagram of a data processing method provided by an embodiment of the present application, including: Step 101: Obtain first data to be processed.
[0028] In this step, the first data to be processed can be various types of data such as transaction volume data, financial statement data, customer historical information table, etc. in a financial scenario.
[0029] In addition, the source of the first data to be processed can be various types of data sources.
[0030] Step 102: Process the first data into second data with a unified data format; the unified data format is applicable to both the batch processing and the stream processing.
[0031] In this step, the metadata in the first data can be automatically parsed according to the data source type to generate second data with a unified data format, which can be used in the subsequent batch processing and stream processing modes.
[0032] For example, when this application collects the customer historical information table in the MySQL database, the table structure of this table can be automatically parsed and mapped to the Avro format. It should be noted that the Avro format is an example of the unified data format in this application, and this application is not limited to this data format.
[0033] Step 103: Determine a first processing mode for processing the second data according to the data characteristics of the first data and / or the instant state characteristics of the system; the first processing mode is one of the batch processing and the stream processing.
[0034] After converting the first data into second data with a unified data format in the previous step 102, in this step, it is necessary to determine the processing mode for the second data with a unified data format. Among them, there are two processing modes in this application: batch processing and stream processing. For how to determine the processing mode for the second data, this application determines it by combining factors in multiple dimensions. For example, it can be determined according to the data characteristics of the first data, or it can be determined according to the instant state characteristics of the system where the batch processing and the stream processing are located, or it can be jointly determined according to the data characteristics of the first data and the instant state characteristics of the system.
[0035] Step 104: Process the second data through the first processing mode.
[0036] In this step, after determining the first processing mode, the first processing mode can be used to process the second data with a unified data format. For example, if the first processing mode is determined to be batch processing, then batch processing can be used to process the second data; if the first processing mode is determined to be stream processing, then stream processing can be used to process the second data.
[0037] In the above solution, by constructing a system where batch processing and stream processing coexist, and in this system, by converting the first data to be processed into the second data in a unified data format that is suitable for both batch processing and stream processing, then after determining the first processing mode for processing the second data, the second data can be directly processed according to the first processing mode. Through this method, the effect of quickly processing offline data and real-time data in a system where batch processing and stream processing coexist is achieved. At the same time, since there is no need to maintain two sets of code logics, the operation and maintenance cost is relatively low. For step 101 above, in a possible implementation method, the obtaining of the first data to be processed includes: obtaining data from different data sources through a pluggable unified interface; wherein, obtaining the first data to be processed from a real-time data source according to a time window or a quantity window; or, obtaining the first data to be processed from an offline data source according to primary key time range sharding or file size chunking.
[0038] For example, a unified data source access specification can be defined first. This unified data source access specification can support data access from various types of data sources such as databases, files, message queues, and API interfaces. Then, by forming the unified data source access specification into a pluggable unified interface, based on this unified interface, data to be processed can be obtained from different types of data sources such as databases, files, message queues, and API interfaces. These data to be processed are the first data to be processed. Among them, the database can be, for example, a MySQL database, the file can be, for example, an Nginx Log or an HDFS file, and the message queue can be, for example, Kafka. Of course, the specific content of the database, file, and message queue is not limited to the examples.
[0039] In addition, for the two types of data sources, namely the message queue and the API interface in the example, they can be regarded as real-time data sources. The data existing in the real-time data sources in this application is called streaming data. For streaming data, data can be obtained by sharding according to a time window or a quantity window. Among them, the time window can be, for example, 1 minute, that is, streaming data is obtained every minute, and the quantity window can be, for example, every 1000 pieces, that is, streaming data is obtained every time 1000 pieces of data are filled.
[0040] For the two types of data sources, namely the database and the file in the example, they can be regarded as offline data sources. The data existing in the offline data sources in this application is called batch data. For batch data, data can be obtained by automatically sharding according to data characteristics. For example, for a MySQL database, data can be obtained by sharding according to the primary key time range, and for a file, data can be obtained by chunking according to the file size.
[0041] For step 103 above, in a possible implementation method, determining a first processing mode for processing the second data according to the data characteristics of the first data and / or the immediate state characteristics of the system includes: if the data volume of the first data is lower than a first value and the data latency is lower than a second value, determining that the first processing mode is the stream processing; otherwise, determining that the first processing mode is the batch processing; or if the resource status of the system is higher than a third value, determining that the first processing mode is any one of the batch processing and the stream processing; or if the resource status of the system is lower than the third value, determining the first processing mode for processing the second data according to the priority of the service to which the first data belongs.
[0042] In a possible implementation method, the data characteristics of the first data include at least one of the data volume and the data latency; the immediate state characteristics of the system include at least one of the resource status of the system and the priorities of each service in the system.
[0043] For example, when the data volume of the first data is less than 10,000 and the latency requirement is less than 1 second, the stream processing can be selected as the first processing mode. Here, the exemplified 10,000 is an exemplary illustration of the first value, and 1 second is an exemplary illustration of the second value. The present application does not limit the magnitudes of the first value and the second value.
[0044] Otherwise, when the data volume exceeds 1 million or complex aggregation is required, the present application can select the batch processing as the first processing mode.
[0045] In addition, by collecting in real time metrics such as the CPU and memory of a system where batch processing and stream processing coexist, and determining the resource status of the system based on the data of each metric, it is possible to allow both the stream processing and the batch processing as the first processing mode when the system resources are sufficient, that is, when the system load is small. When the system resources are tight, that is, when the system load is large, it is necessary to further determine the first processing mode for processing the second data in combination with the priority of the service to which the first data belongs. For example, an example for the case of tight system resources can be that during a promotion activity of a certain bank, there will be a sharp increase in the real-time customer transaction data volume. For this situation, the present application can automatically suspend the offline report service with a lower business priority and allocate system resources to the stream processing to process the customer transaction volume data with a higher business priority. It should be noted that the present application does not give a specific numerical description of the third value, and this value can be determined by those skilled in the art according to actual needs, and the present application does not make a limitation.
[0046] To further illustrate how to determine the first processing mode, the present application can also provide the following examples for scheme illustration: (1) According to the data volume in the data characteristics: when the data volume exceeds the threshold, stream processing is preferred to quickly digest real-time data, and the backlogged part is transferred to batch processing; when the data volume is lower than the threshold, batch processing is used to save resources.
[0047] (2) According to the data latency in the data characteristics (data latency can also be referred to as real-time requirement or delay tolerance): for risk control data with low latency tolerance, this application can enforce stream processing, while for monthly report data with high latency tolerance, this application prefers batch processing. For example, in the scenario where a user swipes a credit card, since it is necessary to determine whether it is a fraudulent swipe within 100 milliseconds, stream processing is required to analyze the characteristics of this transaction, and if an anomaly is found, it will be intercepted immediately; of course, if the stream processing times out due to network latency, batch processing can be switched to supplement historical data.
[0048] (3) According to the calculation accuracy in the data characteristics (calculation accuracy can also be referred to as the error range allowed by the business): for financial statement data with a high-precision requirement that the error must be less than 0.1%, this application can enforce batch processing, while for real-time card product popularity ranking data with low-precision requirements that allow some errors, this application allows stream processing.
[0049] (4) According to the priorities of each business in the system: for example, for the Double Eleven transaction data with high priority, this application preferentially allocates stream processing computing resources, while for the low-priority log data archiving, it queues up for batch processing.
[0050] For step 104 above, a possible implementation is that after processing the second data through the first processing mode, the method further includes: during the process of processing the second data through the first processing mode, based on the processing status of the second data and / or the immediate state characteristics of the system, determine to switch the first processing mode to a second processing mode, where the second processing mode is a processing mode different from the first processing mode; store the intermediate processing result of the second data, and after obtaining the intermediate processing result through the second processing mode, continue to process the second data in the second processing mode. Among them, the processing status of the second data includes at least one of processing duration and the amount of processed data.
[0051] For example, taking the statistics of bank transaction amounts as an example, when the traffic is stable during the day, stream processing can be used as the first processing mode to statistically calculate the transaction amount per minute in real time; while after a marketing campaign is launched in the evening, the data volume may surge from 10,000 records per minute to 1 million records per minute (exceeding the stream processing threshold) or the CPU usage rate reaches 90% (indicating resource tension). In this regard, this application will comprehensively consider the data volume and the resource status of the system, switch the first processing mode of stream processing to the second processing mode of batch processing, such as changing to statistically calculate the transaction amount every 5 minutes.
[0052] Meanwhile, stream processing persists the intermediate results of the current window (such as the 800,000 records that have been statistically calculated) to HDFS, and then batch processing loads the HDFS data and continues to process the remaining data (for example, the remaining data is 200,000 records) to generate the complete result. After the marketing campaign ends, switch back to stream processing, and the historical data generated by batch processing flows back to the real-time layer to correct the real-time statistical value.
[0053] When stream processing is switched to batch processing, the intermediate state of real-time calculation, such as the window aggregation result, is persisted to storage, such as stored in Redis, for batch processing to load and continue the calculation; when batch processing is switched to stream processing, the breakpoint is restored from storage and incremental data is subscribed. For another example, when the bank customer behavior analysis task needs to be switched from stream processing to batch processing, the independent visitor value of real-time statistics can be saved to Apache Doris, and batch processing reads from Doris and continues to perform deduplication calculation.
[0054] In a possible implementation method, the processing of the second data in the second processing mode includes: after converting the calculation logic in the first processing mode into the calculation logic of the second processing mode, processing the second data in the second processing mode.
[0055] For example, this application can first convert the calculation logic in the first processing mode into the calculation logic of the second processing mode according to the preset conversion calculation logic, and then process the second data according to the second processing mode. For example, for the case of switching the processing mode where the first processing mode is stream processing event-time window and the second processing mode is batch processing group by date, this application defines the logic of "card growth amount per hour". Under stream processing, the Apache Flink time window is used, and under batch processing, it is automatically switched to GROUP BY hour(order_time) of Spark SQL. In this way, after defining the switching method of the calculation logic of the processing mode, when it is necessary to perform the switching of the processing mode, the calculation logic of the processing mode can be quickly converted and the latest processing mode, that is, the second processing mode, can be used to process the data.
[0056] In a possible implementation method, the third data from the same data source is respectively processed through stream processing to obtain a first processing result and through batch processing to obtain a second processing result; the first processing result is stored in the real-time partition of the transactional data lake, and the second processing result is stored in the offline partition of the transactional data lake; wherein, the second processing result in the offline partition is used to correct the first processing result in the real-time partition.
[0057] In the prior art, when using a stream-batch separation architecture mode to process the same data, since there is an expiration mechanism for Kafka messages relied on during stream processing, it is difficult to keep the stream processing result consistent with the batch processing result that depends on complete historical data. To address this problem, this application adopts the transactional data lake format (Apache Iceberg) as the unified storage, enabling the real-time computing layer and the batch processing layer to share the same data, ensuring ACID characteristics, and being compatible with Flink and Spark.
[0058] For example, for a piece of data from the same data source, that is, the third data, through stream processing, the generated result can be written into the real-time partition of Iceberg, and through batch processing, the generated result is written into the offline partition of Iceberg. Different partitions are distinguished by different version numbers, and the output result data automatically selects the latest valid partition according to the time range. For example, it preferentially reads the real-time partition and switches to the offline partition after timeout. For instance, Flink reads Kafka stream data, and according to the business logic, such as adding up account data by day and writing it into the Iceberg real-time partition, and Spark reads the complete data every day at midnight and writes it into the Iceberg offline partition according to the same business logic; if 10 accounts are undercounted in the real-time result due to network packet loss on a certain day, the records of the missing 10 accounts can be automatically extracted from the offline partition, backfilled to Kafka to trigger the stream processing to recalculate, so as to correct the result in the real-time partition.
[0059] In addition, this application can also version the data lake and use Iceberg tables to store the processing results. Each version contains the stream processing result (version_stream) and the batch processing correction result (version_batch). In this way, by retaining each historical version, it is convenient for subsequent problem backtracking.
[0060] Based on the same concept, an embodiment of this application also provides a data processing device, as Figure 2 shown. The device includes: A data acquisition unit 201, configured to acquire first data to be processed; A data format conversion unit 202 is configured to process the first data into second data with a unified data format; the unified data format is applicable to the batch processing and the stream processing; A processing mode determination unit 203 is configured to determine a first processing mode for processing the second data according to the data characteristics of the first data and / or the instant state characteristics of the system; the first processing mode is one of the batch processing and the stream processing; A processing unit 204 is configured to process the second data through the first processing mode.
[0061] Further, for this device, a switching unit 205 is further included; the switching unit 205 is configured to determine to switch the first processing mode to a second processing mode based on the processing state of the second data and / or the instant state characteristics of the system during the process of processing the second data through the first processing mode, and the second processing mode is a processing mode different from the first processing mode; the processing unit 204 is further configured to store the intermediate processing result of the second data, and after obtaining the intermediate processing result through the second processing mode, continue to process the second data through the second processing mode.
[0062] Further, for this device, the processing unit 204 is specifically configured to convert the calculation logic in the first processing mode into the calculation logic of the second processing mode, and then process the second data through the second processing mode.
[0063] Further, for this device, the data characteristics of the first data include at least one of the data volume and the data latency; the instant state characteristics of the system include at least one of the resource status of the system and the priorities of each service in the system; the processing state of the second data includes at least one of the processing duration and the processed data volume.
[0064] Further, for this device, the processing mode determination unit 203 is specifically configured to determine that the first processing mode is the stream processing if the data volume of the first data is lower than a first value and the data latency is lower than a second value, otherwise determine that the first processing mode is the batch processing; or determine that the first processing mode is any one of the batch processing and the stream processing if the resource status of the system is higher than a third value; or determine the first processing mode for processing the second data according to the priority of the service to which the first data belongs if the resource status of the system is lower than the third value.
[0065] Further, for the device, the data acquisition unit 201 is specifically configured to obtain data from different data sources through a plug-in unified interface; wherein, the first data to be processed is obtained from a real-time data source according to a time window or a quantity window; or, the first data to be processed is obtained from an offline data source according to a primary key time range sharding or a file size chunking.
[0066] Further, for the device, it further includes a storage unit 206; the processing unit 204 is further configured to respectively obtain a first processing result through stream processing and a second processing result through batch processing for the third data from the same data source; the storage unit 206 is configured to store the first processing result in the real-time partition of the transactional data lake and store the second processing result in the offline partition of the transactional data lake; wherein, the second processing result in the offline partition is used to correct the first processing result in the real-time partition.
[0067] An embodiment of the present application further provides a computing device, which may specifically be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), etc. The computing device may include a central processing unit (Center Processing Unit, CPU), a memory, an input / output device, etc. The input device may include a keyboard, a mouse, a touch screen, etc., and the output device may include a display device, such as a liquid crystal display (Liquid Crystal Display, LCD), a cathode ray tube (Cathode Ray Tube, CRT), etc.
[0068] The memory may include a read-only memory (ROM) and a random access memory (RAM), and provide program instructions and data stored in the memory to the processor. In the embodiment of the present application, the memory may be used to store program instructions of the data processing method; The processor is configured to call the program instructions stored in the memory and execute the data processing method according to the obtained program.
[0069] As Figure 3 shown, it is a schematic diagram of a computing device provided by an embodiment of the present application. The computing device includes: A processor 301, a memory 302, a transceiver 303, and a bus interface 304; wherein, the processor 301, the memory 302, and the transceiver 303 are connected through a bus 305; The processor 301 is configured to read the program in the memory 302 and execute the above data processing method; The processor 301 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. It may also be a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0070] The memory 302 is used to store one or more executable programs and can store the data used by the processor 301 when performing operations.
[0071] Specifically, the program may include program code, and the program code includes computer operation instructions. The memory 302 may include a volatile memory, such as a random-access memory (RAM); the memory 302 may also include a non-volatile memory, such as a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 302 may also include a combination of the above types of memories.
[0072] The memory 302 stores the following elements, executable modules, or data structures, or subsets thereof, or extended sets thereof: Operation instructions: including various operation instructions for implementing various operations.
[0073] Operating system: including various system programs for implementing various basic services and processing hardware-based tasks.
[0074] The bus 305 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 it is only represented by a thick line in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0075] The bus interface 304 can be a wired communication access port, a wireless bus interface, or a combination thereof. Among them, the wired bus interface can be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless bus interface can be a WLAN interface.
[0076] The embodiments of the present application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a data processing method.
[0077] The embodiments of the present application also provide a computer program product including computer-executable instructions for causing a computer to execute a data processing method.
[0078] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the function specified in the flowchart Figure 1 a flowchart or multiple flowcharts and / or blocks Figure 1 the function specified in a block or multiple blocks.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in the flowchart Figure 1 a flowchart or multiple flowcharts and / or blocks Figure 1 the function specified in a block or multiple blocks.
[0082] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data processing method, characterized in that, A system applicable to the coexistence of batch processing and stream processing; the method includes: Obtain first data to be processed; Process the first data into second data with a unified data format; the unified data format is applicable to the batch processing and the stream processing; Determine a first processing mode for processing the second data according to the data characteristics of the first data and / or the immediate state characteristics of the system; the first processing mode is one of the batch processing and the stream processing; Process the second data through the first processing mode.
2. The method according to claim 1, wherein After processing the second data through the first processing mode, the method further includes: During the process of processing the second data through the first processing mode, determine to switch the first processing mode to a second processing mode based on the processing state of the second data and / or the immediate state characteristics of the system, and the second processing mode is a processing mode different from the first processing mode; Store the intermediate processing result of the second data, and after obtaining the intermediate processing result through the second processing mode, continue to process the second data in the second processing mode.
3. The method according to claim 2, wherein The processing of the second data in the second processing mode includes: After converting the calculation logic in the first processing mode into the calculation logic of the second processing mode, process the second data in the second processing mode.
4. The method according to any one of claims 1 to 3, wherein The data characteristics of the first data include at least one of the data volume and the data latency; The immediate state characteristics of the system include at least one of the resource status of the system and the priorities of each service in the system; The processing state of the second data includes at least one of the processing duration and the amount of processed data.
5. The method according to claim 4, wherein The determining of the first processing mode for processing the second data according to the data characteristics of the first data and / or the immediate state characteristics of the system includes: If the data volume of the first data is lower than a first value and the data latency is lower than a second value, determine that the first processing mode is the stream processing, otherwise determine that the first processing mode is the batch processing; or If the resource status of the system is higher than a third value, determine that the first processing mode is any one of the batch processing and the stream processing; or If the resource status of the system is lower than the third value, determine the first processing mode for processing the second data according to the priority of the service to which the first data belongs.
6. The method according to any one of claims 1 to 3, characterized in that The obtaining of the first data to be processed includes: Obtain data from different data sources through a pluggable unified interface; wherein, obtain the first data to be processed from a real-time data source according to a time window or a quantity window; or, obtain the first data to be processed from an offline data source according to a primary key time range sharding or a file size chunking.
7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The third data from the same data source is respectively processed by stream processing to obtain a first processing result and by batch processing to obtain a second processing result; The first processing result is stored in the real-time partition of the transactional data lake, and the second processing result is stored in the offline partition of the transactional data lake; wherein, the second processing result in the offline partition is used to correct the first processing result in the real-time partition.
8. A data processing device, characterized in that, Applicable to a system where batch processing and stream processing coexist; the device includes: A data acquisition unit for acquiring first data to be processed; A data format conversion unit for processing the first data into second data with a unified data format; the unified data format is applicable to the batch processing and the stream processing; A processing mode determination unit for determining a first processing mode for processing the second data according to the data characteristics of the first data and / or the instant state characteristics of the system; the first processing mode is one of the batch processing and the stream processing; A processing unit for processing the second data through the first processing mode.
9. A computing device, characterized in that, Including: A memory for storing program instructions; A processor for calling the program instructions stored in the memory and executing the method according to any one of claims 1-7 according to the obtained program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method according to any one of claims 1-7.
Citation Information
Cited By
Stream batch collaborative data governance method
CN121858579A