Small file processing method, device, apparatus and storage medium

CN116775574BActive Publication Date: 2026-10-09CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310745976.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-10-09
Estimated Expiration
2043-06-21

AI Technical Summary

Benefits of technology

[0028] The small file processing method, device, apparatus, and storage medium provided in this application obtain the pending event data and the number of tasks during the operation of the Flink framework. Based on the pending event data, the event occurrence time is obtained, and the data type (e.g., normal data or late data) is determined according to the event occurrence time and the current physical time. The corresponding task index value is obtained based on the event time field and the number of tasks in the pending event data. The target subtask is then obtained based on the task index value. This achieves the processing of pending data using appropriate target subtasks according to different data types. The processed event data is stored in the directory file corresponding to the target subtask. In other words, by using an appropriate number of subtasks to process the data, the visibility time of the partitioned directory file to the downstream system is avoided due to task backpressure. Thanks to Flink's operator chain mechanism, the allocated pending data is written to the same subtask without worrying about data being partitioned again in the sink operator, causing data inconsistency and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116775574B_ABST
    Figure CN116775574B_ABST
Patent Text Reader

Abstract

The application provides a small file processing method and device, an apparatus and a storage medium, and relates to the technical field of big data stream computing. The method obtains to-be-processed event data and the number of tasks in the running process of the Flink framework, obtains the event occurrence time based on the to-be-processed event data, judges the data type as normal data or late data according to the event occurrence time, obtains the corresponding task index value according to the data type, obtains the target sub-task to process the to-be-processed data according to the task index value, and stores the processed data into the directory file corresponding to the target sub-task, that is, the data is processed by using a proper number of sub-tasks, the situation that the partition directory file is visible to the downstream system for a long time due to task back pressure is avoided, the inconsistency of data is avoided, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data streaming computing technology, specifically to a method, device, apparatus, and storage medium for processing small files. Background Technology

[0002] When Flink computes and outputs streaming results, it generates a large number of small files. Generally speaking, the higher the degree of parallelism used by Flink when writing streaming results to the downstream file system, that is, the more subtasks are used, the more small files are generated. Furthermore, the size of the small files generated by Flink is also related to the amount of data traffic and the data writing time. An excessive number of small files will slow down the retrieval speed, which in turn will slow down the speed of reading and writing data within the small files.

[0003] To address the increased operational burden caused by the number and size of small files, existing technologies generally employ two approaches: "in-process optimization" and "post-file-to-disk optimization." In-process optimization reduces the number of files in the entire partition directory by increasing the rolling strategy, checkpoints, and reducing Flink parallelism. Post-file-to-disk optimization uses an automatic merging strategy to merge a batch of small files into one or a few large files, which can also reduce the number of small files. However, both existing solutions, due to the increased number of operations, prolong the time it takes for partition directory files to become visible to downstream systems, and also introduce data consistency issues.

[0004] Therefore, existing technologies still fall short in time-partitioned business scenarios, particularly in terms of extending the visibility time of partitioned directory files to downstream systems due to the number and size of small files being processed, and in terms of data consistency. Summary of the Invention

[0005] This application provides a small file processing method, device, apparatus, and storage medium to solve the problems in the prior art where the visibility time of partitioned directory files to downstream systems is prolonged and data consistency is still lacking due to the number and size of small files being processed.

[0006] Firstly, this application provides a method for processing small files, including:

[0007] Obtain the pending event data and the number of tasks during the Flink operation, and obtain the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks under the Flink framework used to process the event data.

[0008] The time difference is obtained based on the time of the event occurrence and the current physical time.

[0009] If the time difference value is greater than the late time difference threshold, the current pending event data is confirmed to be late data, and the task index value is obtained according to the event time field and the number of tasks.

[0010] Based on the task index value, a target subtask is obtained, the event data to be processed is processed through the target subtask, and the processed data is stored in the directory file corresponding to the target subtask.

[0011] In one possible design, obtaining the task index value based on the event time field and the number of tasks includes: extracting the event days and event hours from the event time field; obtaining the actual number of hours based on the event days and event hours; and taking the remainder of the actual number of hours divided by the number of tasks to obtain the task index value.

[0012] In one possible design, the method further includes: if the time difference value is not greater than the late arrival time difference threshold, confirming that the current pending event data is normal data, extracting the event minutes from the event time field, determining the target number of minutes based on the time range to which the event minutes belong, obtaining the ideal parallelism based on the target number of minutes and the number of tasks, the ideal parallelism being used to indicate the maximum number of the target subtasks processing the pending event data under the target number of minutes, and determining the task index value based on the ideal parallelism.

[0013] In one possible design, obtaining the ideal parallelism based on the target number of minutes and the number of tasks includes: obtaining a first ratio of the target number of minutes to a preset number of minutes, and determining a first parallelism based on the product of the first ratio and the number of tasks, wherein the preset number of minutes is determined based on the time partition boundary; obtaining the sum of the first parallelism and the number of tasks; obtaining a second ratio of the difference between the preset number of minutes and the target number of minutes to the preset number of minutes, and determining a second parallelism based on the product of the second ratio and half of the sum; and obtaining the ideal parallelism based on the sum of the first parallelism and the second parallelism.

[0014] In one possible design, determining the task index value based on the ideal parallelism includes: generating a random number based on the ideal parallelism, and determining the task index value based on the random number.

[0015] In one possible design, if the time range is the first half hour, then the target number of minutes is the event minutes; if the time range is the second half hour, then the target number of minutes is the difference between the hourly minutes and the event minutes.

[0016] In one possible design, if the time range is the latter half hour, determining the task index value based on the random number includes: obtaining the candidate index corresponding to the target subtask based on the difference between the number of tasks and the calibration value, obtaining the difference between the candidate index and the random number, and obtaining the task index value.

[0017] In one possible design, determining the task index value based on the ideal parallelism includes: using multiple values ​​prior to the ideal parallelism as an index set, wherein the N-1 value in the index set is the Nth task index value; when new data to be processed is detected and is still normal data, taking the remainder of the sum of the Nth task index value and the calibration value with respect to the ideal parallelism to obtain the N+1th task index value, until no new data to be processed is detected or the new data to be processed is abnormal data.

[0018] In one possible design, obtaining the event time field from the event data to be processed includes: extracting the event time field from the event data to be processed using a keyword extraction strategy set in the Keyselector operator.

[0019] Secondly, this application provides a small file processing device, comprising:

[0020] The acquisition module is used to acquire the pending event data and the number of tasks during the Flink operation, and to acquire the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks under the Flink framework used to process the event data.

[0021] The first processing module is used to obtain the time difference based on the time of the event occurrence and the current physical time;

[0022] The second processing module is used to confirm that the current event data to be processed is late data if the time difference value is greater than the late time difference threshold, and to obtain the task index value according to the event time field and the number of tasks.

[0023] The execution module is used to obtain the target subtask based on the task index value, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask.

[0024] Thirdly, this application provides a small file processing device, including: a processor, and a memory communicatively connected to the processor;

[0025] The memory stores computer-executed instructions;

[0026] The processor executes computer execution instructions stored in the memory to implement the small file processing method.

[0027] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement a small file processing method.

[0028] The small file processing method, device, apparatus, and storage medium provided in this application obtain the pending event data and the number of tasks during the operation of the Flink framework. Based on the pending event data, the event occurrence time is obtained, and the data type (e.g., normal data or late data) is determined according to the event occurrence time and the current physical time. The corresponding task index value is obtained based on the event time field and the number of tasks in the pending event data. The target subtask is then obtained based on the task index value. This achieves the processing of pending data using appropriate target subtasks according to different data types. The processed event data is stored in the directory file corresponding to the target subtask. In other words, by using an appropriate number of subtasks to process the data, the visibility time of the partitioned directory file to the downstream system is avoided due to task backpressure. Thanks to Flink's operator chain mechanism, the allocated pending data is written to the same subtask without worrying about data being partitioned again in the sink operator, causing data inconsistency and improving the user experience. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 Flowchart of the small file processing method provided in this application embodiment Figure 1 ;

[0031] Figure 2 Flowchart of the small file processing method provided in this application embodiment Figure 2 ;

[0032] Figure 3 This is a schematic diagram of the structure of the small file processing device provided in the embodiments of this application;

[0033] Figure 4 This is a schematic diagram of the structure of a small file processing device provided in an embodiment of this application. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0035] First, the relevant concepts or terms involved in this application will be explained:

[0036] Apache Flink is an open-source stream processing framework. At its core is a computing framework and distributed processing engine for performing stateful computations on unbounded and bounded data streams. Flink executes arbitrary streaming data programs in a data-parallel and pipelined manner. Flink's pipeline runtime system can execute batch and stream processing programs.

[0037] Stream computing is a high-frequency, incremental, and real-time data processing mode, mainly suitable for scenarios with high performance requirements such as small memory footprint, fast single-processing, and low system latency, as well as scenarios that require continuous computation within a task. Flink is a common framework for implementing stream computing tasks.

[0038] When the Flink framework computes the output streaming results, it generates a large number of small files. The number of small files is related to the degree of parallelism used by Flink. That is, the higher the degree of parallelism used, the more small files are generated. At the same time, the size of the small files generated by Flink is also related to the amount of data traffic and the data writing time. An excessive number of small files will lead to excessive index resource consumption and slower retrieval speed, resulting in slower reading and writing of data within the small files.

[0039] Existing technologies generally control the number and size of small files generated through two approaches: "in-process optimization" and "post-file write optimization." In-process optimization reduces the parallelism of Flink's file writing to downstream file systems by increasing the rolling strategy and checkpoint frequency, thereby reducing the number of files in the entire partition directory. Post-file write optimization controls the reduction of small file numbers through file merging procedures. However, both methods can cause task backpressure during peak data traffic periods, thus prolonging the visibility time of partition directory files to downstream systems. Therefore, a method is needed that can process data using an appropriate number of subtasks under different data traffic scenarios to avoid the situation where task backpressure prolongs the visibility time of partition directory files to downstream systems, thereby improving the user experience.

[0040] This application provides a small file processing method that obtains pending event data and the number of tasks during the Flink framework's operation. Based on the pending event data, it obtains the event occurrence time and determines the data type (e.g., normal data or late data) according to the event occurrence time and current physical time. Based on the event time field and the number of tasks in the pending event data, it obtains the corresponding task index value and uses the task index value to obtain the target subtask. This achieves the processing of pending data using appropriate target subtasks based on different data types, and the processing of pending event data through target subtasks. The processed data is then stored in the directory file corresponding to the target subtask. In other words, by using an appropriate number of subtasks to process the data, it avoids the situation where task backpressure extends the visibility time of partition directory files to downstream systems, thus improving the user experience.

[0041] The technical solutions of this application and how they solve the aforementioned technical problems are described in detail below using specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0042] Example 1

[0043] Figure 1 Flowchart of the small file processing method provided in this application embodiment Figure 1 .like Figure 1 As shown, the method includes:

[0044] S101. Obtain the pending event data and the number of tasks during the Flink operation, and obtain the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks used to process event data under the Flink framework.

[0045] Specifically, the event data to be processed input through the Flink framework is inspected, and based on the input event data, the event time field related to the event data is obtained. In particular, through a custom keyword retrieval strategy, the value of the event time field can be extracted from the relevant event time field as a key value field, and the event occurrence time can be obtained based on the key value field.

[0046] S102. Obtain the time difference based on the time of the event occurrence and the current physical time;

[0047] Specifically, after obtaining the event occurrence time through a custom keyword acquisition strategy, the difference between the event occurrence time and the current physical time is obtained, i.e., the time difference. The obtained time difference is compared with a preset time difference threshold, and the type of data to be processed, such as late data or normal data, is determined based on the comparison result with the time difference threshold.

[0048] S103. If the time difference value is greater than the late time difference threshold, then the current pending event data is confirmed to be late data, and the task index value is obtained according to the event time field and the number of tasks.

[0049] Specifically, after obtaining the time difference between the event occurrence time and the current physical time, the time difference is compared with a preset time difference threshold. When the obtained time difference is less than the late time difference threshold, the data type of the current data to be processed is confirmed to be normal data. When the obtained time difference is greater than the late time difference threshold, the data type of the current data to be processed is confirmed to be late data. The task index value corresponding to the late data is obtained by extracting the event time field and obtaining the number of tasks.

[0050] S104. Obtain the target subtask according to the task index value, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask.

[0051] Specifically, when the time difference is compared with the preset time difference threshold, the data type of the current data to be processed is confirmed. After obtaining the matching task index value based on the data type, the corresponding target subtask number is indexed according to the task index value. The target subtask that needs to process the data to be processed is confirmed through the corresponding target subtask number. The data after the target subtask is processed is stored in the directory file, where the directory file corresponds to the subtask.

[0052] This application provides a small file processing method that obtains pending event data and the number of tasks during the Flink framework's operation. Based on the pending event data, it obtains the event occurrence time and determines the data type (e.g., normal data or late data) according to the event occurrence time and current physical time. Based on the event time field and the number of tasks in the pending event data, it obtains the corresponding task index value and uses the task index value to obtain the target subtask. This achieves the processing of pending data using appropriate target subtasks based on different data types, and the processing of pending event data through target subtasks. The processed data is then stored in the directory file corresponding to the target subtask. In other words, by using an appropriate number of subtasks to process the data, it avoids the situation where task backpressure extends the visibility time of partition directory files to downstream systems, thus improving the user experience.

[0053] The following specific embodiment will be used to describe the small file processing method of this application in detail.

[0054] Example 2

[0055] Figure 2 Flowchart of the small file processing method provided in this application embodiment Figure 2 .like Figure 2 As shown, the method includes:

[0056] S201. Obtain the pending event data and the number of tasks during the Flink operation, and obtain the event time field based on the pending event data, wherein the event time field includes the event occurrence time;

[0057] Among them, the keyword crawling strategy set in the Keyselector operator based on the Flink framework crawls the event time field from the event data to be processed;

[0058] Specifically, the event data to be processed input through the Flink framework is inspected, and based on the input event data, the keyword crawling strategy set in the Keyselector operator is implemented to extract the event time field related to the event occurrence time from the event data. The value of the event time field is used as the key value field, and the specific event occurrence time is obtained based on the extracted event time field.

[0059] S202. Obtain the time difference based on the time of the event occurrence and the current physical time;

[0060] Specifically, after implementing the keyword crawling strategy set in the Keyselector operator, the time difference is calculated based on the obtained event occurrence time and the current physical time. The obtained time difference is then compared with a preset time difference threshold to determine whether the current data to be processed is late data or normal data.

[0061] S203. If the time difference value is greater than the late time difference threshold, then the current pending event data is confirmed to be late data.

[0062] The late arrival time difference threshold is a calibrated empirical value and can be adjusted according to the actual business situation.

[0063] Specifically, the acquired time difference is compared with a preset time difference threshold. When the acquired time difference is greater than the late time difference threshold, it indicates that the time interval between the event occurrence time of the data to be processed and the current physical time is too far. In this case, the data type of the data to be processed is confirmed to be late data. The traffic of late data is smaller than that of other normal data in the same time partition, so a large number of subtasks are not needed to process the late data.

[0064] S204. Obtain the task index value based on the event time field and the number of tasks;

[0065] Specifically, the event days and event hours are extracted from the event time field; the actual number of hours is obtained based on the event days and the event hours; and the task index value is obtained by taking the remainder of the actual number of hours divided by the number of tasks.

[0066] Specifically, when the acquired time difference value is greater than the late time difference threshold, confirming that the data type of the current data to be processed is late data, the event days and event hours are extracted based on the event time field. Further, adjustments are made to ensure that the extracted event days are not less than one and the event hours are not less than zero. Based on the event days and event hours, the actual number of hours is obtained. The modulo operation is then performed on the acquired task quantity using the actual number of hours to obtain a task index value suitable for late data. This ensures that late data belonging to the same hour partition will calculate the same task index value, expressed by the following formula:

[0067] N=((dayOfYear-1)*24+hourOfDay)%numPartitions;

[0068] Where dayOfYear is the number of days in the event, hourOfDay is the number of hours in the event, numPartitions is the number of tasks, and N is the task index value applicable to late data.

[0069] S205. If the time difference is not greater than the late time difference threshold, confirm that the current event data to be processed is normal data, extract the event minutes from the event time field, and determine the target number of minutes according to the time range to which the event minutes belong.

[0070] Specifically, the time difference value is compared with the preset time difference threshold. When the time difference value is not greater than the late time difference threshold, it means that the interval between the event time of the data to be processed and the current physical time is small, and the data type of the data to be processed is confirmed to be normal data.

[0071] Furthermore, since the trigger times of the data to be processed in the Flink framework streaming computation are nearly evenly distributed on the time axis, the data traffic near the hourly partition boundary time point is less than the data traffic near the hourly partition center time point. This means that the number of subtasks that need to process the data near the hourly partition boundary time point is also less than the number of data near the hourly partition center time point. Therefore, the normal data is further divided into the first half hour data and the second half hour data according to the time range.

[0072] S206. If the time range is the first half hour, then the target number of minutes is the number of minutes of the event;

[0073] Specifically, when the obtained time difference is not greater than the late time difference threshold, and it is confirmed that the current normal data time range is within the first half hour, the number of event minutes related to the current data to be processed and the event occurrence is obtained based on the extracted event time field, and the obtained event minutes are used as the target number of minutes.

[0074] S207. Obtain the ideal parallelism based on the target number of minutes and the number of tasks;

[0075] The process involves obtaining a first ratio between the target number of minutes and a preset number of minutes, where the preset number of minutes is thirty. A first degree of parallelism is determined by multiplying the first ratio by the number of tasks. The preset number of minutes is determined based on the time partition boundaries. The sum of the first degree of parallelism and the number of tasks is obtained. A second ratio between the difference between the preset number of minutes and the target number of minutes and the preset number of minutes is obtained. A second degree of parallelism is determined by multiplying the second ratio by half of the sum. The ideal degree of parallelism is obtained by summing the first and second degrees of parallelism.

[0076] Specifically, after confirming that the current data to be processed is normal data from the first half hour, the ideal parallelism is obtained based on the target number of minutes and the number of tasks. The ideal parallelism indicates the maximum number of target subtasks that can process the data to be processed within the target number of minutes, that is, it describes the range of numbers of the selected target subtasks. Furthermore, the ideal parallelism is expressed by the following formula:

[0077] M=minOfHour / 30.0*numPartitions+(30-minOfHour) / 30.0*(minOfHour /

[0078] 30.0*numPartitions+numPartitions) / 2;

[0079] Where minOfHour is the target number of minutes, numPartitions is the number of tasks, and M is the ideal parallelism.

[0080] S208. Generate a random number based on the ideal parallelism, and determine the task index value based on the random number;

[0081] Specifically, after obtaining the ideal parallelism based on the target number of minutes and the number of tasks, a random integer is selected between zero and the obtained ideal parallelism. That is, a random number is generated based on the ideal parallelism, and the selected random integer is used as the task index value, which is to determine the task index value.

[0082] S209. If the time range is the last half hour, then the target number of minutes is the difference between the hourly minutes and the event minutes.

[0083] Specifically, after confirming that the current data to be processed is normal data in the latter half hour, based on the extracted event time field, the event minutes and hour minutes related to the current data to be processed are obtained, where the hour minutes are sixty. According to the obtained target minutes and the number of tasks, the ideal parallelism is obtained, that is, calculated according to the formula for obtaining the ideal parallelism in S207 above. The target minutes are updated to the difference between the hour minutes and the event minutes. After obtaining the ideal parallelism, a random integer between zero and the obtained ideal parallelism is selected as the random number.

[0084] S210. Based on the difference between the number of tasks and the calibration value, obtain the candidate index corresponding to the target subtask, and obtain the difference between the candidate index and the random number to obtain the task index value;

[0085] Specifically, after obtaining the random number, in order to balance the flow of downstream subtasks processing pending data and prevent subtasks with smaller numbers from accumulating too much pending data, a task index value suitable for the normal data of the latter half hour is obtained based on the random number and the number of tasks, expressed by the following formula:

[0086] Z=numPartitions-1-rand_index;

[0087] Where numPartitions is the number of tasks, rand_index is a random number, and Z is the task index value applicable to the normal data in the second half hour.

[0088] S211. Obtain the target subtask according to the task index value, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask;

[0089] Specifically, after obtaining the matching task index value based on the data type, such as late data or normal data, the target subtask number corresponding to the task index value is indexed, and the target subtask that needs to be processed is confirmed through the target subtask number corresponding to the Sink operator. Appropriate subtasks are allocated according to the processing data of different traffic volumes, and the data after the target subtask is processed is stored in the directory file corresponding to the subtask. In other words, the size of the generated file is controlled by controlling the number of subtasks.

[0090] In another embodiment, when the obtained time difference value is not greater than the late time difference threshold, it is confirmed that the current pending event data is normal data. After obtaining the ideal parallelism based on the obtained target minutes and number of tasks, multiple values ​​before the ideal parallelism are used as an index set. The N-1 value in the index set is the Nth task index value. When new pending data is detected and is still normal data, the N+1th task index value is obtained by taking the remainder of the sum of the Nth task index value and the calibration value with the ideal parallelism. This process continues until no new pending data is detected or the new pending data is abnormal data.

[0091] Specifically, after confirming that the current pending event data is normal data, and obtaining the ideal parallelism based on the target number of minutes and the number of tasks, a polling method is used to sequentially obtain integer values ​​between zero and the ideal parallelism as task index values. That is, multiple values ​​before the ideal parallelism are used as an index set. Since the ideal parallelism calculation method is fixed, the ideal parallelism calculated at different time points is also fixed. The previous task index value is recorded at each time point. When new pending data is detected, the previous task index value (i.e., the Nth task index value plus one) is used as the remainder after dividing by the ideal parallelism as the new task index value (i.e., the N+1th task index value). The subtask number of the current pending data flow is controlled according to the new task index value.

[0092] This application provides a small file processing method that obtains pending event data and the number of tasks during the Flink framework's operation. Based on the pending event data, it obtains the event occurrence time and determines the data type (e.g., normal data or late data) according to the event occurrence time and current physical time. Based on the event time field and the number of tasks in the pending event data, it obtains the corresponding task index value and uses the task index value to obtain the target subtask. This achieves the processing of pending data using appropriate target subtasks based on different data types, and the processing of pending event data through target subtasks. The processed data is then stored in the directory file corresponding to the target subtask. In other words, by using an appropriate number of subtasks to process the data, it avoids the situation where task backpressure extends the visibility time of partition directory files to downstream systems, thus improving the user experience.

[0093] In this embodiment of the invention, electronic devices or main control devices can be divided into functional modules according to the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional module. It should be noted that the module division in this embodiment of the invention is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0094] Figure 3 This is a schematic diagram of the small file processing device provided in an embodiment of this application. Figure 3 As shown, the device 300 includes:

[0095] The acquisition module 301 is used to acquire the pending event data and the number of tasks during the Flink operation, and to acquire the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks under the Flink framework used to process the event data.

[0096] The first processing module 302 is used to obtain a time difference based on the event occurrence time and the current physical time;

[0097] The second processing module 303 is used to confirm that the current event data to be processed is late data if the time difference value is greater than the late time difference threshold, and to obtain the task index value according to the event time field and the number of tasks.

[0098] The execution module 304 is used to obtain the target subtask according to the task index value, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask.

[0099] Furthermore, the second processing module 303 is specifically used to extract the event days and event hours from the event time field, obtain the actual number of hours based on the event days and event hours, and obtain the task index value by taking the remainder of the actual number of hours divided by the number of tasks.

[0100] Furthermore, the second processing module 303 is specifically used to confirm that the current event data to be processed is normal data if the time difference is not greater than the late time difference threshold, extract the event minutes in the event time field, determine the target minutes according to the time range to which the event minutes belong, obtain the ideal parallelism according to the target minutes and the number of tasks, the ideal parallelism is used to indicate the maximum number of the target subtasks for processing the event data to be processed under the target minutes, and determine the task index value according to the ideal parallelism.

[0101] Furthermore, the second processing module 303 is specifically used to obtain a first ratio between the target number of minutes and the preset number of minutes, and determine a first degree of parallelism based on the product of the first ratio and the number of tasks. The preset number of minutes is determined based on the time partition boundary. The module obtains the sum of the first degree of parallelism and the number of tasks, obtains a second ratio between the difference between the preset number of minutes and the target number of minutes and the preset number of minutes, and determines a second degree of parallelism based on the product of the second ratio and half of the sum. The ideal degree of parallelism is obtained based on the sum of the first degree of parallelism and the second degree of parallelism.

[0102] Furthermore, the second processing module 303 is specifically used to generate random numbers based on the ideal parallelism, and to determine the task index value based on the random numbers.

[0103] Furthermore, the second processing module 303 is also used to determine the target number of minutes as the event minutes if the time range is the first half hour, and the target number of minutes as the difference between the hourly minutes and the event minutes if the time range is the second half hour.

[0104] Furthermore, the second processing module 303 is specifically used to obtain the candidate index corresponding to the target subtask based on the difference between the number of tasks and the calibration value, obtain the difference between the candidate index and the random number, and obtain the task index value.

[0105] Furthermore, the second processing module 303 is specifically used to take multiple values ​​before the ideal parallelism as an index set, wherein the N-1 value in the index set is the Nth task index value. When new data to be processed is detected and is still normal data, the N+1th task index value is obtained by taking the remainder of the sum of the Nth task index value and the calibration value with respect to the ideal parallelism, until no new data to be processed is detected or the new data to be processed is abnormal data.

[0106] Furthermore, the first processing module 302 is specifically used to extract the event time field from the event data to be processed through the keyword extraction strategy set in the Keyselector operator.

[0107] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes at least one processor 401 and a memory 402. The electronic device 400 also includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.

[0108] In the specific implementation process, at least one processor 401 executes the computer execution instructions stored in the memory 402, causing at least one processor 401 to execute the small file processing method executed on the electronic device side as described above.

[0109] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0110] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0111] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.

[0112] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0113] The above description of the functions implemented by electronic devices and main control devices has introduced the solutions provided by the embodiments of the present invention. It is understood that, in order to implement the above functions, the electronic device or main control device includes hardware structures and / or software modules corresponding to the execution of each function. By combining the units and algorithm steps of the various examples described in the embodiments of the present invention, the embodiments of the present invention can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of the present invention.

[0114] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-mentioned small file processing method.

[0115] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0116] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in an electronic device or a host device.

[0117] This application also provides a computer program product, comprising: a computer program stored in a readable storage medium, wherein at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the scheme provided in any of the above embodiments.

[0118] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A small file processing method, applied to the Flink framework, characterized in that, The method includes: Obtain the pending event data and the number of tasks during the Flink operation, and obtain the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks under the Flink framework used to process the event data. The time difference is obtained based on the time of the event occurrence and the current physical time. If the time difference value is greater than the late time difference threshold, the current pending event data is confirmed to be late data, and the task index value is obtained according to the event time field and the number of tasks. Based on the task index value, obtain the target subtask, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask; If the time difference is not greater than the late arrival time difference threshold, the current pending event data is confirmed as normal data; extract the event minutes from the event time field; determine the target number of minutes based on the time range to which the event minutes belong; obtain a first ratio of the target number of minutes to a preset number of minutes, and determine a first degree of parallelism based on the product of the first ratio and the number of tasks, wherein the preset number of minutes is determined based on the time partition boundary; obtain the sum of the first degree of parallelism and the number of tasks; obtain a second ratio of the difference between the preset number of minutes and the target number of minutes to the preset number of minutes, and determine a second degree of parallelism based on the product of the second ratio and half of the sum; obtain the ideal degree of parallelism based on the sum of the first degree of parallelism and the second degree of parallelism; the ideal degree of parallelism is used to indicate the maximum number of the target subtasks processing the pending event data under the target number of minutes; determine the task index value based on the ideal degree of parallelism.

2. The method according to claim 1, characterized in that, The step of obtaining the task index value based on the event time field and the number of tasks includes: Extract the event days and event hours from the event time field; Based on the number of days and the number of hours of the event, obtain the actual number of hours; The task index value is obtained by taking the remainder of the actual number of hours divided by the number of tasks.

3. The method according to claim 1, characterized in that, The step of determining the task index value based on the ideal parallelism includes: Generate random numbers based on the ideal parallelism; The task index value is determined based on the random number.

4. The method according to claim 3, characterized in that, If the time range is the first half hour, then the target number of minutes is the number of minutes of the event; If the time range is the latter half hour, then the target number of minutes is the difference between the hourly minutes and the event minutes.

5. The method according to claim 4, characterized in that, If the time range is the latter half hour, determining the task index value based on the random number includes: Based on the difference between the number of tasks and the calibration value, the candidate index corresponding to the target subtask is obtained; The task index value is obtained by obtaining the difference between the candidate index and the random number.

6. The method according to claim 1, characterized in that, The step of determining the task index value based on the ideal parallelism includes: The multiple values ​​preceding the ideal parallelism are used as an index set, wherein the N-1 values ​​in the index set are the index values ​​of the Nth task; When new data to be processed is detected and is still normal data, the N+1th task index value is obtained by taking the remainder of the sum of the Nth task index value and the calibration value with respect to the ideal parallelism, until no new data to be processed is detected or the new data to be processed is abnormal data.

7. The method according to claim 1, characterized in that, The step of obtaining the event time field based on the event data to be processed includes: The event time field is extracted from the event data to be processed using the keyword extraction strategy set in the Keyselector operator.

8. A small file processing device, characterized in that, include: The acquisition module is used to acquire the pending event data and the number of tasks during the Flink operation, and to acquire the event time field based on the pending event data. The event time field includes the event occurrence time, and the number of tasks is the number of subtasks under the Flink framework used to process the event data. The first processing module is used to obtain the time difference based on the time of the event occurrence and the current physical time; The second processing module is used to confirm that the current event data to be processed is late data if the time difference value is greater than the late time difference threshold, and to obtain the task index value according to the event time field and the number of tasks. If the time difference is not greater than the late arrival time difference threshold, the current pending event data is confirmed as normal data; extract the event minutes from the event time field; determine the target number of minutes based on the time range to which the event minutes belong; obtain a first ratio between the target number of minutes and a preset number of minutes, and determine a first degree of parallelism based on the product of the first ratio and the number of tasks, wherein the preset number of minutes is determined based on the time partition boundary; obtain the sum of the first degree of parallelism and the number of tasks; Obtain a second ratio between the difference between a preset number of minutes and a target number of minutes and the preset number of minutes, and determine a second degree of parallelism based on the product of the second ratio and half of the sum; obtain an ideal degree of parallelism based on the sum of the first degree of parallelism and the second degree of parallelism; the ideal degree of parallelism is used to indicate the maximum number of target subtasks for processing the event data to be processed under the target number of minutes; determine the task index value based on the ideal degree of parallelism; The execution module is used to obtain the target subtask based on the task index value, process the event data to be processed through the target subtask, and store the processed data in the directory file corresponding to the target subtask.

9. A small file processing device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data writing method, device and equipment and readable storage medium

    CN115221116A

  • Method, device, and program product for managing index of storage system

    US20220237188A1