Distributed data processing delay factor analysis and improvement method, device, and computer program
The method simulates data processing applications in both actual and virtual environments to quantify delay factors, recommending optimal conditions for improved execution efficiency.
Patent Information
- Application Number
- JP2025050735
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2045-03-25
Smart Images

Figure 2025148313000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, device, and computer program for analyzing and improving factors in distributed data processing delays for users, and more specifically to a method, device, and computer program for quantitatively analyzing factors that cause delays in applications that process large amounts of data in a distributed manner and providing improvement solutions for the same. [Background technology]
[0002] Recently, technologies for efficiently processing large amounts of data based on clustering techniques and the like have been rapidly developing.
[0003] For example, various frameworks for large-scale data distributed processing, such as MapReduce and SPARK, are widely used.
[0004] However, for users who run applications based on MapReduce or SPARK, it has been difficult to accurately grasp the degree of data processing delay that occurs during application execution due to various factors such as allocation delays of executors such as containers, and this has made it even more difficult to build a running environment for efficiently running the application.
[0005] As a result, there is a continuing demand for a method that can quantify and accurately grasp the impact of various delay factors that may occur during the execution of a user application that is driven based on a framework for data distributed processing, and thereby provide an operating environment that can efficiently run the application based on this, but an appropriate solution for this has not yet been provided. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Korean Patent Publication No. 10-2014-0080795 (Published July 1, 2014) Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention has been made to solve the problems of the conventional technology described above, and aims to provide a method, device, and computer program for analyzing and improving the causes of delays in distributed data processing, which can quantify and accurately grasp the impact of various delay factors that may occur during the execution of a user application driven based on a framework for distributed data processing.
[0008] More specifically, the present invention aims to provide a method, device, and computer program for analyzing and improving the causes of delays in distributed data processing, which can recommend or provide an operating environment that enables the application to be executed efficiently based on an analysis of an application that is run based on a framework for distributed data processing.
[0009] Other detailed objects of the present invention will be clearly understood and grasped by experts or researchers in the relevant technical field through the specific details described below. [Means for solving the problem]
[0010] A method according to one aspect of the present invention for solving the above problem includes the steps of: calculating, in an analysis device, simulation information for simulating an application based on a log file for the application that performs distributed processing on data; using the simulation information to perform simulations in multiple operating environments including the actual operating environment of the application and a virtual operating environment in which one or more delay factors have been removed; and analyzing the application based on the simulation results in the multiple operating environments.
[0011] Furthermore, a computer program according to another aspect of the present invention may be a computer program for executing the steps of the above-described methods in combination with hardware.
[0012] Another aspect of the present invention provides an apparatus for analyzing an application that performs distributed data processing, the apparatus including a processor and a memory, wherein the memory includes instructions configured to cause the apparatus to perform specific operations when executed by the processor, the specific operations including calculating simulation information for simulating the application based on a log file for the application, using the simulation information to perform simulations in multiple operating environments including an actual operating environment for the application and a virtual operating environment that removes one or more delay factors, and proceeding with analysis of the application based on the simulation results in the multiple operating environments. [Effects of the Invention]
[0013] Therefore, the method, device, and computer program for analyzing and improving delay factors in distributed data processing according to one embodiment of the present invention can quantify and accurately grasp the impact of various delay factors that may occur during the execution of a user application driven based on a framework for distributed data processing.
[0014] In addition, the method, apparatus, and computer program for analyzing and improving the causes of delays in distributed data processing according to one embodiment of the present invention can recommend or provide an operating environment that can efficiently execute an application based on an analysis of an application that is run based on a framework for distributed data processing.
[0015] The effects obtained by the present invention are not limited to those described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present invention pertains from the contents described in this specification. [Brief explanation of the drawings]
[0016] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present invention, provide embodiments of the present invention and, together with the detailed description, explain the technical concept of the present invention. [Figure 1] 1 is a configuration diagram of an application analysis system according to an embodiment of the present invention. [Figure 2] 1 is a flowchart of an application analysis method according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating an application analysis method according to an embodiment of the present invention. [Figure 4] FIG. 1 is a diagram illustrating an application analysis method according to an embodiment of the present invention. [Figure 5] FIG. 1 is a diagram illustrating an application analysis method according to an embodiment of the present invention. [Figure 6] FIG. 1 is a diagram illustrating an application analysis method according to an embodiment of the present invention. [Figure 7] FIG. 10 is a diagram illustrating a specific flowchart of a step of performing a simulation in an application analysis method according to an embodiment of the present invention. [Figure 8] FIG. 2 illustrates an example of the configuration and operation of an analysis device for an application according to an embodiment of the present invention. [Figure 9]10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 10] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 11] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 12] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 13] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 14] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 15] 10 is a flowchart illustrating a specific operation of the application analysis device according to the embodiment of the present invention. [Figure 16a] 10 is a diagram illustrating an example of an analysis result generated by an application analysis method according to an embodiment of the present invention. [Figure 16b] 10 is a diagram illustrating an example of an analysis result generated by an application analysis method according to an embodiment of the present invention. [Figure 16c] 10 is a diagram illustrating an example of an analysis result generated by an application analysis method according to an embodiment of the present invention. [Figure 17] 1 is a block diagram of an apparatus for analyzing an application according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0017] The present invention can be modified in various ways and can have various embodiments, so a specific embodiment will be described in detail below with reference to the accompanying drawings.
[0018] The following embodiments are provided to facilitate a comprehensive understanding of the methods, devices and / or systems described herein, but are for illustrative purposes only and are not intended to limit the present invention.
[0019] When describing embodiments of the present invention, if a detailed description of known technologies related to the present invention is deemed to unnecessarily obscure the gist of the present invention, the detailed description will be omitted. The terms used below are defined in consideration of their functions in the present invention and may vary depending on the intentions or practices of users or operators. Therefore, their definitions should be based on the overall content of this specification. The terms used in the detailed description are intended to describe embodiments of the present invention only and are not limiting. Unless otherwise specified, singular expressions include plural meanings. In this description, terms such as "include" or "comprise" are intended to refer to certain characteristics, numbers, steps, operations, elements, parts thereof, or combinations thereof, and should not be interpreted as excluding the presence or possibility of one or more other characteristics, numbers, steps, operations, elements, parts thereof, or combinations thereof in addition to those described.
[0020] Furthermore, terms such as "first" and "second" may be used to describe various components, but these components are not limited by the above terms, and these terms are used only to distinguish one component from another.
[0021] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Exemplary embodiments of a method, apparatus, and computer program for analyzing and improving causes of delay in distributed data processing according to the present invention will be described in detail below with reference to the accompanying drawings.
[0022] 1 shows a configuration diagram of an application analysis system 100 according to an embodiment of the present invention. As shown in FIG. 1, the application analysis system 100 according to an embodiment of the present invention may include one or more user terminals 110a, 110b, a data distribution processing system 120 that provides data distribution processing for user applications based on a framework such as MapReduce or SPARK, an analysis device 130 that analyzes delay factors in the user applications, and a communication network 140.
[0023] In this case, the terminals 110a and 110b may be various terminal devices such as personal computers (PCs), notebook computers, etc. that can connect to the data distributed processing system 120 or the analysis device 130 via a communication network 140 to run user applications or make requests for analysis work, and various other wired and wireless terminal devices such as tablet PCs 7, smartphones, and PDAs may also be used as the terminals 110a and 110b.
[0024] Furthermore, the data distributed processing system 120 may be realized using one or more servers or devices, or may be realized based on a clustering or cloud-based system, etc., so as to provide data distributed processing for applications requested by users via terminals 110a and 110b, but the present invention is not necessarily limited to this, and may also be realized in various other forms, such as as a dedicated device.
[0025] Furthermore, the analysis device 130 may be realized using one or more servers or devices, or may be realized based on a clustering or cloud-based system, etc., so as to analyze the application requested by the user via the terminal 110a, 110b, thereby calculating the impact of various delay factors, and ultimately recommending an improved operating environment. However, the present invention is not necessarily limited to this, and may be realized in various other forms, such as as a dedicated device.
[0026] Furthermore, in one embodiment of the present invention, the user terminals 110a, 110b, the data distribution processing system 120, and the analysis device 130 do not necessarily have to be realized in a separate form, and they can also be realized in various forms, such as two or more of the terminals 110a, 110b, the data distribution processing system 120, and the analysis device 130 being realized in an integrated form.
[0027] The communication network 140 connecting the user terminals 110a, 110b, the data distributed processing system 120, and the analysis device 130 may include wired and wireless networks, and more specifically, may include various communication networks such as a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN). The communication network 130 may also include the well-known World Wide Web (WWW). However, the communication network 140 according to the present invention is not limited to the networks listed above, and may also include at least a known wireless data network, a known telephone network, or a known wired or wireless television network.
[0028] FIG. 2 also shows a flowchart of an application analysis method according to an embodiment of the present invention.
[0029] 2 may be performed by, for example, an analysis device 130. Furthermore, the analysis device 130 may be implemented by including a computing device such as that shown in FIG. 17 and the description below with respect to FIG. 17. For example, the analysis device 130 may include a processor 10, which may execute instructions configured to perform an analysis of a user's application.
[0030] More specifically, as shown in FIG. 2, an application analysis method according to an embodiment of the present invention may include the steps of: calculating simulation information for simulating an application based on a log file for an application that performs distributed processing on data in an analysis device 130 (S210); performing simulations in multiple operating environments, including an actual operating environment of the application and a virtual operating environment in which one or more delay factors are removed, using the simulation information (S220); and conducting analysis of the application based on the simulation results in the multiple operating environments (S230).
[0031] At this time, in the calculating step (S210), the dependency relationships between one or more work steps of the application can be calculated based on the log file.
[0032] Furthermore, in the calculating step (S210), information on the execution history of one or more unit tasks included in one or more task steps can be calculated.
[0033] In addition, in the calculation step (S210), one or more pieces of information can be calculated, including the number of unit tasks included in one or more task steps, the time taken to perform each unit task, and the causes of delay that occurred in each unit task.
[0034] Furthermore, in the calculating step (S210), information on one or more executors to which one or more unit tasks are assigned can be calculated.
[0035] Furthermore, in the calculating step (S210), it is possible to calculate information on one or more of the allocated times for one or more executors and the causes of delays that have occurred in one or more executors.
[0036] In addition, the step of performing a simulation (S220) may include a step of performing a first simulation of the application in an actual driving environment (S710), and a step of performing a second simulation of the application in a virtual driving environment in which one or more of the multiple delay factors in the actual driving environment have been removed (S720).
[0037] In this case, in the step of performing the second simulation (S720), the second simulation can be performed for the application in multiple virtual operating environments that are configured by combining and removing one or more of multiple delay factors from the actual operating environment.
[0038] In addition, in the proceeding step (S230), the influence of the one or more delay factors on the application can be calculated based on the first execution completion time of the application calculated through simulation in the actual operating environment and the second execution completion time of the application calculated through simulation in the virtual operating environment.
[0039] At this time, in the proceeding step (S230), the degree of influence of each of the multiple delay factors on the application can be calculated based on the difference between the multiple second execution completion times and the first execution completion time corresponding to each of the multiple delay factors.
[0040] In addition, in the proceeding step (S230), the combined impact of two or more delay factors on the application can be calculated based on the second execution completion time in a virtual operating environment to which two or more of the multiple delay factors are applied together.
[0041] In addition, in the step of performing a simulation (S220), a simulation is performed while changing the setting values for the application, and in the proceeding step (S230), optimal setting values for the application can be calculated based on the analysis of the application.
[0042] Therefore, the method, device, and computer program for analyzing a user's application according to one embodiment of the present invention can quantify and accurately grasp the impact of various delay factors that may occur during the execution of a user's application that is driven based on a framework for data distributed processing, and can recommend or provide an operating environment that can efficiently execute the application based on an analysis of the application that is driven based on a framework for data distributed processing.
[0043] Hereinafter, each step of the application analysis method according to an embodiment of the present invention will be described in detail with reference to FIGS.
[0044] First, in step S210, the analysis device 130 calculates simulation information for simulating the application based on a log file for the application that performs distributed processing on data.
[0045] Here, the application can be run based on various frameworks for distributed data processing, such as SPARK, MapReduce, Tez, and Trino, but the present invention is not necessarily limited to these.
[0046] In addition, the log file may include all types of log files generated during the execution of an application, such as log files stored in storage, but may also include various types of log files stored in memory, remote devices, etc.
[0047] This allows the analysis device to calculate log information for analyzing the application based on the log file for the application.
[0048] In addition, in a distributed data processing system aimed at large-volume processing, the entire work can be divided into stages that must be performed sequentially, taking into account dependencies between tasks, and the stages can be further divided into unit tasks. Each stage can have dependent steps that must be performed before it, and the stage cannot be performed until the dependent steps are completed.
[0049] More specifically, as shown in Figure 3, an application run by a data distribution processing system 310 can include multiple work steps (e.g., Stage 0 (315a), Stage 1 (315b), Stage 2 (315c), and Stage 3 (315d) in Figure 3), and each work step can include one or more unit tasks (Tasks). In Figure 3, the dependent work step of Stage 2 is that Stage 0 and Stage 1 must be completed in order for Stage 0, Stage 1, and Stage 2 to be performed.
[0050] A scheduler 311 of a data distribution processing system 310 references the dependency relationships of task steps according to an application execution plan 312, and assigns unit tasks 317a, 317b of each task step to executors 316a, 316n for execution.
[0051] In this case, the Executor can be driven by a container. The scheduler 311 starts work from executable work steps, and when all unit work of a work step is completed, the work step is completed.
[0052] However, in the data distributed processing system 410, since multiple unit tasks are distributed among multiple executors, some executors or unit tasks may fail. In this case, the system can identify the parts that need to be re-executed and re-execute those parts. For example, if a specific executor fails, the unit work being executed by that executor at the time of the failure may fail. In addition, intermediate results (shuffles) of work steps previously completed by the failed executor may be lost, and unit work that uses the lost intermediate results as inputs will also fail.
[0053] To give a more specific example, in FIG. 4, the data distribution processing system 410 assigns a unit task (Task) of Stage 0 (413a) to each executor (Executor) 416a, 416b, etc., and then performs unit tasks (Tasks) of Stage 1 (413b) and Stage 2 (413c), and then assigns a unit task (Task) of Stage 3 (413d) to each executor (Executor) 416a, 416b, etc., and performs them.
[0054] In this case, in the example of Figure 4, if the third executor (Executor#3) 416b fails during the execution of Stage 3, the unit task 417b performed by that executor must be re-executed, and as a result, the results (shuffle) of the previous task steps are also lost, and the unit tasks already performed (e.g., 415a, 416b, 415c in Figure 4) can also be included in the unit tasks that need to be re-executed. In addition, even if they were not performed by the failed executor, other unit tasks 417a and 417c that used the lost intermediate results as inputs can also be included in the unit tasks that need to be re-executed.
[0055] Referring to Figure 4, work steps that use the intermediate results (Shuffle) of other work steps can be considered to belong to the same Job. The final work step of each Job can derive permanent results (for example, storing them in a shared file system and outputting the results to the user) while completing the work.
[0056] As a result, in this invention, a simulator is implemented that reflects the scheduler and re-execution logic of each data distribution processing framework such as SPARK and MapReduce, and simulation information calculated from the log file of the user's application is applied to simulate the application, and based on this, the impact of various delay factors that may occur during the execution of the application can be quantified and accurately understood, and based on this, an operating environment that can efficiently execute the application can be recommended or provided.
[0057] More specifically, in the calculation step (S210), the analysis device can calculate information (=simulation information) necessary for simulating the application based on the log file generated through the operation of the application, or can receive the information from the user. However, the present invention is not necessarily limited to this, and simulation information can be collected in various other ways, such as by being sent from an external server, etc.
[0058] In this case, the simulation information may include one or more pieces of information such as information on dependencies between work steps, information on the execution history of unit work, such as the number of unit work steps included in the work step, the time taken to execute each unit work, and delay factors that occurred in each unit work, and information on the executor, such as the allocated time for the executor and delay factors that occurred in the executor.
[0059] As a more specific example, FIG. 5(a) illustrates simulation information calculated from the log file of an application. In this case, the simulation information may include the configuration of each stage and task, the time taken to execute each task, the configuration and allocation time of each executor, the time when a delay occurred due to a failure, and the cause of the delay (for example, in FIG. 5(a), a delay occurred 300 seconds after the third executor (Executor#3) started execution due to a failure caused by out-of-memory (OOM)).
[0060] FIG. 5(b) shows an example of reconstructing the entire task execution process, including the re-execution process after the failure of the third executor (Executor#3), for the case of FIG. 5(a) based on the log file.
[0061] Next, in step S220, the calculated simulation information is used to perform simulations in a plurality of operating environments, including the actual operating environment of the application and a virtual operating environment in which one or more delay factors have been removed.
[0062] In this case, step S220 may include a step (S710) of performing a first simulation of the application in an actual operating environment, and a step (S720) of performing a second simulation of the application in a virtual operating environment in which one or more of the multiple delay factors in the actual operating environment have been removed, as shown in FIG. 7.
[0063] Furthermore, in step S720, a second simulation of the application can be performed in a plurality of virtual operating environments configured by combining and removing one or more of the plurality of delay factors from the actual operating environment.
[0064] More specifically, in the present invention, delay factors that delay the execution of an application may include executor allocation delay, execution delay (skewness) of a specific unit task (Task) of an execution step, termination of an executor due to resource preemption, termination of an executor due to out-of-memory, etc. However, the present invention is not limited to these, and some delay factors may be added or removed depending on the application field, and even multiple delay factors may act in combination to cause delay.
[0065] In this way, in step S220, simulations are performed in various operating environments while applying or excluding these delay factors individually or in combination.
[0066] More specifically, as shown in FIG. 6, first, a simulation can be performed in the actual operating environment captured in the application log file (FIG. 6(a)).
[0067] Next, it is possible to perform a simulation in a virtual driving environment in which each delay factor is eliminated one by one, or in which a plurality of delay factors are further eliminated ((b) to (e) of FIG. 6).
[0068] In this case, as shown in Figure 6, when compared with the time required to execute an application in the actual operating environment (T0 in Figure 6), in the virtual operating environment where one or more delay factors are removed, a simulation result can be calculated in which the execution time is shortened according to each delay factor (T1 to T4 in Figure 6).
[0069] Accordingly, in step S230, analysis of the application is carried out based on the simulation results in a plurality of operating environments.
[0070] Here, in step S230, the influence of the one or more delay factors on the application can be calculated based on the first execution completion time of the application calculated through simulation in the actual operating environment and the second execution completion time of the application calculated through simulation in the virtual operating environment.
[0071] At this time, in step S230, the degree of influence of each of the multiple delay factors on the application can be calculated based on the difference between the multiple second execution completion times and the first execution completion time due to each of the multiple delay factors. For example, the difference between the first execution completion time, which is the result of a simulation when the delay factor is included as is, and the second execution completion time, which is the result of a simulation when a specific delay factor is excluded, can be calculated as the degree of influence of the delay factor.
[0072] More specifically, referring to FIG. 6, the impact of each delay factor on an application can be calculated based on the first execution completion time T0 calculated through a simulation in the actual operating environment and the second execution completion time calculated through a simulation in the virtual operating environment.
[0073] In this case, the delay due to Executor Allocation delay can be T1, the delay due to skewness of a specific task in the execution step can be T2, and the delay due to resource preemption of the Executor can be T3. Based on this, the degree of contribution (%) of each delay factor can be calculated and the impact of each delay factor can be calculated.
[0074] Furthermore, in step S230, the combined impact of two or more delay factors on an application can be calculated based on the second execution completion time in a virtual operating environment in which two or more of the multiple delay factors are applied together. For example, if the difference between the first execution completion time, which is the result of a simulation when the delay factors are included as they are, and the second execution completion time, which is the result of a simulation when all delay factors are excluded, is greater than the sum of the impacts when the delay factors are each excluded, the difference can be calculated as the combined impact of the delay factors.
[0075] More specifically, in a situation where only allocation delay factors and execution delay factors exist, if the first execution completion time calculated through a simulation in an actual operating environment including all delay factors is 60 seconds, the second execution completion time in a virtual operating environment excluding executor allocation delay is 55 seconds, the second execution completion time in a virtual operating environment excluding a specific unit task execution delay (skewness) is 50 seconds, and the second execution completion time in a virtual operating environment excluding all delay factors (allocation, skewness) is 30 seconds, the impact of allocation delay is calculated to be approximately 8% (60 seconds - 55 seconds = 5 seconds), the impact of unit task execution delay (skewness) is calculated to be approximately 16% (60 seconds - 50 seconds = 10 seconds), and the combined impact of allocation delay and execution delay factors can be calculated to be approximately 25% ((60 seconds - 30 seconds) - 5 seconds - 10 seconds = 15 seconds).
[0076] In addition, the present invention can quantitatively analyze and provide the impact of each delay factor in various aspects such as the time it takes to execute an application and resource usage through simulations in various operating environments, and can also provide guidance on how to deal with major delay factors.
[0077] Furthermore, in the present invention, a simulation is performed while changing the setting values for the application in step S220, and optimal setting values for the application can be calculated and provided based on an analysis of the application in step S230.
[0078] As a more specific example, the present invention can perform simulations while changing various delay factors and application settings (e.g., the number of concurrent executors) to calculate or provide settings that are optimal for the user's situation. For example, through analysis of an application performed with 40 executors, it can propose a solution of changing the number of executors to 80 if shortening execution time is important, to 25 if reducing resource usage is important, or to 55 if efficiency of execution time is more important than resource usage.
[0079] FIG. 8 also illustrates the configuration and operation of an analysis device for an application according to an embodiment of the present invention.
[0080] As shown in FIG. 8, the analysis device 830 for the application 810 may include an application log file acquisition and status confirmation module 831, a simulation information extraction module 832, multiple simulation modules 833a, 833b, 833c, 833d, 833n, a result compilation module 834, an improvement method derivation module 835, and a report generation module 836.
[0081] At this time, when the application log file acquisition and status confirmation module 831 receives an analysis request for an application from the user 840, it can collect log files for the application to be analyzed from the work log repository 820, etc., and confirm whether the work was completed normally and was successful.
[0082] More specifically, referring to FIG. 9, the application log file acquisition and status confirmation module 831 first attempts to acquire a log file for the application to be analyzed from the work log repository 820 or the like (S910).
[0083] It is checked whether a log file can be secured (S920), and if a log file can be secured, it can be checked whether the application corresponds to the work of the user who requested the analysis for security reasons (S930).
[0084] In addition, the application log file securing and status checking module 831 checks whether the application to be analyzed has not abnormally terminated and is a task that has been successfully completed (S940).
[0085] This is to compare and analyze the simulation results for various virtual driving environments based on the actual driving environment of the work successfully completed in the present invention.
[0086] As a result, in the application log file allocation and status confirmation module 831, if all of the conditions of the above three steps (S920, S930, S940) are met, the module proceeds to analyze the application (S960), and if even one of the conditions is not met, the module terminates without performing the analysis (S950).
[0087] Next, as shown in FIG. 10, the simulation information extraction module 832 first extracts simulation information for executing a simulation of the application from the secured log file (S1010).
[0088] Next, the simulation information extraction module 832 extracts a topology for dependencies between work steps for the application to be analyzed (S1020).
[0089] More specifically, the simulation information extraction module 832 calculates the submission and completion time or order of each work step from the log file, and based on this, can extract a topology such as the dependency relationships between each work step.
[0090] As a more specific example, consider the case of extracting a topology for an application to be analyzed based on the following log file:
[0091] <Example of a log file> Get started 10:00 Stage 0 submission 10:05 Stage done 10:06 Stage 1 submitted 10:07 Stage 2 submitted 10:08 Stage 2 done 10:10 Stage 1 done 10:11 Stage 3 submission 10:15 Stage 3 done Work completed At this point, it can be seen that the work is composed of four work steps, Stage 0 to Stage 3. Since there are no work steps preceding Stage 0 at the start, it can be seen that there are no work steps that must be completed before Stage 0. Since the submission times of Stage 1 and Stage 2 are after the completion of Stage 0, it can be judged that these are work steps that will be carried out after the completion of Stage 0. Since the submission time of Stage 3 is after the completion of Stages 1 and 2, it can be judged that this is a work step that will be carried out after the completion of Stages 1 and 2.
[0092] As a result, as shown in FIG. 11, a topology including the dependency relationships between Stage 0 (1110), Stage 1 (1120), Stage 2 (1130), and Stage 3 (1140) can be extracted from the log file.
[0093] Furthermore, the simulation information extraction module 832 extracts information about the unit tasks (Tasks), such as the number of unit tasks, the execution time of each task, whether or not there was a failure, and the cause of the failure (S1030).
[0094] Furthermore, the simulation information extraction module 832 extracts information about the executors, such as the number of executors assigned, the required time for assignment, whether or not there was a failure, and the cause of the failure (S1040).
[0095] As a result, the simulation information extraction module 832 proceeds to analyze the application based on the simulation information calculated based on the log file (S1050).
[0096] Next, as shown in FIG. 12, a plurality of simulation modules 833a, 833b, 833c, 833d, and 833n perform simulations for the actual driving environment and each virtual driving environment.
[0097] More specifically, first, the simulation module performs a simulation based on the calculated simulation information (S1210), while recording information about the progress of the simulation, such as the execution time and the amount of resources used (S1220).
[0098] Next, the simulation module determines whether the ongoing unit task (Task) has succeeded or failed (S1230), and if a success or failure has occurred, it performs a sub-processing routine for the success or failure of the unit task (Task) as follows.
[0099] First, in the sub-processing routine for the success or failure of a unit task (Task), it is determined whether the unit task (Task) is a success or a failure (S1231).
[0100] At this time, if it is determined to be successful, the information regarding the success of the unit task is updated (S1232).
[0101] On the other hand, if it is determined to be a failure, it is determined whether the failure was caused by a delay factor to be excluded from the simulation (S1233). If it is not a delay factor to be excluded, the impact of the failure on the unit task is grasped and whether or not re-execution is necessary is updated (S1234).
[0102] In addition, the simulation module determines whether an executor has been added or failed (S1240), and if an addition or failure has occurred, performs a sub-processing routine for the addition or failure of an executor as follows.
[0103] First, in the sub-processing routine for adding or failing an executor, it is determined whether an executor has been added or failed (S1241).
[0104] At this time, if it is determined that the executor has been added, the information regarding the addition of the executor is updated (S1242).
[0105] On the other hand, if it is determined to be a failure, it is determined whether the failure was caused by a delay factor that is to be excluded from the simulation (S1243). If the delay factor is not one that is to be excluded, the impact of the failure on the Executor is grasped and whether or not re-execution is necessary is updated (S1244).
[0106] The simulation module determines whether additional scheduling is possible due to the updated information and proceeds with the scheduling (S1250).
[0107] Next, the simulation module determines whether all tasks have been completed (S1260), and ends the simulation (S1270), or if not, returns to step S1220 and repeats the processing steps.
[0108] Next, a result summarizing module 834 summarises the results of the simulations carried out by the multiple simulation modules 833a, 833b, 833c, 833d, and 833n, and based on this, analyzes the degree of influence of each delay factor, etc.
[0109] More specifically, as shown in FIG. 13, the result summarizing module 834 first obtains the time taken for execution, the amount of resources used, and the like from the results of each simulation (S1310).
[0110] Next, the result summarization module 834 compares the results of the simulation in the virtual driving environment, in which each delay factor is excluded, with the results of the simulation in the actual driving environment, in which all delays are included, to determine how much the required time has been reduced (S1320).
[0111] Next, the result summarizing module 834 determines the degree of influence of each delay factor based on the reduction in the required time in each simulation (S1330).
[0112] In addition, the result summarization module 834 calculates how much the required time reduction in the simulation in the virtual driving environment excluding all delay factors compares with the simulation result when all delay factors are included (S1340), determines whether the reduction in the simulation excluding all delay factors is greater than the total required time reduction in the simulation in the virtual driving environment excluding each delay factor (S1350), and based on this, determines the degree of combined impact of multiple delay factors (S1370) or determines that there is no combined impact (S1360), and completes the result summarization and impact analysis (S1380).
[0113] Next, the improvement method derivation module 835 derives an improvement plan for the application based on the influence of each delay factor calculated through the simulation.
[0114] More specifically, as shown in FIG. 14, the improvement method derivation module 835 first checks the degree of influence of each delay factor (S1410).
[0115] Next, the improvement method derivation module 835 determines whether the influence of a specific delay factor is equal to or greater than a predetermined reference value (S1420).
[0116] At this time, if the degree of influence of a specific delay factor is equal to or greater than a reference value, the improvement method deriving module 835 can provide the user with a guide for improving the delay factor (S1430).
[0117] Next, the improvement method derivation module 835 performs various simulations while changing setting values such as the number of executors (S1440), and based on the results, determines whether there is an improvement compared to the existing system for each objective such as the time taken for execution, the amount of resources required, and efficiency (S1450).If there is an improvement, it conveys a guide for the setting values along with the expected improvement effect (S1460), and generates a report for the user (S1470).
[0118] As a result, as shown in FIG. 15, a user can transmit an analysis request 1510 for his / her application to an analysis device 1520 to have the application analyzed.
[0119] In this case, the user's analysis request 1510 may include user identification information such as a user ID that can identify the user, application identification information such as an application ID that can identify the user's application, the type of framework in which the application is run (e.g., MapReduce, SPARK, etc.), and may even be requested to include information that is difficult to calculate from the log file of the application to be analyzed (e.g., a set value for the number of executors that can actually be used in SPARK, a set value for the number of mappers / reducers that can be executed simultaneously in MapReduce, etc.).
[0120] In this case, as shown in FIG. 15, the analysis device 1520 can be realized to include one or more application analysis (App.Analysis) pods that perform analysis of distributed processing applications according to the present invention based on Kubernetes, but this is merely one example and the present invention is not necessarily limited to this.
[0121] This allows the analysis device 1520 to record information about the user request 1510 in a database 1530 or the like, and store the allocation state and the like.
[0122] Next, when the analysis device 1520 completes the analysis of the application to be analyzed, an analysis report 1540 can be provided to the user via email or the like.
[0123] At this time, the analysis report 1540 may include the analysis results of the degree of influence of each delay factor in terms of the time required to execute the application, as shown in FIG. 16a.
[0124] The analysis report may also include the analysis results on the impact of each delay factor in terms of the amount of resources used to execute the application, as shown in FIG. 16b.
[0125] Furthermore, the analysis report can calculate and recommend optimal settings for the application based on objectives such as the time taken to execute the application, the amount of resources required, and efficiency, as shown in FIG. 16c.
[0126] Furthermore, a computer program according to one aspect of the present invention is a computer program for executing each step of the above-described application analysis method in combination with hardware, and may be continuously stored in a computer-readable medium or temporarily stored for execution or download. The computer program may be a computer program including machine language code created by a compiler, or a computer program including high-level language code that can be executed on a computer using an interpreter, etc. In this case, the computer is not limited to a personal computer (PC) or a laptop, but may also include integrated information processing devices equipped with a central processing unit (CPU) and capable of executing a computer program, such as a server, smartphone, tablet PC, PDA, or mobile phone. The computer-readable medium may also include integrated computer-readable storage media such as electronic storage media (e.g., ROM, flash memory, etc.), magnetic storage media (e.g., floppy disks, hard disks, etc.), and optically readable media (e.g., CD-ROM, DVD, etc.).
[0127] FIG. 17 also illustrates the configuration and operation of an apparatus 50 for analyzing an application according to an embodiment of the present invention.
[0128] Referring to FIG. 17, said device 50 can be an analysis device for analyzing an application according to the proposed method of the present invention.
[0129] For example, the device 50 to which the proposed method of the present invention can be applied may include network devices such as repeaters, hubs, bridges, switches, routers, and gateways, computer devices such as desktop computers and workstations, mobile terminals such as smartphones, portable devices such as laptop computers, home appliances such as digital TVs, and means of transportation such as automobiles. As another example, the device 200 to which the present invention can be applied may be included as part of an ASIC (Application Specific Integrated Circuit) realized in the form of an SoC (System On Chip).
[0130] The memory 20 may be connected to the processor 10 during operation, and may store programs and / or commands for processing and controlling the processor 10, as well as data and information used in the present invention, control information necessary for data and information processing according to the present invention, temporary data generated during data and information processing, etc. The memory 20 may be implemented as a storage device such as a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a static RAM (SRAM), a hard disk drive (HDD), or a solid state drive (SSD).
[0131] The processor 10 may be operatively connected to the memory 20 and / or the network interface 30 and controls the operation of each module within the device 50. In particular, the processor 10 may perform various control functions for implementing the proposed method of the present invention. The processor 10 may also be referred to as a controller, microcontroller, microprocessor, microcomputer, etc. The proposed method of the present invention may be implemented by hardware, firmware, software, or a combination thereof. When implementing the present invention using hardware, the processor 10 may include an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), etc., configured to implement the present invention. On the other hand, when the proposed method of the present invention is implemented using firmware or software, the firmware or software may include instructions relating to modules, procedures, functions, etc. that perform the functions or operations necessary to implement the proposed method of the present invention, and the instructions may be stored in memory 20 or in a computer-readable recording medium (not shown) separate from memory 20 and, when executed by processor 10, may configure device 50 to implement the proposed method of the present invention.
[0132] The apparatus 50 may also include a network interface device 30. The network interface device 30 is operatively connected to the processor 10, and the processor 10 controls the network interface device 30 to transmit or receive wireless / wired signals carrying information and / or data, signals, messages, etc. over a wireless / wired network. The network interface device 30 supports various communication standards, such as IEEE 802 series, 3GPP LTE(-A), and 3GPP 5G, and can transmit and receive control information and / or data signals in accordance with the communication standards. The network interface device 30 may be implemented as the apparatus 50 alone, if necessary.
[0133] Therefore, the method, device, and computer program for analyzing and improving delay factors in distributed data processing according to one embodiment of the present invention can quantify and accurately grasp the impact of various delay factors that may occur during the execution of a user application driven based on a framework for distributed data processing.
[0134] In addition, the method, apparatus, and computer program for analyzing and improving the causes of delays in distributed data processing according to one embodiment of the present invention can recommend or provide an operating environment that can efficiently execute an application based on an analysis of an application that is run based on a framework for distributed data processing.
[0135] The above description merely exemplifies the technical concept of the present invention, and those skilled in the art can make various modifications and variations without departing from the essential characteristics of the present invention. Therefore, the embodiments described in the present invention are for illustrative purposes only and are not intended to limit the technical concept of the present invention. The scope of protection of the present invention should be interpreted by the following claims, and all technical concepts within the scope equivalent thereto should be interpreted as being within the scope of the present invention. [Explanation of symbols]
[0136] 10 processors 20 memory 30 Network Interfaces 50 equipment 100 Application Analysis Systems 110, 110a, 110b terminals 120 Data Distributed Processing System 130 Analyzer 140 Communication Network
Claims
1. calculating, in the analysis device, simulation information for simulating an application that performs distributed processing on data based on a log file for the application; using the simulation information to perform simulations in a plurality of operating environments, including an actual operating environment for the application and a virtual operating environment in which one or more delay factors have been removed; and progressing analysis of the application based on simulation results in the plurality of operating environments.
2. In the calculating step, The method of claim 1 , further comprising calculating information regarding dependencies between one or more work steps of the application based on the log files.
3. In the calculating step, The method according to claim 2 , wherein information about an execution history of one or more unit operations included in the one or more operation steps is calculated.
4. In the calculating step, The method according to claim 3, further comprising calculating one or more pieces of information among the number of unit tasks included in the one or more task steps, the time taken to perform each unit task, and delay factors occurring in each unit task.
5. In the calculating step, The method of claim 3 , further comprising calculating information about one or more performers to whom the one or more unit tasks are assigned.
6. In the calculating step, The method according to claim 5 , further comprising calculating information on one or more of the allocated times for the one or more performers and delay factors caused by the one or more performers.
7. The performing step is performing a first simulation of the application in the actual operating environment; and performing a second simulation of the application in the virtual driving environment in which one or more of a plurality of delay factors relative to the actual driving environment have been removed.
8. In the step of performing the second simulation, The method according to claim 7 , further comprising: performing a second simulation of the application in a plurality of virtual driving environments configured by combining and removing one or more of a plurality of delay factors from the actual driving environment.
9. In the proceeding step, a first execution completion time of the application calculated through a simulation in the actual operating environment; based on a second execution completion time of the application calculated through a simulation in the virtual operating environment; The method of claim 1 , further comprising calculating an impact of the one or more delay factors on the application.
10. In the proceeding step, The method of claim 9 , further comprising calculating an impact of each of the plurality of delay factors on the application based on a difference between a plurality of second execution completion times caused by each of the plurality of delay factors and the first execution completion time.
11. In the proceeding step, The method of claim 10, further comprising calculating a combined impact of the two or more delay factors on the application based on a second execution completion time in a virtual operating environment to which two or more of the plurality of delay factors are applied together.
12. In the step of performing A simulation is performed while changing the setting values for the application. In the proceeding step, The method of claim 1 , wherein optimal settings for the application are calculated based on an analysis of the application.
13. A computer program for performing the steps of the method according to any one of claims 1 to 12 in combination with hardware.
14. An apparatus for performing analysis on an application that performs distributed processing of data, the apparatus including a processor and a memory, The memory includes instructions that, when executed by the processor, cause the device to perform specific operations, the specific operations including: calculating simulation information for simulating the application based on a log file for the application; using the simulation information to perform simulations in a plurality of operating environments, including an actual operating environment for the application and a virtual operating environment in which one or more delay factors have been removed; and progressing analysis of the application based on simulation results in the plurality of operating environments.
15. By the calculation, The apparatus of claim 14 , further comprising: a processor configured to calculate information about dependencies between one or more work steps of the application based on the log file.
16. The above-mentioned steps include: conducting a first simulation of the application in the actual operating environment; and The apparatus of claim 14 further comprising performing a second simulation for the application in the virtual driving environment in which one or more of a plurality of delay factors relative to the actual driving environment are removed.
17. In the above-mentioned progression, a first execution completion time of the application calculated through a simulation in the actual operating environment; based on a second execution completion time of the application calculated through a simulation in the virtual operating environment; 17. The apparatus of claim 14, 15 or 16, configured to calculate an impact of the one or more delay causes on the application.
18. In the above-mentioned steps, A simulation is performed while changing the setting values for the application. In the above-mentioned progression, The apparatus of claim 14 , further comprising: an apparatus for calculating optimal settings for the application based on an analysis of the application.
Citation Information
Patent Citations
Policy simulator for autonomous management system
JP2005196601A
Trace based method for the analysis, benchmarking and tuning of object oriented databases and applications
US6145121A
Load balancing method and load balancing system for hadoop mapreduce under virtual environment
KR1020140080795A