Execution time prediction device, execution time prediction method, and program

WO2026203427A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/026700
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2025-07-28
Publication Date
2026-10-01

Smart Images

  • Figure JP2025026700_01102026_PF_FP_ABST
    Figure JP2025026700_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A training data generation unit (11) generates training data that constitute sets of the processing type, the size of an argument passed to the processing, and the execution time for when executing the processing on the basis of the argument of said size. A machine learning model selection function unit (121) selects a machine learning model from among a plurality of machine learning model candidates on the basis of variation tendencies in the execution times associated with the processing types and argument sizes included in the training data generated by the training data generation unit. A machine learning model training function unit (122) uses the training data generated by the training data generation unit to train the machine learning model selected by the machine learning model selection function unit and construct an execution time prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Execution time prediction apparatus, execution time prediction method, and program

[0001] The present invention relates to an execution time prediction apparatus, an execution time prediction method, and a program.

[0002] In so-called cloud services, in order to quickly and efficiently respond to execution requests for application programs, it is required to select optimal computing resources. Note that a cloud service is a system that enables use of services such as data and programs via the Internet.

[0003] Here, predicting the execution time of a program with high accuracy is important for ensuring quality of service (QoS) and reducing waste of computing resources.

[0004] In the technology described in Non-Patent Document 1, execution time is predicted by constructing a dedicated execution time prediction model for each program. When using the technology described in Non-Patent Document 1, there is a problem that in order to support a new program, the execution time cannot be predicted unless a model for that program is reconstructed. To solve this problem, as one highly versatile execution time prediction model, a method of analyzing a program at the source code level to predict execution time has been studied.

[0005] In the technology described in Non-Patent Document 2, the structure and processing content of a program are specified by analyzing the source code of the program. Then, based on the number of occurrences of processing during execution and the total size of generated data, a single model for predicting the execution time of the entire program is constructed. Once the model is constructed, the execution time of a program can be predicted using the model. The technology described in Non-Patent Document 2 does not construct a dedicated model for a specific program, and has high versatility. In other words, with the technology described in Non-Patent Document 2, it is possible to predict the execution time of a newly introduced program using a single model.

[0006] Chao Wu, Shingo Horiuchi, Kenichi Tayama, “A Resource Design Framework to Realize Intent-based Cloud Management”, 2019 IEEE International Conference on Cloud Computing Technology and Science, p.37-44, 2019. Weihua Liu, Erh-Wen Hu, Bogong Su, Jian Wang, “Using machine learning techniques for DSP software performance prediction at source code level”, Connection Science, 2021, Vol.33, No.1, pp26-41.

[0007] The conventional technologies described above also have their drawbacks. In the technology described in Non-Patent Literature 2, the prediction of execution time is based on the frequency of processing and the amount of data being processed. The prediction of execution time in the conventional technology does not take into account the specific content of the processing, the data type of the arguments, or the data size of the arguments. Therefore, the conventional technology has the problem of estimating execution time without distinguishing between differences in specific processing.

[0008] For example, conventional techniques predict execution time without distinguishing between simple arithmetic operations (e.g., addition) and computationally intensive operations (e.g., sorting within a loop). Furthermore, even for operations represented by the operator "+", the processing content can differ depending on the data type of the arguments being processed. The operator "+" represents arithmetic addition when the argument is an integer, and string concatenation when the argument is a string. In other words, operations represented by the same symbol can have different processing content. Because of these differences, prediction errors based on processing content can occur at the processing unit level. If these prediction errors accumulate, they can lead to large errors.

[0009] For example, if a 0.1 microsecond error occurs in a process that takes several microseconds (μs, 1 μs = 10^(-6) seconds), then 1 million calculations will accumulate an error of 100 milliseconds (m seconds, 1 m second = 10^(-3) seconds). Errors of this magnitude can impact the selection and scaling of cloud resources.

[0010] This invention has been made in consideration of the above circumstances, and aims to provide an execution time prediction device, execution time prediction method, and program that can select a type of machine learning model according to the content (type) of processing and the data size of the arguments, train the selected machine learning model, and construct a machine learning model that can accurately predict execution time.

[0011] [1] To solve the above problems, an execution time prediction device according to one aspect of the present invention includes: a training data generation unit that generates training data which is a set of a type of process, the size of an argument passed to the process, and the execution time when the process is executed based on the size of the argument; a machine learning model selection function unit that selects a machine learning model from a plurality of candidate machine learning models based on the trend of variation in the execution time corresponding to the type of process and the size of the argument in the training data generated by the training data generation unit; and a machine learning model training function unit that constructs an execution time prediction model by training the machine learning model selected by the machine learning model selection function unit using the training data generated by the training data generation unit.

[0012] [2] In another embodiment, the execution time prediction device of [1] further comprises an execution time prediction unit that inputs the type of processing included in a given program and the size of the arguments passed to the processing to an execution time prediction model constructed by the machine learning model training function unit, thereby predicting the execution time of the processing corresponding to the type of processing and the size of the arguments.

[0013] [3] In another embodiment, the execution time prediction device described in [2] above includes a source code analysis unit that analyzes the source code of the program and outputs information on the type of processing included in the program and the size of the arguments passed to the processing, and inputs the type of processing and the size of the arguments passed to the processing output by the source code analysis unit into the execution time prediction model.

[0014] [4] In another embodiment, in the execution time prediction device described in [3] above, the source code analysis unit further outputs information on the data attributes of the arguments in the source code by analyzing the source code, and the execution time prediction unit corrects the execution time of the process predicted by inputting the type of process and the size of the arguments passed to the process into the execution time prediction model, based on the data attributes output by the source code analysis unit.

[0015] [5] In another embodiment, in the execution time prediction device described in [4] above, the data attribute is at least an attribute that indicates whether the argument is a constant or a variable.

[0016] [6] In another embodiment, the execution time prediction device described in [5] above, the data attribute is an attribute that, when the argument is a variable, indicates whether the variable is a local scope variable that is valid only within a predetermined scope of the source code, or a global scope variable that is valid outside the predetermined scope.

[0017] [7] In another embodiment, the execution time prediction device of any of the above [4] to [6] is further comprising a correction value calculation unit which performs the process using sample data with modified data attributes different from the data attributes of the training data used by the machine learning model training function unit when training the machine learning model, measures the execution time thereof, determines the difference between the execution time when the process is performed with the original training data and when the process is performed with the sample data with modified data attributes, and calculates a correction value based on the difference, and the execution time prediction unit corrects the execution time of the process using the correction value calculated by the correction value calculation unit.

[0018] [8] In another embodiment, in the execution time prediction device described in [7] above, the correction value calculation unit receives information on the data attributes of the arguments output by the source code analysis unit and calculates the correction value based on said data attributes.

[0019] [9] Another embodiment is an execution time prediction method that includes the process of a training data generation unit generating training data which is a set of a type of process, the size of an argument passed to the process, and the execution time when the process is executed based on the size of the argument; the process of a machine learning model selection function unit selecting a machine learning model from a plurality of candidate machine learning models based on the trend of variation in the execution time corresponding to the type of process and the size of the argument in the training data generated by the training data generation unit; and the process of a machine learning model training function unit constructing an execution time prediction model by training the machine learning model selected by the machine learning model selection function unit using the training data generated by the training data generation unit.

[0020]

[10] Another embodiment is a program for causing a computer to execute the execution time prediction method described in [9] above.

[0021] According to the present invention, a machine learning model selection function unit can select a machine learning model suitable for the training data based on the training data generated by the training data generation unit, and a machine learning model training function unit can train the selected machine learning model.

[0022] This is a block diagram showing the schematic functional configuration of the execution time prediction device according to the first embodiment. This is a schematic diagram showing examples of multiple types of machine learning models that can be selected by the machine learning model selection function unit according to the first embodiment. This is an example of a graph showing the relationship between parameter values ​​(size of processing arguments) and the execution time of a predetermined process in the first embodiment. This is another example of a graph showing the relationship between parameter values ​​(size of processing arguments) and the execution time of a predetermined process in the first embodiment. This is a schematic diagram showing an example of data representing the relationship between parameter size and measured execution time in the first embodiment. This is another schematic diagram showing another example of data representing the relationship between parameter size and measured execution time in the first embodiment. This is a schematic diagram showing the result of the machine learning model selection function unit selecting a machine learning model in the first embodiment. This is a schematic diagram showing an example of source code stored in the source code storage unit in the first embodiment. This is a schematic diagram showing an example of the result of source code analysis by the source code analysis unit of the execution time prediction unit in the first embodiment. This is a block diagram showing an example of the internal configuration of the execution time prediction device (computer) according to the first embodiment. This is a block diagram showing the schematic functional configuration of the execution time prediction device according to the second embodiment. This graph shows the relationship between the size of the arguments (number of elements in the array) and the execution time measured by the execution time measurement function unit when the process "sum" is executed using arguments of two different data types in the second embodiment. This schematic diagram shows the type of machine learning model selected by the machine learning model selection function unit according to the type of process and data type in the second embodiment. This block diagram shows the schematic functional configuration of the execution time prediction device according to the third embodiment. This block diagram shows the schematic functional configuration of the execution time prediction device according to the fourth embodiment. This schematic diagram shows an example of source code to be analyzed in the processing example of the fourth embodiment. This schematic diagram shows an example of analysis results generated by the data structure analysis function unit by analyzing the source code shown in Figure 16. This schematic diagram shows an example of analysis results generated by the data attribute analysis function unit by analyzing the source code shown in Figure 16. This schematic diagram shows the result of the execution time calculation unit predicting the execution time of the process in the source code based on the analysis results in Figure 17.This is a schematic diagram showing a list of correction values ​​obtained by the execution time prediction correction unit for each process in the source code, based on the analysis results in Figure 18, and the correction results using those correction values. This is a schematic diagram showing another example of analysis results generated by the data attribute analysis function unit by analyzing the source code shown in Figure 16. This is a schematic diagram showing a list of correction values ​​obtained by the execution time prediction correction unit for each process in the source code, based on the analysis results in Figure 21, and the correction results using those correction values. This is a block diagram showing the schematic functional configuration of the execution time prediction device according to the fifth embodiment. This is a block diagram showing the schematic functional configuration of the execution time prediction device according to the sixth embodiment.

[0023] Next, several embodiments of the present invention will be described with reference to the drawings. The embodiments described below are execution time prediction devices for predicting processing time on a computer. The execution time prediction device of the embodiments predicts execution time based on the type of processing and the size of the arguments in that processing. In other words, the basis for predicting execution time is the type of processing and the size of the arguments in that processing. Furthermore, the data type of the arguments and the performance of the processor that executes the processing may also be used as the basis for predicting execution time. The execution time prediction device of the embodiments predicts the execution time for a specific processing with a specific argument size using a machine learning model. The execution time prediction device of the embodiments has a function for training a machine learning model. In addition, the execution time prediction device of the embodiments generates training data and allows for the selection of a machine learning model that is suitable for the training data from among several types of machine learning models.

[0024] It should be noted that, as a prerequisite for the embodiment, the execution time differs depending on the arguments passed to the computer program. In general programming languages, multiple data types can be used as arguments to be processed. In this case, the amount of memory used differs depending on the data type. This difference in memory size can affect the execution time of the process. Furthermore, even within a specific data type, the amount of memory used differs depending on the size of the data (for example, the number of elements in an array). If the amount of memory used by the program differs, the time required for processing such as memory access and memory management may change.

[0025] In the embodiments described below, the target programming language is a dynamically typed language. Programming languages ​​are broadly classified into statically typed languages ​​such as C, C++, Java, and Pascal, and dynamically typed languages ​​such as Python, JavaScript, Ruby, Perl, Lisp, and PHP.

[0026] In dynamically typed languages, even with the same code, the processing content differs depending on the data type of the arguments. In other words, even with the same code, the execution time of the processing differs depending on the data type of the arguments. In each embodiment described below, we take into account that the processing time differs depending on the data type of the arguments and attempt to make an accurate prediction of the execution time.

[0027] [First Embodiment] Figure 1 is a block diagram illustrating the schematic functional configuration of an execution time prediction device according to this embodiment. As shown in the figure, the execution time prediction device 1 includes a training data generation unit 11, an execution time prediction model construction unit 12, a source code storage unit 13, an execution time prediction unit 14, and a prediction result output unit 15. The execution time prediction device 1 performs the processing of an execution time prediction method, including the processing steps of each unit. The functions of the execution time prediction device 1 can be realized, for example, by a computer and a program. Each functional unit also has storage means as needed. The storage means are, for example, variables in the program or memory allocated by the execution of the program. Non-volatile storage means such as a magnetic hard disk drive or a solid-state drive (SSD) may also be used as needed. At least some of the functions of each functional unit may be realized as a dedicated electronic circuit instead of a program.

[0028] As shown in the figure, the training data generation unit 11 includes a sample data generation function unit 111 and an execution time measurement function unit 112. The execution time prediction model construction unit 12 includes a machine learning model selection function unit 121 and a machine learning model training function unit 122. The execution time prediction unit 14 includes a source code analysis unit 141 and an execution time calculation unit 142.

[0029] The training data generation unit 11 generates training data which is a set of the type of processing, the size of the argument passed to the processing, and the execution time when the processing is executed based on the size of the argument.

[0030] The sample data generation function unit 111 generates sample data for measuring the execution time of the process.

[0031] The sample data generation function unit 111 takes data type information as input and outputs that data type information and the generated sample data. The sample data generation function unit 111 passes the data type information and the generated sample data to the execution time measurement function unit 112. Specifically, the sample data generation function unit 111 uses the data type information to identify variable points and parameters for each data type and generates sample data. Parameters may be, for example, the number of elements in an array or the length of a string (number of characters). To identify variable points (points that can become parameters) for each data type, the sample data generation function unit 111 extracts data from usage examples (source code, etc.) to analyze the characteristics of the type or analyzes documentation.

[0032] Specifically, the sample data generation function unit 111 generates sample data by performing the following processing. The sample data generation function unit 111 acquires data type information. For example, the sample data generation function unit 111 reads data type information from internal memory, etc. The data types acquired by the sample data generation function unit 111 here are, for example, strings (str, string) and lists (list). Here, a list is an array of integer (int) type data.

[0033] The sample data generation function unit 111 then identifies variable parameters for each data type. For example, if the data type is a string (str), the variable parameter is the length of the string. For example, if the string is "Hello", the parameter, i.e., the length of the string, is 5. Also, for example, if the data type is a list (list), the variable parameter is the size of the array. For example, if the array of integers is [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], the parameter, i.e., the size of the array, is 10.

[0034] The sample data generation function unit 111 generates sample data by increasing or decreasing the values ​​of the above parameters according to the data type. In other words, if the data type is a string (str), the sample data generation function unit 111 generates strings of various lengths as sample data. If the data type is a list (list), the sample data generation function unit 111 generates arrays of integers of various sizes as sample data.

[0035] The sample data generation function 111 generates data of a predetermined parameter size by adding together the unit values ​​of a data type, for example, in the Python language. For example, when the data type is a string (str), the unit data is "a" (a string with length 1). When the data type is a string (str), the sample data generation function 111 adds up multiple "a" units to generate sample data such as "aa", "aaa", ... and so on. Also, for example, when the data type is a list (list), the unit data is [1] (an array of integers with one element). When the data type is a list (list), the sample data generation function 111 adds up multiple [1] units to generate sample data such as [1,1], [1,1,1], ... and so on.

[0036] The sample data generation function unit 111 passes the generated sample data to the execution time measurement function unit 112. The sample data generation function unit 111 also passes the generated sample data to the machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12.

[0037] The correspondence between data types and variable parameters may be stored in the internal memory of the sample data generation function unit 111 in advance, for example. Alternatively, the sample data generation function unit 111 may extract data from existing program source code (usage examples) to identify what the variable parameters are, by analyzing the characteristics of the data types or analyzing documentation.

[0038] The execution time measurement function unit 112 measures the execution time when a predetermined process is executed using the sample data generated by the sample data generation function unit 111.

[0039] The execution time measurement function unit 112 receives information on the type of process, data type information, and sample data as input. The execution time measurement function unit 112 then outputs a set of pairs of information on the type of process, data type information, sample data, and the measured execution time. The execution time measurement function unit 112 passes the output data to the machine learning model selection function unit 121 and the machine learning model training function unit 122. The execution time measurement function unit 112 has the processor execute a process using the sample data received as input data as arguments, and measures the execution time.

[0040] Specifically, the execution time measurement function unit 112 acquires information about the process. This information includes, for example, "sum" and "print". Here, "sum" is the process of calculating the sum of the integers contained in the list (list, the array of integers mentioned above) passed as an argument. In other words, the "sum" process calculates the sum of the elements of the array. "Print" is the process of outputting the string (str) passed as an argument to the screen (this can also be rephrased as the process of outputting to a predetermined output stream). The information about the process may be stored in the internal memory of the execution time measurement function unit 112 in advance.

[0041] The execution time measurement function unit 112 actually executes a predetermined process (for example, the above "sum", "print", etc.) and measures the execution time thereof. Specifically, the execution time measurement function unit 112 executes the process using various sample data generated by the sample data generation function unit 111 as arguments, and measures the execution time of the process corresponding to each argument. To measure the execution time, the execution time measurement function unit 112 can use, for example, an internal clock provided in a computer system.

[0042] The execution time measurement function unit 112 outputs a set including information on the type of process (for example, the type such as the above "sum" and "print"), a parameter of the argument used in the process (size of the argument), and the actually measured execution time. The execution time measurement function unit 112 passes this data set of execution results to the machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12.

[0043] The execution time prediction model construction unit 12 constructs an execution time prediction model for predicting the execution time of a process based at least on the type of process and the size of the argument. To this end, the machine learning model selection function unit 121 selects a machine learning model, and the machine learning model training function unit 122 trains the selected machine learning model using training data. Note that the type of process depends not only on the type of function to be executed (sum, print, etc.) but also on the data type of the argument.

[0044] The machine learning model selection function unit 121 selects a machine learning model from a plurality of candidate machine learning models based on the variation tendency of execution time corresponding to the process type and the argument size included in the training data generated by the training data generation unit 11.

[0045] Specifically, the machine learning model selection function unit 121 analyzes the relationship between the size of the argument and the execution time for each piece of process information and each piece of data type information of the argument. Then, the machine learning model selection function unit 121 selects an appropriate machine learning model from the plurality of candidate machine learning models based on the analysis result.

[0046] The machine learning model selection function unit 121 receives information on the type of processing, information on the data type, characteristics of sample data, and a set of execution times as inputs. The machine learning model selection function unit 121 also previously stores information on a set of machine learning model candidates to be selected. The machine learning model selection function unit 121 outputs information on the type of processing, information on the data type, characteristics of sample data, a set of execution times, and information on the selected machine learning model. The machine learning model selection function unit 121 transfers these output data to the machine learning model training function unit 122.

[0047] The machine learning model selection function unit 121 analyzes variations in execution time with respect to parameter variations for each processing type, and selects an appropriate machine learning model. For example, the machine learning model selection function unit 121 selects a machine learning model having the highest correlation based on the equation corresponding to each candidate machine learning model to be selected and the graph of execution time corresponding to parameter variations. In other words, the machine learning model selection function unit 121 selects a machine learning model for which the relationship between the parameter value of the argument to be processed and the variation in execution time measured by the execution time measurement function unit 112 is the most consistent. That is, the machine learning model selection function unit 121 selects a machine learning model that constitutes an equation close to the measured variation in execution time.

[0048] For each candidate machine learning model to be selected, the machine learning model selection function unit 121 may store a template of the machine learning model and information related to the equation corresponding to the machine learning model (e.g., an example of the equation). In this case, the machine learning model selection function unit 121 may calculate the degree of correlation between the equation represented by the training data transferred from the training data generation unit 11 (the relationship between the execution time and the size of the argument) and the equation of each model, and select the machine learning model with a high degree of correlation.

[0049] Alternatively, the machine learning model selection function unit 121 may perform a preliminary training of each candidate machine learning model using at least a portion of the training data passed from the training data generation unit 11, and calculate the difference between the execution time calculated by the machine learning model after preliminary training and the execution time of the training data. This difference is, for example, the mean squared error. The machine learning model with the smallest difference may then be selected.

[0050] The machine learning model training function unit 122 trains the machine learning model selected by the machine learning model selection function unit 121. In other words, the machine learning model training function unit uses the training data generated by the training data generation unit 11 to train the machine learning model selected by the machine learning model selection function unit 121 and construct an execution time prediction model.

[0051] The machine learning model training unit 122 receives information about the type of processing, data type information, a set of pairs of sample data and execution time, and information about the selected machine learning model. The machine learning model training unit 122 outputs a trained execution time prediction model. In other words, the machine learning model training unit 122 outputs a set of intrinsic parameter values ​​for the machine learning model obtained as a result of training.

[0052] Specifically, the machine learning model training function unit 122 trains the machine learning model based on the relationship between the parameter size values ​​and the execution time of the process measured by the execution time measurement function unit 112, for each type of process ("sum", "print", etc.) and each data type (list (list), string (str), etc.). The data used by the machine learning model training function unit 122 for training the machine learning model is the data generated by the training data generation unit 11. In other words, the machine learning model training function unit 122 trains the machine learning model based on the relationship between the sample data generated by the sample data generation function unit 111 and the execution time of the process measured by the execution time measurement function unit 112 using that sample data.

[0053] The machine learning model training function unit 122 performs the above training, thereby adjusting the internal parameters of the machine learning model selected by the machine learning model selection function unit 121. In other words, the internal parameters of the machine learning model selected by the machine learning model selection function unit 121 are adjusted to be optimal using the training data generated by the training data generation unit 11. The machine learning model training function unit 122 stores the values ​​of the model's internal parameters obtained as a result of the training in memory or the like.

[0054] The source code storage unit 13 stores the source code of the program. The source code stored by the source code storage unit 13 is the source code that the execution time prediction unit 14 uses to predict the execution time.

[0055] The execution time prediction unit 14 inputs the types of processes included in a given program and the sizes of the arguments passed to those processes into an execution time prediction model constructed by the machine learning model training function unit 122, thereby predicting the execution time of the processes corresponding to the types of processes and the sizes of the arguments. To this end, the source code analysis unit 141 analyzes the source code of the program. The execution time calculation unit 142 inputs the results analyzed by the source code analysis unit 141 into a machine learning model trained by the machine learning model training function unit 122, thereby predicting the execution time of the processes. The execution time calculation unit 142 also calculates the execution time of the program by summing the execution times (predicted values) for all processes included in the program.

[0056] In other words, the execution time calculation unit 142 of the execution time prediction unit 14 predicts the execution time by inputting the type of processing output by the source code analysis unit 141 and the size of the arguments passed to the processing into the execution time prediction model.

[0057] The source code analysis unit 141 analyzes the source code read from the source code storage unit 13 and extracts information necessary for predicting the execution time of the process.

[0058] Specifically, the source code analysis unit 141 takes the source code to be analyzed as input. The source code analysis unit 141 extracts information on the type of processing and argument information that appear in the source code and outputs the analysis result information. The analysis result information output by the source code analysis unit 141 includes information on the type of processing and information on the data type and size of the arguments of the processing. The source code analysis unit 141 passes the above output information to the execution time calculation unit 142. In other words, by analyzing the source code of a program, the source code analysis unit 141 outputs information on the type of processing included in the program and the size of the arguments passed to the processing. The source code analysis process performed by the source code analysis unit 141 can be carried out using existing technologies.

[0059] The execution time calculation unit 142 receives information about the type of process and the size of the arguments for each process included in the source code. Based on this input data, the execution time calculation unit 142 uses a trained machine learning model to calculate a predicted execution time for each process. The execution time calculation unit 142 then sums up the predicted execution times for all processes. The execution time calculation unit 142 passes the calculated predicted execution times to the prediction result output unit 15. When predicting execution time, the execution time calculation unit 142 uses a machine learning model that has been selected by the machine learning model selection function unit 121 and trained by the machine learning model training function unit 122, at least according to the type of process and the data type of the arguments.

[0060] The prediction result output unit 15 outputs to the outside the predicted execution time calculated by the execution time calculation unit 142 in relation to a specific source code.

[0061] In the aforementioned "print" process, the output data is buffered before being displayed on the screen. If the output data exceeds the buffer size, the "print" process processes the data up to the buffer size and then clears the buffer. Therefore, if the size of the output data is large, the execution time will increase. Also, in the aforementioned "sum" process, each element of the array is added up from top to bottom, and the sum is stored in memory. However, if the sum cannot be stored in the available memory, additional memory is allocated. Therefore, if the size of the calculation result data is large, the execution time will increase.

[0062] Figure 2 is a schematic diagram showing examples of multiple types of machine learning models that can be selected by the machine learning model selection function unit 121. As shown in the figure, the machine learning models that can be selected by the machine learning model selection function unit 121 may include simple regression models, multiple regression models, support vector regression models, polynomial regression models, regression tree models, and random forest regression models. In the figure, the relationship between the sample points and the graphs represented by each model is shown. The horizontal axis of each graph represents the parameter value (size of the argument), and the vertical axis represents the execution time of the process. For example, a simple regression model regresses to a straight line (a line represented by a linear equation). A polynomial regression model regresses to a curve represented by a polynomial.

[0063] The machine learning model selection function unit 121 may store information about the mathematical formulas for each of the selectable machine learning models. These formulas represent regression curves.

[0064] Figure 3 is an example of a graph showing the relationship between parameter values ​​(argument sizes) and execution time for a given process. The type of process in Figure 3 is the aforementioned "sum" process. The unit of the vertical axis in this graph is seconds (sec). The graph shown in Figure 3 represents a set of pairs of parameter values ​​(argument sizes) for each sample data generated by the sample data generation function unit 111 and the execution time measured by the execution time measurement function unit 112 based on that sample data. In other words, the graph shown in Figure 3 can be drawn based on these sets of pairs. The machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12 each receive the data of these pairs from the training data generation unit 11. As this graph shows, execution time generally increases as the argument size increases.

[0065] Figure 4 is another example of a graph showing the relationship between parameter values ​​(argument sizes) and execution time for a given process. The type of process in Figure 4 is the aforementioned "print". The unit of the vertical axis in this graph is seconds (sec). The graph shown in Figure 4 represents a set of pairs of parameter values ​​(argument sizes) for each sample data generated by the sample data generation function unit 111 and the execution time measured by the execution time measurement function unit 112 based on that sample data. In other words, the graph shown in Figure 4 can be drawn based on a set of such pairs. The machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12 each receive the data of these pairs from the training data generation unit 11. As this graph shows, although there are local increases and decreases, the execution time generally increases as the argument size increases.

[0066] Figure 5 is a schematic diagram showing an example of data representing the relationship between parameter size and measured execution time. As shown in the figure, this data pertains to the aforementioned "sum" process. This data also pertains to the case where the argument data type is "list" (an array of integers). As shown in the figure, this data indicates that the execution time is 4 milliseconds (ms) when the parameter size is 1, 4 milliseconds when the parameter size is 2, and 5 milliseconds when the parameter size is 3. This data may also contain information on the execution time when the parameter size is 4 or greater. The machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12 each receive data like the one shown in this figure from the training data generation unit 11.

[0067] Figure 6 is a schematic diagram showing another example of data representing the relationship between parameter size and measured execution time. As shown in the figure, this data pertains to the "print" process described above. This data also pertains to the case where the argument data type is "str" ​​(string). As shown in the figure, this data indicates that the execution time is 10 milliseconds (ms) when the parameter size is 1, 14 milliseconds when the parameter size is 2, and 29 milliseconds when the parameter size is 3. This data may also contain information on the execution time when the parameter size is 4 or greater. The machine learning model selection function unit 121 and the machine learning model training function unit 122 of the execution time prediction model construction unit 12 each receive data like the one shown in this figure from the training data generation unit 11.

[0068] Figure 7 is a schematic diagram showing the results of the machine learning model selection performed by the machine learning model selection function unit 121. Note that row numbers are added in this figure for convenience. The machine learning model selection function unit 121 selects a machine learning model based on data such as those in Figures 5 and 6. The first row of data in Figure 7 indicates that when the process is "sum" and the argument data type is "list" (an array of integers), the machine learning model selected by the machine learning model selection function unit 121 is a simple linear regression model. This was selected by the machine learning model selection function unit 121 based on the data shown in Figure 5. The second row of data in Figure 7 indicates that when the process is "print" and the argument data type is "str" ​​(a string), the machine learning model selected by the machine learning model selection function unit 121 is a regression tree model. This was selected by the machine learning model selection function unit 121 based on the data shown in Figure 6. The machine learning model selection function unit 121 passes information about the type of machine learning model selected for a specific data type of argument in a specific process to the machine learning model training function unit 122.

[0069] Figure 8 is a schematic diagram showing an example of source code stored in the source code storage unit 13. For convenience, line numbers are included in this diagram. As shown in the diagram, the first line of this source code represents the process of executing the "sum" operation and assigning the result to the variable a. The argument passed to the "sum" operation is of type list (list), and is an array of integers [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]. The second line of this source code represents the execution of the "print" operation. The argument passed to the "print" operation is of type string (str), and is "Hello".

[0070] Figure 9 is a schematic diagram showing an example of the source code analysis results by the source code analysis unit 141 of the execution time prediction unit 14. The analysis result of the first line in Figure 9 corresponds to the first line of the source code in Figure 8. In other words, by analyzing the first line of the source code in Figure 8, the source code analysis unit 141 outputs the analysis result that the process is "sum", the data type of the argument is "list" (an array of integers), and the size of the parameter is 10. The analysis result of the second line in Figure 9 corresponds to the second line of the source code in Figure 8. In other words, by analyzing the second line of the source code in Figure 8, the source code analysis unit 141 outputs the analysis result that the process is "print", the data type of the argument is "str" ​​(a string), and the size of the parameter is 5. The source code analysis unit 141 can output these analysis results by performing lexical analysis of the source code.

[0071] Figure 10 is a block diagram showing an example of the internal configuration of the execution time prediction device 1. The execution time prediction device 1 can be implemented using a computer. As shown in the figure, the computer is composed of a central processing unit 901, RAM 902, input / output ports 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions contained in programs read from RAM 902, etc. The central processing unit 901 writes data to RAM 902, reads data from RAM 902, and performs arithmetic and logical operations according to each instruction. RAM 902 stores data and programs. Each element contained in RAM 902 has an address and can be accessed using that address. RAM stands for "Random Access Memory". Input / output ports 903 are ports for the central processing unit 901 to exchange data with external input / output devices, etc. Input / output devices 904 and 905 exchange data with the central processing unit 901 via input / output ports 903. Bus 906 is a common communication channel used within the computer. For example, the central processing unit 901 reads and writes data to RAM 902 via bus 906. Also, for example, the central processing unit 901 accesses input / output ports 903 via bus 906.

[0072] Furthermore, at least some of the functions of the execution time prediction device 1 in this embodiment can be realized by a computer and a program. In that case, the program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Here, "computer system" includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, DVD-ROMs, USB memory, and storage devices such as hard disks built into a computer system. In other words, "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Moreover, "computer-readable recording medium" may also include those that temporarily and dynamically hold programs, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside a computer system that acts as a server or client in that case. Furthermore, the above program may be for realizing some of the functions described above, and may also be able to realize the above functions in combination with a program already recorded in the computer system.

[0073] The method by which the machine learning model training function unit 122 trains a machine learning model is as follows:

[0074] A machine learning model calculates output data (in this embodiment, predicted execution time) based on input data. The machine learning model utilizes the values ​​of internal parameters when calculating the output data. These internal parameters can be updated and optimized through training. Training data is used when training a machine learning model. The training data includes input data for the machine learning model and the correct output data calculated based on that input data. During training, the machine learning model reads the input data included in the training data and calculates output data based on that input data and using the internal parameter values ​​at that time. This output data is an estimate obtained based on the internal parameters at that time and does not necessarily coincide with the correct answer. To update the internal parameters, the difference between the estimate calculated and output by the machine learning model based on the input data at that time and the correct answer corresponding to that input data is calculated. This difference is called error, loss, etc. The difference calculated here may be, for example, the absolute value of the difference between scalars, a squared error, a cross-entropy error, or a difference calculated by other methods. Based on the calculated difference, the values ​​of the internal parameters can be updated using backpropagation. This operation adjusts the values ​​of the internal parameters in a direction that reduces the error. By performing the above operation multiple times (many times) using a predetermined amount of training data, the values ​​of the internal parameters are optimized. In other words, the machine learning model is adjusted to perform the processing exemplified by the given training data. After a sufficient amount of training, training may be terminated. By storing the set of learned internal parameter values ​​in a memory device, it becomes possible to perform estimations based on the training results. In other words, a machine learning model includes internal parameters. The training of a machine learning model may also be called "learning".

[0075] As described above, according to the first embodiment, the machine learning model selection function unit selects a machine learning model that is suitable for the sample data from among several candidate machine learning models. Furthermore, the machine learning model training function unit can train the machine learning model selected by the machine learning model selection function unit and construct a model for predicting the execution time of the program.

[0076] [Second Embodiment] Next, a second embodiment of the present invention will be described. Note that matters already described in the previous embodiment may be omitted below. Here, we will focus on matters specific to this embodiment.

[0077] Figure 11 is a block diagram illustrating the schematic functional configuration of the execution time prediction device according to this embodiment. The execution time prediction device 2 is based on the configuration of the execution time prediction device 1 described above, but has further features. As shown in the figure, the execution time prediction device 2 is composed of a training data generation unit 11, an execution time prediction model construction unit 22, a source code storage unit 13, an execution time prediction unit 24, and a prediction result output unit 15. The functions of the execution time prediction device 2 can be realized, for example, by a computer and a program, as in the previous embodiment.

[0078] As shown in the figure, the training data generation unit 11 includes a sample data generation function unit 111 and an execution time measurement function unit 112. The execution time prediction model construction unit 22 includes a machine learning model selection function unit 221 and a machine learning model training function unit 222. The execution time prediction unit 24 includes a source code analysis unit 241 and an execution time calculation unit 242.

[0079] In this embodiment, the machine learning model selection function unit 221 selects a machine learning model according to the content of the sample data included in the training data passed from the training data generation unit 11.

[0080] The content of the sample data is, for example, the size of each element in an array. For example, consider two types of data types, "list1" and "list2". Here, data type "list1" is an array of integers (int), and its unit value is [1]. In other words, the amount of memory occupied by one element in an array of data type "list1" is relatively small. On the other hand, data type "list2" is an array of integers (int), and its unit value is [1000000000]. In other words, the amount of memory occupied by one element in an array of data type "list2" is relatively large.

[0081] As an example of a type of process, let's consider the process "sum". The process "sum" calculates the sum of the elements in an array of integers that is passed as an argument. In other words, when an argument of data type "list1" is passed, the process "sum" calculates the sum of the elements of that array (each element occupies a relatively small amount of memory). When an argument of data type "list2" is passed, the process "sum" calculates the sum of the elements of that array (each element occupies a relatively large amount of memory). In other words, if the number of elements in the arrays is the same, the execution time of the process "sum" for data type "list2" tends to be longer than the execution time of the process "sum" for data type "list1". This is because there is a difference in the time required for processing to allocate working memory during the execution of the process "sum".

[0082] In this embodiment, the sample data generation function unit 111 identifies variable parameters for each data type, similar to the first embodiment. The sample data generation function unit 111 then generates sample data by increasing or decreasing the values ​​of the above parameters according to the data type.

[0083] In the example above (data types "list1" and "list2"), the sample data generation function unit 111 generates sample data for each of the data types "list1" and "list2".

[0084] The execution time measurement function unit 112 takes each of the sample data generated by the sample data generation function unit 111 as an argument and executes a process such as "sum", and measures the execution time at that time.

[0085] Figure 12 is a graph showing the relationship between the size of the arguments (number of elements in the array) and the execution time measured by the execution time measurement function 112 when the process "sum" is executed using two types of data type arguments in this embodiment. In the figure, G1 is the graph when the process "sum" is executed with an argument of data type "list1". G2 is the graph when the process "sum" is executed with an argument of data type "list2". Overall, when comparing with the same argument size (number of elements in the array), the processing time shown in graph G2 is longer than that shown in graph G1. In addition, although there are local increases and decreases in graphs G1 and G2, the overall trend is that the execution time is longer as the size of the arguments increases.

[0086] Similar to the first embodiment, the machine learning model selection function 221 selects an appropriate machine learning model based on the relationship between the argument size and execution time shown in the graph in Figure 12. As also shown in Figure 12, when the argument is of data type "list1" (in the case of graph G1), the machine learning model selection function 221 selects a simple linear regression model as the machine learning model. When the argument is of data type "list2" (in the case of graph G2), the machine learning model selection function 221 selects two types of machine learning models according to the size of the argument. Specifically, for arguments of data type "list2" with a size in the range from 0 to 400, the machine learning model selection function 221 selects a regression tree model. For arguments of data type "list2" with a size of 400 or more (from 400 to 1000), the machine learning model selection function 221 selects a simple linear regression model.

[0087] As described above, the machine learning model selection function 221 may select different machine learning models for the same type of processing (e.g., processing "sum") depending on the data type of the argument. Alternatively, the machine learning model selection function 221 may select different machine learning models for a single type of processing (e.g., processing "sum") and a single data type of argument (e.g., data type "list2") depending on the range of the argument size (e.g., less than 400 or 400 or more).

[0088] Figure 13 is a schematic diagram showing the types of machine learning models selected by the machine learning model selection function unit 221 according to the type of processing and data type. As shown in the figure, when the type of processing is "sum" and the data type is "list1", the machine learning model selection function unit 221 selects a simple linear regression model. When the type of processing is "sum" and the data type is "list2", the machine learning model selection function unit 221 selects both a regression tree model and a simple linear regression model. The selection results by the machine learning model selection function unit 221 are consistent with the graph in Figure 12. Note that when the type of processing is "sum" and the data type is "list2", the machine learning model selection function unit 221 selects a regression tree model for argument sizes from 0 to 400, and selects a simple linear regression model for argument sizes from 400 to 1000.

[0089] The processing of the machine learning model training function 222 and the execution time prediction unit 24 after the machine learning model selection function 221 has selected a machine learning model is basically the same as in the first embodiment. However, the machine learning model training function 222 trains the machine learning model selected by the machine learning model selection function 221 for each of the argument data types using appropriate training data. In addition, the source code analysis unit 241 of the execution time prediction unit 24 determines whether the source code to be predicted is a process that uses an argument of data type "list1" or a process that uses an argument of data type "list2". In other words, the source code analysis unit 241 identifies the data type of the argument. The source code analysis unit 241 also identifies which region the size of the argument belongs to. In addition, the execution time calculation unit 242 calculates (predicts) the execution time using the machine learning model selected by the machine learning model selection function 221 and trained by the machine learning model training function 222 according to the data type and size of the argument. The prediction result output unit 15 outputs a predicted value for the execution time of the target source code.

[0090] In other words, in the second embodiment, the machine learning model selection function 221 selects a machine learning model according to the size of the arguments (or the region to which that size belongs). When the machine learning model training function 222 trains the selected machine learning model, it uses data from the training data in which the size of the arguments (or the region to which that size belongs) matches.

[0091] As described above, according to the second embodiment, for example, for one type of process (process "sum", etc.), the machine learning model selection function unit selects a machine learning model that is suitable for the sample data from among multiple candidate machine learning models for each argument of a different data type. Furthermore, the machine learning model training function unit can train the machine learning model selected by the machine learning model selection function unit and construct a model for predicting the execution time of the program.

[0092] [Third Embodiment] Next, a third embodiment of the present invention will be described. Note that matters already described in the previous embodiments may be omitted below. Here, the focus will be on matters specific to this embodiment.

[0093] Figure 14 is a block diagram illustrating the schematic functional configuration of the execution time prediction device according to this embodiment. The execution time prediction device 3 is based on the configurations of the execution time prediction device 1 and execution time prediction device 2 described above, but has further features. As shown in the figure, the execution time prediction device 3 is composed of a training data generation unit 31, an execution time prediction model construction unit 32, a source code storage unit 13, an execution time prediction unit 34, and a prediction result output unit 15. The functions of the execution time prediction device 3 can be realized, for example, by a computer and a program, as in the previous embodiments.

[0094] As shown in the figure, the training data generation unit 31 includes a sample data generation function unit 311 and an execution time measurement function unit 312. The execution time prediction model construction unit 32 includes a machine learning model selection function unit 321 and a machine learning model training function unit 322. The execution time prediction unit 34 includes a source code analysis unit 341 and an execution time calculation unit 342.

[0095] In this embodiment, the machine learning model selection function 321 selects a machine learning model not only based on the type of processing and the size of the arguments, but also based on the performance of the processor, such as the CPU (Central Processing Unit), when executing the processing.

[0096] Here, as in the first embodiment, we consider "sum" and "print" as examples of processing types. We also consider "int" (integer) and "str" ​​(string) as argument data types. The contents of "sum" and "print" are as already explained. In this embodiment, as an example of processor performance, we consider the clock frequency for driving the processor. The processor frequency is expressed in units such as gigahertz (GHz) or terahertz (THz).

[0097] In this embodiment, the sample data generation function 311 identifies variable parameters for each data type, similar to the case in the first embodiment. The sample data generation function 311 then generates sample data by increasing or decreasing the values ​​of the above parameters according to the data type. The sample data generation function 311 also generates the above sample data for each of the multiple processors (CPU, etc.). For example, the sample data generation function 311 generates sample data for each processor performance. The processor performance can be represented, for example, by the processor's clock frequency (e.g., F1, F2, ...).

[0098] The execution time measurement function unit 312 executes a process for each processor using the sample data arguments generated by the sample data generation function unit 311, and measures the execution time. As a result, the execution time measurement function unit 312 outputs data consisting of a set of the type of process, the size of the arguments, the processor performance (represented, for example, by the clock frequency), and the measured execution time corresponding to these.

[0099] The machine learning model selection function unit 321 of the execution time prediction model construction unit 32 selects a machine learning model based on the data output by the execution time measurement function unit 312. In other words, the machine learning model selection function unit 321 analyzes the variation in execution time corresponding to changes in argument size for each processor performance (clock frequency, etc.) and each type of processing. The machine learning model selection function unit 321 selects a machine learning model that constructs an equation similar to the variation in execution time.

[0100] The machine learning model training unit 322 uses the training data provided by the training data generation unit 31 to train the machine learning model selected by the machine learning model selection unit 321. The machine learning model training unit 322 trains the machine learning model selected by the machine learning model selection unit 321 for each type of processing and for each processor performance (clock frequency, etc.). The machine learning model training unit 322 stores the internal parameters of the machine learning model obtained as a result of the training in memory or the like.

[0101] The execution time calculation unit 342 calculates (predicts) the execution time of the process based on the results of the source code analysis unit 341 and the performance of the processor that will execute the process. The prediction result output unit 15 outputs the predicted execution time when the target source code is executed on a processor with a specific performance level to the outside.

[0102] In other words, in the third embodiment, the machine learning model calculates (predicts) the execution time of the process in accordance with the processor's performance information. In this case, the training data for training the machine learning model includes the processor's performance information. Furthermore, when calculating the execution time using the trained machine learning model, the processor's performance information is also used as input to the machine learning model.

[0103] As described above, according to the third embodiment, for each processor having different performance characteristics, the machine learning model selection function unit selects a machine learning model that is suitable for the sample data from among multiple candidate machine learning models. Furthermore, the machine learning model training function unit can train the machine learning model selected by the machine learning model selection function unit and construct a model for predicting the execution time of a program.

[0104] [Fourth Embodiment] Next, a fourth embodiment of the present invention will be described. Note that matters already described in the previous embodiments may be omitted below. Here, the focus will be on matters specific to this embodiment.

[0105] This embodiment adds further functionality to the configurations of the first to third embodiments. Similar to the problems that the first to third embodiments aim to solve, cloud services require the selection of optimal resources in order to respond quickly and efficiently to application program execution requests. Predicting program execution time with high accuracy is an important element for ensuring service quality, etc.

[0106] In the first to third embodiments, there is a problem in that when the attributes of the data used for training prediction differ from the attributes of the data in the actual program being predicted, this can cause a discrepancy with the actual execution time. In other words, there are cases where predictions cannot be made with high accuracy.

[0107] In other words, in order to execute a process, it is necessary to prepare data such as argument data. The time required to prepare the data can vary greatly depending on the data attributes.

[0108] On the other hand, it is conceivable to construct a model that predicts the execution time of a process, taking into account all data attributes. However, this would present problems such as the extremely large time required to construct the prediction model and the large memory size needed to store it.

[0109] This embodiment focuses on the fact that data preparation affects the prediction results. This embodiment is characterized by applying a correction value that reflects the differences in data preparation to the predicted processing execution time. As a result, this embodiment makes it possible to predict the execution time of the entire program with even greater accuracy.

[0110] Figure 15 is a block diagram illustrating the schematic functional configuration of the execution time prediction device according to the fourth embodiment. The execution time prediction device 4 is based on the configurations of the execution time prediction devices 1 to 3 described above, but has further features. As shown in the figure, the execution time prediction device 4 is composed of a training data generation unit 41, an execution time prediction model construction unit 42, a source code storage unit 13, an execution time prediction unit 44, and a prediction result output unit 15. The functions of the execution time prediction device 4 can be realized, for example, by a computer and a program, as in the previous embodiments.

[0111] As shown in the figure, the training data generation unit 41 includes a sample data generation function unit 411 and an execution time measurement function unit 412. The execution time prediction model construction unit 42 includes a machine learning model selection function unit 421 and a machine learning model training function unit 422.

[0112] A key feature of this embodiment lies in the functional configuration of the execution time prediction unit 44. The execution time prediction unit 44 includes a source code analysis unit 441, an execution time calculation unit 442, an execution time prediction value correction unit 443, and an aggregation unit 444. The source code analysis unit 441 also includes a data structure analysis unit 4411 and a data attribute analysis unit 4412.

[0113] The program's execution time includes data preparation time and internal processing time. Of these times, data preparation time depends on data attributes, while internal processing time does not. This embodiment corrects the predicted execution time using a correction value related to data preparation time, which depends on data attributes. In other words, this embodiment allows data attributes to be reflected in the execution time prediction model.

[0114] In this embodiment, data type information is the format of the data to be processed. Examples of data types include "str" ​​and "list," which will be described later. Processing information is information that identifies the processing content, and is represented, for example, by the processing name (function name, etc.) or processing code. Machine learning model candidate information is information that shows the template of a model for a specific machine learning algorithm and the characteristics of the equations it forms. Machine learning model candidate information may also be, for example, an example of an equation. Data attribute information is information about the attributes of the target data. In this embodiment, data attributes are, for example, the distinction between constants and variables. Alternatively, variables may be further distinguished as data attributes, for example, by local scope variables and global scope variables.

[0115] The general functions of each component of the execution time prediction device are as follows:

[0116] The training data generation unit 41 generates training data for machine learning.

[0117] The sample data generation function unit 411 within the training data generation unit 41 generates sample data based on the given data type information. Specifically, the sample data generation function unit 411 identifies the variable locations and parameters for each data type and generates sample data. To identify the variable locations and parameters, the sample data generation function unit 411 extracts data from usage examples, analyzes the characteristics of the data type, and analyzes documentation. The sample data generation function unit 411 passes the data type information and the generated sample data to the execution time measurement function unit 412.

[0118] The execution time measurement function unit 412 within the training data generation unit 41 receives data type information and sample data generated by the sample data generation function unit 411 from the sample data generation function unit 411. The execution time measurement function unit 412 measures the execution time of a process that uses the received sample data as an argument. The execution time measurement function unit 412 passes information identifying the process, data type information, information representing the characteristics of the sample data, and the measured execution time information to the machine learning model selection function unit 421. The execution time measurement function unit 412 may perform multiple execution time measurements based on multiple sample data. The execution time measurement function unit 412 passes the execution time information, which is the result of all those trials, to the machine learning model selection function unit 421.

[0119] The execution time prediction model construction unit 42 trains a machine learning model using the training data generated by the training data generation unit 41. In other words, the execution time prediction model construction unit 42 constructs a processing execution time prediction model.

[0120] The machine learning model selection function unit 421 of the execution time prediction model construction unit 42 selects a machine learning model. Specifically, the machine learning model selection function unit 421 receives information identifying the process, data type information, information representing the characteristics of the sample data, and execution time information measured by the execution time measurement function unit 412 from the execution time measurement function unit 412. The machine learning model selection function unit 421 also has information on candidate machine learning models in advance. The machine learning model selection function unit 421 analyzes the execution time for each process and selects an appropriate machine learning model for the variation in execution time based on the information on candidate machine learning models.

[0121] For example, if the machine learning model selection function 421 has information on the generated equations as information on candidate machine learning models, the machine learning model selection function 421 may select the one that has the most overlap with the execution time graph. In other words, the machine learning model selection function 421 may select a machine learning model based on the degree of overlap between the set of measured execution times and the execution time graph it has in place.

[0122] The machine learning model selection function unit 421 passes processing information, data type information, sample data feature information, execution time information, and selected machine learning model information to the machine learning model training function unit 422.

[0123] The machine learning model training function unit 422 of the execution time prediction model construction unit 42 trains a machine learning model. Specifically, the machine learning model training function unit 422 receives processing information, data type information, sample data characteristics, execution time information, and information on the machine learning model selected by the machine learning model selection function unit 421 from the machine learning model selection function unit 421. The machine learning model training function unit 422 trains a machine learning model using the sample data characteristics and execution time for each piece of processing information and each piece of data type information. In other words, the machine learning model training function unit 422 constructs a processing execution time prediction model.

[0124] The machine learning model training function unit 422 passes processing information, data type information, and a processing execution time prediction model to the execution time calculation unit 442 of the execution time prediction unit 44.

[0125] The execution time prediction unit 44 predicts the execution time of the process using a trained process execution time prediction model.

[0126] The source code analysis unit 441 of the execution time prediction unit 44 analyzes the program source code read from the source code storage unit 13.

[0127] The data structure analysis unit 4411 of the source code analysis unit 441 performs an analysis of the data structure in the source code read from the source code storage unit 13. Specifically, the data structure analysis unit 4411 reads the source code from the source code storage unit 13. The data structure analysis unit 4411 analyzes the read source code and extracts the processes and their arguments that appear in the source code. The data structure analysis unit 4411 outputs the process information and the data type information of the arguments corresponding to that process. As an analysis result, the data structure analysis unit 4411 passes the process information and the data type information of the arguments corresponding to that process to the execution time calculation unit 442. An example of the analysis result by the data structure analysis unit 4411 will be explained later with reference to the diagram.

[0128] In other words, the data structure analysis unit 4411 of the source code analysis unit 441 analyzes the source code of the program and outputs information about the type of processing included in the program and the size of the arguments passed to the processing.

[0129] The data attribute analysis unit 4412 of the source code analysis unit 441 performs analysis of the data attributes in the source code read from the source code storage unit 13. Specifically, the data attribute analysis unit 4412 reads the source code from the source code storage unit 13. The data attribute analysis unit 4412 analyzes the read source code, extracts the data attributes in the processes that appear, and records only those data attributes. The data attribute analysis unit 4412 outputs a set of pairs of information about the processes in the source code and the data attribute information corresponding to those processes as the analysis result. The data attribute analysis unit 4412 passes the data attribute information, which is the analysis result, to the execution time prediction value correction unit 443.

[0130] In other words, when the data attribute analysis unit 4412 analyzes the source code read from the source code storage unit 13, it records and classifies the data state based on the usage of functions and arguments. The data state refers to distinctions such as whether the data is a constant or a variable. In short, the data attribute analysis unit 4412 analyzes the attributes of the data. The data state obtained by the data attribute analysis unit 4412 analyzing the source code is useful information for correcting the predicted processing execution time.

[0131] In other words, the data attribute analysis unit 4412 of the source code analysis unit 441 outputs information about the data attributes of the arguments in the source code by analyzing the source code. An example of the analysis result by the data attribute analysis unit 4412 will be explained later with reference to the drawings.

[0132] The execution time calculation unit 442 of the execution time prediction unit 44 estimates (predicts) the processing time using the processing execution time prediction model (machine learning model) trained by the machine learning model training function unit 422 and the information passed from the data structure analysis unit 4411. The execution time calculation unit 442 receives the processing execution time prediction model (machine learning model) trained in the machine learning model training function unit 422. The execution time calculation unit 442 also receives information about the source code processing and information about the data types of the arguments of that processing from the data structure analysis unit 4411. Specifically, the execution time calculation unit 442 receives a processing execution time prediction model (machine learning model) from the machine learning model training function unit 422 that corresponds to the processing information and the information about the data types of the arguments of that processing received from the data structure analysis unit 4411. The execution time calculation unit 442 predicts the execution time of the processing by applying the processing information and the information about the data types of the arguments of that processing to the processing execution time prediction model (machine learning model). The execution time calculation unit 442 passes the predicted execution time to the execution time prediction value correction unit 443.

[0133] In other words, the execution time calculation unit 442 of the execution time prediction unit 44 predicts the execution time by inputting the type of processing output by the data structure analysis unit 4411 of the source code analysis unit 441 and the size of the arguments passed to the processing into the execution time prediction model.

[0134] The execution time prediction value correction unit 443 of the execution time prediction unit 44 corrects the predicted execution time of the process passed from the execution time calculation unit 442 based on data attributes. Specifically, the execution time prediction value correction unit 443 receives information about the source code process and the data attributes of the data corresponding to that process from the data attribute analysis unit 4412. The execution time prediction value correction unit 443 compares the data attributes of the sample data used to train the processing execution time prediction model (machine learning model) with the data attributes of the relevant data passed from the data attribute analysis unit 4412. The execution time prediction value correction unit 443 determines a correction value according to this comparison result. The execution time prediction value correction unit 443 corrects the predicted execution time of the process passed from the execution time calculation unit 442 using the determined correction value.

[0135] In other words, the execution time prediction value correction unit 443 reflects a correction value in the execution time prediction value calculated by the execution time calculation unit 442, based on the data attributes of the function and arguments which are the analysis results of the data attribute analysis unit 4412 described above.

[0136] In other words, the execution time prediction value correction unit 443 of the execution time prediction unit 44 corrects the execution time of the process predicted by inputting the type of process and the size of the arguments passed to the process into the execution time prediction model, based on the data attributes output by the data attribute analysis unit 4412 of the source code analysis unit 441.

[0137] The correction values ​​used for correction are calculated based on the difference in execution time for each attribute of the relevant data. The execution time for each data attribute is measured at some point by the execution time measurement function unit 412, the execution time prediction value correction unit 443, or any other processing unit.

[0138] The execution time prediction value correction unit 443 passes the corrected execution time prediction value to the aggregation unit 444.

[0139] If the source code contains multiple processes, the execution time calculation unit 442 calculates an estimated execution time for each process. The execution time estimated value correction unit 443 then corrects the estimated execution time for each process. The execution time estimated value correction unit 443 then passes all of the corrected estimated execution time values ​​to the aggregation unit 444.

[0140] The aggregation unit 444 of the execution time prediction unit 44 aggregates the corrected predicted values ​​of the execution times of the processes included in the source code. In other words, the aggregation unit 444 receives a set of corrected predicted execution times for each process included in the source code from the execution time prediction value correction unit 443. The aggregation unit 444 sums up the corrected predicted execution times of each process to obtain the predicted execution time of the source code. The aggregation unit 444 passes the calculated predicted execution time of the source code to the prediction result output unit 15.

[0141] The prediction result output unit 15 outputs the calculated predicted execution time of the source code (after correction and aggregation), which is received from the aggregation unit 444, to an external source.

[0142] [First Processing Example of the Fourth Embodiment] Next, the first processing example in the fourth embodiment will be described. In this example, "str" ​​and "list" are used as data types. The data type "str" ​​is string data, and its unit value is a. The data type "list" is an array of data types "int" (integer), and its unit value is [1]. The processing included in the source code consists of the function "print" and the function "sum". The "print" process is a function that outputs str type data to the screen. The "sum" process is a function that calculates the sum of all elements (int type) of the list type data. In this example, there are two types of data attributes: "variable" and "constant". In other words, a data attribute is an attribute that indicates at least whether the data (argument) is a constant or a variable.

[0143] Figure 16 is a schematic diagram showing an example of the source code in this example. In this figure, each line of the source code is numbered. In other words, the source code in this example consists of 6 lines. The first line of the source code assigns the string "Hello" to the str type variable message. The second line of the source code performs the print function with the str type constant "I want to say" as an argument. The third line of the source code performs the print function with the message variable, which was assigned a value in the first line, as an argument. The fourth line of the source code assigns the 3-element array [1, 2, 3] to the list type variable set. The fifth line of the source code performs the sum function with the list type constant [1, 1, 1, 1, 1, 1] as an argument. The sixth line of the source code takes the variable `set`, which was assigned a value in the fourth line, as an argument and performs the operation of the `sum` function.

[0144] In other words, the source code in this example includes four operations: a print operation that takes a constant of type str as an argument, a print operation that takes a variable of type str as an argument, a sum operation that takes a constant of type list as an argument, and a sum operation that takes a variable of type list as an argument.

[0145] The sample data generation function unit 411 of the training data generation unit 41 identifies parameters that can be varied for each data type and generates sample data by increasing or decreasing those parameters. For example, in the case of the str type, the parameter that can be varied is the length of the string. In the case of the list type, the parameter that can be varied is the number of elements in the array. In other words, the sample data generation function unit 411 generates str type data of various lengths and list type data of various array element counts as sample data for the argument.

[0146] The sample data generation function 411 extracts data from usage examples in various source codes to identify variable parameters, and analyzes the characteristics of each or the documentation. For example, when Python is used as the programming language, the sample data generation function 411 adds up the unit values ​​for each data type and uses the result as the parameter size.

[0147] The execution time measurement function unit 412 of the training data generation unit 41 takes the sample data generated by the sample data generation function unit 411 as an argument, performs the processing, and measures the execution time.

[0148] The machine learning model selection function 421 of the execution time prediction model construction unit 42 analyzes the execution time for each piece of processing information and each piece of data type information, and selects a machine learning model that constructs an equation that closely matches the variation in processing time. For example, the machine learning model selection function 421 selects a machine learning model that corresponds to the equation that is closest in relationship between the size of the sample data and the measurement result of the execution time (i.e., has a high degree of overlap or a high correlation coefficient).

[0149] The machine learning model selection function unit 421 may also select a machine learning model using prediction results based on execution time measured at a different time.

[0150] The machine learning model training function unit 422 of the execution time prediction model construction unit 42 trains a machine learning model using the characteristics and execution time of the input sample data, for each piece of processing information and each piece of data type information. In other words, the machine learning model training function unit 422 constructs a processing execution time prediction model.

[0151] The data structure analysis unit 4411 analyzes the source code of this example shown in Figure 16, extracts the processes and their arguments that appear, and records that information.

[0152] If a constant that is an argument to a process in the source is enclosed in double quotes, the data structure analysis unit 4411 determines that the data is of type str (for example, the second line of Figure 16). Also, if a constant that is an argument to a process in the source is a comma-separated sequence of numbers enclosed in square brackets, the data structure analysis unit 4411 determines that the data is of type list (for example, the fifth line of Figure 16). Furthermore, if the argument to a process in the source is a variable, the data structure analysis unit 4411 searches for the most recent assignment operator "=" that exists before the process and assigns a value to that variable, and identifies the type of data assigned to that variable. For example, for the variable message in the third line of Figure 16, since the right-hand side of the assignment operation in the first line is of type str, the data structure analysis unit 4411 determines that the variable message on the left-hand side is also of type str. Furthermore, for example, regarding the variable set in the sixth row of Figure 16, since the right-hand side of the assignment operation in the fourth row is of type list, the data structure analysis unit 4411 determines that the variable set on the left-hand side is also of type list.

[0153] Furthermore, the data structure analysis unit 4411 analyzes the size of the parameters. In the case of str type data, the data structure analysis unit 4411 uses the length of the string enclosed in double quotes as the size of the parameter. In the case of list type data, the data structure analysis unit 4411 uses the number of numbers (int type) separated by commas within square brackets (the number of elements in the array) as the size of the parameter.

[0154] As a result of analyzing such examples, the data structure analysis unit 4411 generates information about the processing content and the structure of the arguments as analysis results.

[0155] Figure 17 is a schematic diagram showing an example of the analysis results generated by the data structure analysis unit 4411 by analyzing the source code shown in Figure 16. As shown in the figure, these analysis results may be in tabular format representing the relationship between the ID (identification information of the process in the source code), the processing content, and the structure of the argument data.

[0156] In Figure 17, ID=1 corresponds to the processing on the second line of the source code in Figure 16. The information about the processing content is the function "print". The structure of its argument data is the string "I want to say" enclosed in double quotes. In other words, the structure of the argument data is of type str and has a size of 13.

[0157] In Figure 17, ID=2 corresponds to the processing on the third line of the source code in Figure 16. The information about the processing content is the function "print". Furthermore, the structure of its argument data is the string "Hello" assigned to the variable message. In other words, the structure of the argument data is of type str and has a size of 5.

[0158] In Figure 17, ID=3 corresponds to the processing on the 5th line of the source code in Figure 16. The information about the processing content is the function "sum". The structure of its argument data is a sequence of numbers enclosed in square brackets [1,1,1,1,1,1]. In other words, the structure of the argument data is of type list and has a size of 6.

[0159] In Figure 17, ID=4 corresponds to the processing on line 6 of the source code in Figure 16. The information about the processing content is the function "sum". Furthermore, the structure of its argument data is the array [1, 2, 3] assigned to the variable set. In other words, the structure of the argument data is of type list and has a size of 3.

[0160] The data attribute analysis unit 4412 analyzes the source code of this example shown in Figure 16, extracts the processes and their arguments that appear, and records that information. Specifically in this example, the data attribute analysis unit 4412 analyzes whether the attributes of each data are constants or variables.

[0161] Specifically, the data attribute analysis unit 4412 determines that argument data is a constant if it is a str type data enclosed in double quotes, or if it is a comma-separated sequence of numbers enclosed in square brackets. Furthermore, if the argument is a string not enclosed in double quotes in the source code (for example, message or set), it determines that the data is a variable.

[0162] As a result of analyzing such examples, the data attribute analysis unit 4412 generates information about the processing content and the attributes of the arguments as analysis results.

[0163] Figure 18 is a schematic diagram showing an example of the analysis results generated by the data attribute analysis unit 4412 by analyzing the source code shown in Figure 16. As shown in the figure, these analysis results may be in tabular format representing the relationship between ID (identification information of processing in the source code), processing content, and attribute (constant or variable) of argument data.

[0164] In Figure 18, IDs 1, 2, 3, and 4 correspond to IDs 1, 2, 3, and 4 in Figure 17, respectively.

[0165] In Figure 18, ID=1 corresponds to the processing on the second line of the source code in Figure 16. The information about the processing content is the function "print". Furthermore, the attribute of its argument data is the string "I want to say" enclosed in double quotes. In other words, the attribute of this argument data is a constant.

[0166] In Figure 18, ID=2 corresponds to the processing on the third line of the source code in Figure 16. The information about the processing content is the function "print". Furthermore, the attribute of its argument data is the attribute of message. In other words, the attribute of this argument data is a variable.

[0167] In Figure 18, ID=3 corresponds to the processing on the 5th line of the source code in Figure 16. The information about the processing content is the function "sum". Furthermore, the attribute of its argument data is the attribute of a sequence of integers enclosed in square brackets [1,1,1,1,1,1]. In other words, the attribute of this argument data is a constant.

[0168] In Figure 18, ID=4 corresponds to the processing on line 6 of the source code in Figure 16. The information about the processing content is the function "sum". Furthermore, the attribute of its argument data is the attribute of set. In other words, the attribute of this argument data is a variable.

[0169] The data attribute analysis unit 4412 may also analyze the bit sequence of the data to identify the attributes of each piece of data.

[0170] In this example, constants and variables of two data types, str and list, are used as analysis results. However, for other data types as well, the data structure analysis unit 4411 and the data attribute analysis unit 4412 may perform analysis using analysis methods corresponding to each data type.

[0171] The execution time calculation unit 442 of the execution time prediction unit 44 predicts the execution time of the process based on the analysis results (Figure 17) passed from the data structure analysis unit 4411, using a trained processing execution time prediction model.

[0172] Figure 19 is a schematic diagram showing the results of the execution time calculation unit 442 predicting the execution time for each of the IDs 1, 2, 3, and 4 based on the analysis results in Figure 17. Each of the IDs 1, 2, 3, and 4 in Figure 19 corresponds to the IDs 1, 2, 3, and 4 in Figure 17. Furthermore, the processing content and argument data structure shown in Figure 19 are consistent with those in Figure 17.

[0173] When ID=1 and when ID=2, the processing content is "print", and the respective argument data are of type str. Based on this combination of processing content and argument data structure, the execution time calculation unit 442 selects the "print execution time prediction model" as the execution time prediction model. Then, as a result of calculations performed by the execution time calculation unit 442 using the "print execution time prediction model", the predicted execution time for ID=1 is predicted execution time value 1, and the predicted execution time for ID=2 is predicted execution time value 2. Note that predicted execution time value 1 and predicted execution time value 2 are actually obtained as numerical values ​​in, for example, seconds, and this continues to be the case hereafter.

[0174] For ID=3 and ID=4, the processing content is "sum", and the respective argument data are of type list. Based on this combination of processing content and argument data structure, the execution time calculation unit 442 selects the "sum execution time prediction model" as the execution time prediction model. Then, as a result of calculations performed by the execution time calculation unit 442 using the "sum execution time prediction model", the predicted execution time for ID=3 is predicted execution time value 3, and the predicted execution time for ID=4 is predicted execution time value 4.

[0175] The execution time prediction value correction unit 443 of the execution time prediction unit 44 corrects the execution time prediction values ​​for each process obtained by the execution time calculation unit 442. To do this, the execution time prediction value correction unit 443 first determines whether the attributes of the sample data used to train the processing model match the attributes of the actual argument data described in the source code (Figure 16) for each of IDs = 1, 2, 3, and 4. In this example, the attributes of the sample data used to train the processing model are variables.

[0176] Furthermore, the execution time prediction value correction unit 443 determines a correction value according to the result of the data attribute match or mismatch and corrects the execution time prediction value.

[0177] Figure 20 is a schematic diagram showing the correction values ​​obtained by the execution time prediction correction unit 443 and the correction results using those correction values ​​for each of the IDs 1, 2, 3, and 4. The execution time prediction correction unit 443 performs processing based on the data attribute analysis results (Figure 18) received from the data attribute analysis unit 4412. In Figure 20, each of the IDs 1, 2, 3, and 4 corresponds to the IDs 1, 2, 3, and 4 in Figures 18 and 19. Also, the processing content and argument data attributes shown in Figure 20 are the same as those in Figure 18.

[0178] In the example shown in Figure 20, the attributes of the argument data for ID=1 and ID=3 are constants and do not match the attributes of the sample data used during training. Therefore, the execution time prediction correction unit 443 calculates correction values ​​for ID=1 and ID=3. The correction value for ID=1 is correction value 1, and the correction value for ID=3 is correction value 2. Also, the attributes of the argument data for ID=2 and ID=4 are variables and match the attributes of the sample data used during training. Therefore, the execution time prediction correction unit 443 sets the correction values ​​for ID=2 and ID=4 to 0. A correction value of 0 here represents no correction.

[0179] In the example shown in Figure 20, the execution time prediction correction unit 443 performs the correction by adding a correction value to the execution time prediction value before correction. However, the execution time prediction correction unit 443 may also perform the correction by multiplying the execution time prediction value before correction by a coefficient as the correction value. If the correction value is used as the coefficient for multiplication, setting the correction value to 1.0 means no correction is performed.

[0180] The correction value used by the execution time prediction correction unit 443 may be the difference between the execution time measured by changing the attribute of the sample data (in this example, "variable") (for example, changing it to "constant") and the execution time of the training data.

[0181] Alternatively, if the correction value is to be used as a coefficient for multiplication, the correction value used by the execution time prediction correction unit 443 may be the ratio obtained by dividing the execution time measured by changing the attribute of the sample data (in this example, "variable") (for example, changing it to "constant") by the execution time of the training data.

[0182] Furthermore, the correction value used by the execution time prediction correction unit 443 may be a function value that takes the magnitude of the sample data parameters as an argument.

[0183] Alternatively, the correction value used by the execution time prediction value correction unit 443 may be the predicted value of the prediction model used as a feature.

[0184] In the example shown in Figure 20, the corrected execution time predictions are as follows: Corrected execution time prediction 1 = Uncorrected execution time prediction 1 + Correction value 1 Corrected execution time prediction 2 = Uncorrected execution time prediction 2 + 0 (no correction) Corrected execution time prediction 3 = Uncorrected execution time prediction 3 + Correction value 2 Corrected execution time prediction 4 = Uncorrected execution time prediction 4 + 0 (no correction)

[0185] The aggregation unit 444 sums the corrected predicted execution times of the processes included in the source code.

[0186] Furthermore, for example, arithmetic operations such as addition ("+") may have their overall execution time reduced by pipeline processing in the CPU (Central Processing Unit). Considering this, it may be advisable to use the CPU clock time instead of the estimated execution time.

[0187] [Second Processing Example of the Fourth Embodiment] Next, the second processing example of the fourth embodiment will be described. In this example, processing targeting an interpreted language such as Python will be described. In this example as well, the data attribute analysis unit 4412 analyzes the attributes of the data, and the execution time prediction value correction unit 443 corrects the execution time prediction value based on the analysis results. In this example, "str" ​​and "list" are used as data types. These data types "str" ​​and "list" have already been described. The processing included in the source code is the function "print" and the function "sum". These functions "print" and "sum" have also already been described.

[0188] In this example, there are three types of data attributes: "constants," "L variables," and "G variables." L variables are local scope variables. G variables are global scope variables.

[0189] The concepts of local scope and global scope are used, for example, in programming languages ​​with a block structure. Examples of programming languages ​​with a block structure include, but are not limited to, Python, Pascal, ADA, PL / I, Lisp, C, and C++. Local scope variables are defined within a specific range (scope) of a program and are only valid within that scope. Scopes include, for example, the scope of a function definition or the scope of a block. Local scope variables can only be accessed from within their scope. Global scope variables are uniformly accessible from the entire program, regardless of whether they are inside or outside a scope. Local scope variables are also simply called "local variables," and global scope variables are also simply called "global variables."

[0190] Local scope variables and global scope variables store in different memory areas. The memory area for local scope variables needs to be dynamically allocated, for example, when processing enters the scope, whereas the memory area for global scope variables is not dynamically allocated or deallocated during program execution. In other words, when data is a variable, the difference between an L variable and a G variable can cause a difference in processing time. This embodiment can correct the predicted execution time of processing while distinguishing whether the data attribute is an L variable or a G variable.

[0191] The data attribute analysis unit 4412 of this embodiment can distinguish whether the data to be processed (arguments, etc.) is a constant, an L variable, or a G variable by analyzing the source code.

[0192] In other words, in the second processing example, the data attribute is an attribute that indicates at least whether the data (argument) is a constant or a variable, and further distinguishes the attributes of that variable in more detail. That is, when the data (argument) is a variable, the data attribute is an attribute that indicates whether the variable is a local scope variable that is valid only within a predetermined scope of the source code, or a global scope variable that is valid outside the predetermined scope as well.

[0193] The source code used for predicting execution time in this example is the source code shown in Figure 16.

[0194] Both the "print" and "sum" operations can process constant data, L variable data, and G variable data, respectively.

[0195] Similar to the first processing example, the sample data generation function unit 411 of the training data generation unit 41 identifies parameters that can be varied for each data type and generates sample data by increasing or decreasing those parameters. In other words, the sample data generation function unit 411 generates str-type data of various lengths and list-type data of various array element counts as sample data for the argument.

[0196] The execution time measurement function unit 412 measures the execution time of a process that uses the sample data generated by the sample data generation function unit 411 as an argument to the variable L.

[0197] Similar to the first processing example, the machine learning model selection function unit 421 analyzes the execution time for each piece of processing information and each piece of data type information, and selects a machine learning model that constructs an equation close to the variation in execution time.

[0198] Similar to the first processing example, the machine learning model training unit 422 trains a machine learning model using the characteristics and execution time of the input sample data for each processing information and each data type information. This constructs an execution time prediction model.

[0199] The data structure analysis unit 4411 analyzes the source code read from the source code storage unit 13 to extract the processing content and the structure of the data (arguments) targeted by that processing. As in the first processing example, the analysis results output by the data structure analysis unit 4411 are as explained in Figure 17.

[0200] The data attribute analysis unit 4412 analyzes the source code read from the source code storage unit 13 to extract the processing content and the attributes of the data (arguments) targeted by that processing. In this example, the attributes are constants, L variables, or G variables, as described above. Various methods can be considered for the data attribute analysis unit 4412 to determine the data attributes based on the source code.

[0201] For example, the data attribute analysis unit 4412 determines whether the data is a constant or a variable using the method described in the first processing example. Then, for data determined to be a variable, the data attribute analysis unit 4412 checks, based on the variable name, whether the variable is defined (variable declaration, etc.) within the scope (for example, within a function). If the variable is defined within the scope, the data attribute analysis unit 4412 determines that the variable is an L variable. If the variable is not defined within the scope, the data attribute analysis unit 4412 determines that the variable is a G variable.

[0202] Alternatively, the data attribute analysis unit 4412 can use the `compile` function. In the Python language, the `compile` function is provided as part of the standard functionality. The data attribute analysis unit 4412 uses the `compile` function to convert the source code written in Python into the system's internal representation. Then, the data attribute analysis unit 4412 analyzes the obtained internal representation to determine how the symmetric data is loaded, thereby determining whether a particular variable is an L variable or a G variable.

[0203] Figure 21 is a schematic diagram showing an example of the analysis results generated by the data attribute analysis unit 4412 by analyzing the source code shown in Figure 16. However, in this example, there are three types of data attributes: constants, L variables, and G variables. As shown in the figure, these analysis results may be in tabular format, representing the relationship between the ID (identification information of the process in the source code), the processing content, and the attributes of the argument data (constants, L variables, or G variables).

[0204] In Figure 21, IDs 1, 2, 3, and 4 correspond to the operations of "print" on the second line, "print" on the third line, "sum" on the fifth line, and "sum" on the sixth line in the source code of Figure 16.

[0205] In Figure 21, the row with ID=1 has the operation "print". Furthermore, the attribute of its argument data is the string "I want to say" enclosed in double quotes. In other words, the attribute of this argument data is a constant.

[0206] In Figure 21, the row with ID = 2 has the processing content "print". Furthermore, the attribute of its argument data is the attribute of "message". In other words, the data attribute analysis unit 4412 determined that the attribute of this argument data is a G constant.

[0207] In Figure 21, the row with ID = 3 has the operation "sum". Furthermore, the attribute of its argument data is the attribute of a sequence of integers enclosed in square brackets: [1, 1, 1, 1, 1, 1]. In other words, the attribute of this argument data is a constant.

[0208] In Figure 18, the row with ID = 4 has the processing content "sum". Furthermore, the attribute of its argument data is the attribute of set. Therefore, the data attribute analysis unit 4412 determined that the attribute of this argument data is a G constant.

[0209] The execution time calculation unit 442 of the execution time prediction unit 44 predicts the execution time of the process based on the analysis results (the same analysis results as in Figure 17) passed from the data structure analysis unit 4411, using a trained processing execution time prediction model.

[0210] Similar to the first processing example, in the second processing example, the execution time calculation unit 442 predicts the execution time for each of the IDs 1, 2, 3, and 4. The results of the execution time calculation unit 442's prediction of the execution time are the same as those shown in Figure 19, which was explained in the first processing example.

[0211] In other words, the execution time calculation unit 442 calculates a predicted execution time for the process with ID=1 using the print execution time prediction model, and the result is predicted execution time 1. The execution time calculation unit 442 also calculates a predicted execution time for the process with ID=2 using the print execution time prediction model, and the result is predicted execution time 2. The execution time calculation unit 442 also calculates a predicted execution time for the process with ID=3 using the sum execution time prediction model, and the result is predicted execution time 3. The execution time calculation unit 442 also calculates a predicted execution time for the process with ID=4 using the sum execution time prediction model, and the result is predicted execution time 4.

[0212] The execution time prediction value correction unit 443 of the execution time prediction unit 44 corrects the execution time prediction values ​​for each process obtained by the execution time calculation unit 442. To do this, the execution time prediction value correction unit 443 first determines whether the attributes of the sample data used to train the processing model match the attributes of the actual argument data in the analysis results of the data attribute analysis unit 4412 (Figure 21) for each of IDs = 1, 2, 3, and 4. In this example, the attributes of the sample data used to train the processing model are L variables.

[0213] Figure 22 is a schematic diagram showing a list of the correction values ​​used by the execution time prediction correction unit 443 and the corrected execution time prediction values. In other words, Figure 22 is a schematic diagram showing a list of the correction values ​​obtained by the execution time prediction correction unit 443 for each process in the source code based on the analysis results in Figure 21, and the corrected execution time prediction values ​​using those correction values. Each of the IDs = 1, 2, 3, and 4 in Figure 22 corresponds to the IDs = 1, 2, 3, and 4 in Figure 21. That is, as follows:

[0214] For the print process with ID=1, the attribute of the argument data identified by the data attribute analysis unit 4412 was a constant. The execution time prediction correction unit 443 corrects the execution time prediction value 1 using the constant correction value 1. The constant correction value 1 is a correction value for when the attribute of the data to be processed is a constant, based on the case where the attribute of the data to be processed is an L variable (attribute of the data used to train the machine learning model). The method by which the execution time prediction correction unit 443 determines the correction value is as already described. The constant correction value 1 is a correction value when the process is print.

[0215] Regarding the print process with ID=2, the attribute of the argument data identified by the data attribute analysis unit 4412 was a G variable. The execution time prediction correction unit 443 corrects the execution time prediction value 2 using the correction value 1 of the G variable. The correction value 1 of the G variable is the correction value when the attribute of the data to be processed is a G variable, based on the case where the attribute of the data to be processed is an L variable (attributes of the data used to train the machine learning model). The correction value 1 of the G variable is the correction value when the process is print.

[0216] For the sum process of ID=3, the attribute of the argument data identified by the data attribute analysis unit 4412 was constant. The execution time prediction correction unit 443 corrects the execution time prediction value 3 using the constant correction value 2. The constant correction value 2 is the correction value when the attribute of the data to be processed is constant, based on the case where the attribute of the data to be processed is L variable (attribute of the data used to train the machine learning model). The constant correction value 3 is the correction value when the process is sum.

[0217] For the sum process of ID=4, the attribute of the argument data identified by the data attribute analysis unit 4412 was a G variable. The execution time prediction correction unit 443 corrects the execution time prediction value 4 using the correction value 2 of the G variable. The correction value 2 of the G variable is the correction value when the attribute of the data to be processed is a G variable, based on the case when the attribute of the data to be processed is an L variable (attributes of the data used to train the machine learning model). The correction value 2 of the G variable is the correction value when the process is sum.

[0218] Specifically, the execution time prediction value correction unit 443 performs the correction calculation as follows: Corrected execution time prediction value 1 = Predicted execution time prediction value 1 before correction + Constant correction value 1 Corrected execution time prediction value 2 = Predicted execution time prediction value 2 before correction + G variable correction value 1 Corrected execution time prediction value 3 = Predicted execution time prediction value 3 before correction + Constant correction value 2 Corrected execution time prediction value 4 = Predicted execution time prediction value 4 before correction + G variable correction value 2

[0219] In the second processing example, the method of calculating the corrected execution time prediction is not limited to adding the correction value to the uncorrected execution time prediction; other methods may also be used. For example, the corrected execution time prediction may be calculated by multiplying the uncorrected execution time prediction by the correction value (coefficient).

[0220] The aggregation unit 444 sums the corrected predicted execution times of the processes included in the source code.

[0221] As explained above, in the fourth embodiment, the execution time prediction can be corrected according to the attributes of the data targeted by the processing described in the source code. In other words, a more accurate execution time prediction can be obtained.

[0222] In other words, the execution time prediction device of the fourth embodiment can predict the execution time for each process using an appropriate machine learning model, taking into account the differences in processing content and the data to be processed (arguments, etc.) (including differences in data attributes) for a new program. In other words, it can improve the accuracy of execution time prediction. Highly accurate execution time prediction allows for efficient resource selection and scaling in, for example, a cloud environment, thereby improving service quality.

[0223] The fifth and sixth embodiments, which will be described next, may be considered as variations of the fourth embodiment described above.

[0224] [Fifth Embodiment] Next, a fifth embodiment of the present invention will be described. Note that matters already described in the previous embodiments may be omitted below. Here, the focus will be on matters specific to this embodiment.

[0225] Figure 23 is a block diagram illustrating the schematic functional configuration of the execution time prediction device according to the fifth embodiment. As shown in the figure, the execution time prediction device 5 includes a training data generation unit 41, an execution time prediction model construction unit 42, a source code storage unit 13, an execution time prediction unit 54, a prediction result output unit 15, and a correction value calculation unit 56. The execution time prediction unit 54 also includes a source code analysis unit 441, an execution time calculation unit 442, an execution time prediction value correction unit 543, and an aggregation unit 444. The functions of the execution time prediction device 5 can be realized, for example, by a computer and a program, as in the previous embodiments.

[0226] A key feature of this embodiment is that the execution time prediction device 5 has a correction value calculation unit 56. The correction value calculation unit 56 calculates a correction value and passes this correction value to the execution time prediction value correction unit 543. The execution time prediction value correction unit 543 uses the correction value passed from the correction value calculation unit 56 to correct the execution time prediction value before correction and passes the result to the aggregation unit 444.

[0227] Specifically, the correction value calculation unit 56 generates sample data with data attributes different from those of the sample data for each data type, executes processing using the modified sample data, and measures the execution time. The correction value calculation unit 56 then calculates the difference in execution time between executing the processing with the sample data's data attributes and executing the processing with the modified data attributes, and calculates a correction value for each data type and parameter size.

[0228] In other words, if the data attribute of the sample data is a variable and there is a constant as a data attribute other than a variable, the correction value calculation unit 56 calculates a correction value for the data attribute "constant" for each process and each data type using the method described above (see also Figure 20).

[0229] Furthermore, if the data attribute of the sample data is an L variable, and there are constants and G variables as data attributes other than this L variable, the correction value calculation unit 56 calculates correction values ​​for each of the data attributes "constant" and "G variable" for each process and data type using the method described above (see also Figure 22).

[0230] In other words, the correction value calculation unit 56 executes the process using sample data with modified data attributes that are different from the data attributes of the training data used by the machine learning model training function unit when training the machine learning model, and measures the execution time. The correction value calculation unit 56 finds the difference in execution time between when the process is executed with the original training data and when the process is executed with the sample data with modified data attributes. The correction value calculation unit 56 calculates a correction value based on this difference.

[0231] The execution time prediction value correction unit 543 of the execution time prediction unit 54 corrects the execution time of the process using the correction value calculated by the correction value calculation unit 56.

[0232] Other aspects of the fifth embodiment are the same as those of the fourth embodiment.

[0233] As explained above, the fifth embodiment, like the fourth embodiment, can correct the execution time prediction value according to the attributes of the data targeted by the processing described in the source code. In other words, it is possible to obtain a more accurate execution time prediction value.

[0234] [Sixth Embodiment] Next, a fifth embodiment of the present invention will be described. Note that matters already described in the previous embodiments may be omitted below. Here, the focus will be on matters specific to this embodiment.

[0235] Figure 24 is a block diagram illustrating the schematic functional configuration of the execution time prediction device according to the sixth embodiment. As shown in the figure, the execution time prediction device 6 includes a training data generation unit 41, an execution time prediction model construction unit 42, a source code storage unit 13, an execution time prediction unit 64, and a prediction result output unit 15. The execution time prediction unit 64 also includes a source code analysis unit 641, an execution time calculation unit 442, an execution time prediction value correction unit 643, and an aggregation unit 444. The source code analysis unit 641 includes a data structure analysis unit 4411 and a data attribute analysis unit 6412. The functions of the execution time prediction device 6 can be implemented, for example, by a computer and a program, as in the previous embodiments.

[0236] A key feature of this embodiment is that the execution time prediction unit 64 includes a correction value calculation unit 66. In this embodiment, the data attribute analysis unit 6412 analyzes the source code and passes information about the attributes of the data to be processed (see Figures 18 and 21) to the correction value calculation unit 66. The correction value calculation unit 66 calculates a correction value and passes this correction value to the execution time prediction value correction unit 643. The execution time prediction value correction unit 643 uses the correction value passed from the correction value calculation unit 66 to correct the execution time prediction value before correction and passes the result to the aggregation unit 444.

[0237] Specifically, the correction value calculation unit 66 receives analysis result information from the data attribute analysis unit 6412. This analysis result information records the attributes of the data (arguments, etc.) targeted by each process included in the source code. Based on the information included in the source code analysis results, the correction value calculation unit 66 determines whether the attributes of each data match the attributes of the sample data. If the attributes of the data to be processed match the attributes of the sample data, no correction is required for the attributes of the data to be processed (the correction value is 0). If the attributes of the data to be processed do not match the attributes of the sample data, the correction value calculation unit 66 generates data attributes different from the data attributes of the sample data for each data type, executes the process using the modified sample data, and measures the execution time. The correction value calculation unit 56 finds the difference in execution time between executing the process with the sample data with the original data attributes (original training data) and executing the process with the modified data attributes, and calculates a correction value for each data type and parameter size based on this difference.

[0238] The correction value calculation unit 66 passes the correction value data (including cases where no correction is applied) to the execution time prediction value correction unit 643. The execution time prediction value correction unit 643 uses the correction value passed from the correction value calculation unit 66 to correct the execution time prediction value before correction, and passes the result to the aggregation unit 444.

[0239] In other words, the correction value calculation unit 66 receives the data attribute information of the argument output by the data attribute analysis unit 6412 of the source code analysis unit 641, and calculates the correction value based on the data attribute. The correction value calculation unit 66 executes the process using sample data with modified data attributes that are different from the data attributes of the training data used by the machine learning model training function unit when training the machine learning model, and measures the execution time. The correction value calculation unit 66 finds the difference in execution time between executing the process with the original training data and executing the process with the sample data with modified data attributes. The correction value calculation unit 66 calculates the correction value based on this difference.

[0240] The execution time prediction value correction unit 643 of the execution time prediction unit 64 corrects the execution time of the process using the correction value calculated by the correction value calculation unit 66.

[0241] Other aspects of the sixth embodiment are the same as those of the fourth embodiment.

[0242] As explained above, the sixth embodiment, like the fourth embodiment, can correct the execution time prediction value according to the attributes of the data targeted by the processing described in the source code. In other words, it is possible to obtain a more accurate execution time prediction value.

[0243] Although several embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments and includes designs and the like that do not depart from the spirit of this invention.

[0244] The present invention can be used, for example, in services that plan and provide computing resources such as so-called cloud services. However, the scope of use of the present invention is not limited to those exemplified herein.

[0245] 1, 2, 3, 4, 5, 6 Execution Time Prediction Device 11 Training Data Generation Unit 12 Execution Time Prediction Model Construction Unit 13 Source Code Storage Unit 14 Execution Time Prediction Unit 15 Prediction Result Output Unit 22 Execution Time Prediction Model Construction Unit 24 Execution Time Prediction Unit 31 Training Data Generation Unit 32 Execution Time Prediction Model Construction Unit 34 Execution Time Prediction Unit 41 Training Data Generation Unit 42 Execution Time Prediction Model Construction Unit 44 Execution Time Prediction Unit 54 Execution Time Prediction Unit 56 Correction Value Calculation Unit 64 Execution Time Prediction Unit 111 Sample Data Generation Function Unit 112 Execution Time Measurement Function Unit 121 Machine Learning Model Selection Function Unit 122 Machine Learning Model Training Function Unit 141 Source Code Analysis Unit 142 Execution Time Calculation Unit 221 Machine Learning Model Selection Function Unit 222 Machine Learning Model Training Function Unit 241 Source Code Analysis Unit 242 Execution Time Calculation Unit 311 Sample data generation function unit 312 Execution time measurement function unit 321 Machine learning model selection function unit 322 Machine learning model training function unit 341 Source code analysis unit 342 Execution time calculation unit 411 Sample data generation function unit 412 Execution time measurement function unit 421 Machine learning model selection function unit 422 Machine learning model training function unit 441 Source code analysis unit 442 Execution time calculation unit 443 Execution time prediction value correction unit 444 Aggregation unit 543 Execution time prediction value correction unit 641 Source code analysis unit 643 Execution time prediction value correction unit 901 Central processing unit 902 RAM 903 Input / output ports 904, 905 Input / output devices 906 Bus 4411 Data structure analysis unit 4412 Data attribute analysis unit 6412 Data attribute analysis unit

Claims

1. An execution time prediction device comprising: a training data generation unit that generates training data which is a set of a type of process, the size of an argument passed to the process, and the execution time when the process is executed based on the size of the argument; a machine learning model selection function unit that selects a machine learning model from a plurality of candidate machine learning models based on the trend of variation in the execution time corresponding to the type of process and the size of the argument in the training data generated by the training data generation unit; and a machine learning model training function unit that constructs an execution time prediction model by training the machine learning model selected by the machine learning model selection function unit using the training data generated by the training data generation unit.

2. An execution time prediction device according to claim 1, further comprising: an execution time prediction unit that inputs the type of processing included in a given program and the size of the arguments passed to the processing to an execution time prediction model constructed by the machine learning model training function unit, thereby predicting the execution time of the processing corresponding to the type of processing and the size of the arguments.

3. The execution time prediction device according to claim 2, wherein the execution time prediction unit includes a source code analysis unit that analyzes the source code of the program and outputs information on the type of processing included in the program and the size of the arguments passed to the processing, and the type of processing and the size of the arguments passed to the processing output by the source code analysis unit are input to the execution time prediction model.

4. The execution time prediction device according to claim 3, wherein the source code analysis unit analyzes the source code and further outputs information on the data attributes of the arguments in the source code, and the execution time prediction unit inputs the type of process and the size of the arguments passed to the process into the execution time prediction model to correct the predicted execution time of the process based on the data attributes output by the source code analysis unit.

5. The execution time prediction device according to claim 4, wherein the data attribute is at least an attribute that indicates whether the argument is a constant or a variable.

6. The execution time prediction device according to claim 5, wherein, when the argument is a variable, the data attribute is an attribute that indicates whether the variable is a local scope variable that is valid only within a predetermined scope of the source code, or a global scope variable that is valid outside the predetermined scope as well.

7. The execution time prediction device according to claim 4, further comprising: a machine learning model training function unit that performs the process using sample data with modified data attributes different from the data attributes of the training data used when training the machine learning model, measures the execution time thereof, determines the difference between the execution time when the process is performed with the original training data and when the process is performed with the sample data with modified data attributes, and calculates a correction value based on the difference, wherein the execution time prediction unit corrects the execution time of the process using the correction value calculated by the correction value calculation unit.

8. The execution time prediction device according to claim 7, wherein the correction value calculation unit receives information on the data attributes of the argument output by the source code analysis unit and calculates the correction value based on said data attributes.

9. An execution time prediction method comprising: a process in which a training data generation unit generates training data which is a set of a type of process, the size of an argument passed to the process, and the execution time when the process is executed based on the size of the argument; a process in which a machine learning model selection function unit selects a machine learning model from a plurality of candidate machine learning models based on the trend of variation in the execution time corresponding to the type of process and the size of the argument in the training data generated by the training data generation unit; and a process in which a machine learning model training function unit constructs an execution time prediction model by training the machine learning model selected by the machine learning model selection function unit using the training data generated by the training data generation unit.

10. A program for causing a computer to execute the execution time prediction method described in claim 9.