HiveSQL Script Parameter Optimization Method, Device, Computer Equipment and Storage Medium
By generating target encoding parameters and training fitness calculation models, and optimizing HiveSQL script parameters in combination with evolutionary algorithms, the problems of low efficiency and unstable optimization of HiveSQL script parameters in the existing technology are solved, and more efficient and stable parameter optimization is achieved.
Patent Information
- Application Number
- CN202210936917.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-05
AI Technical Summary
Existing HiveSQL script parameter optimization and verification methods are inefficient and unstable.
By obtaining HiveSQL scripts, execution plan and execution time, the target encoding parameters and target training sample data are generated using preset encoding methods, the random forest fitness calculation model is trained, and the script parameters are optimized using evolutionary algorithms.
It significantly reduces the system CPU resource consumption, avoids the execution time of HiveSQL scripts, improves parameter optimization efficiency, and avoids the instability of manual empirical optimization.
Smart Images

Figure CN115269640B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and particularly to a method, apparatus, computer device and storage medium for optimizing HiveSQL script parameters. Background Art
[0002] In the field of big data technology, Hive is a data warehouse processing tool with Hadoop encapsulated at the bottom layer, and uses the HiveSQL scripting language similar to SQL scripting language to implement data query.
[0003] Hive comes with a large number of optimization parameters. By adjusting these optimization parameters, the execution results of HiveSQL can be affected. However, big data developers need to set reasonable optimization parameters based on experience to achieve the optimization effect, resulting in low efficiency and instability in the existing methods for optimizing and validating HiveSQL script parameters. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, computer device and storage medium for optimizing HiveSQL script parameters to solve the problems of low efficiency and instability in the existing methods for optimizing and validating HiveSQL script parameters.
[0005] In the first aspect of this application, a method for optimizing HiveSQL script parameters is provided, including:
[0006] Obtain the HiveSQL script of the target system, as well as the execution plan and execution time corresponding to the HiveSQL script;
[0007] Obtain the target encoding parameter, where the target encoding parameter is obtained by encoding the first target parameter using a preset first encoding method, and the first target parameter is obtained from the HiveSQL script;
[0008] Obtain the target training sample data, where the target training sample data is obtained by encoding the first target parameter, the execution plan and the execution time of the HiveSQL script using a preset second encoding method;
[0009] Obtain the target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data. The preset fitness calculation model is constructed using the random forest algorithm. The preset fitness calculation model receives the target training sample data and outputs the fitness data corresponding to the target training sample data;
[0010] Input the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model, where the optimized target HiveSQL script is obtained by optimizing the second target parameter of the target HiveSQL script using a preset evolutionary algorithm;
[0011] Obtain the parameter optimization result of the target HiveSQL script, where the parameter optimization result is the second target parameter corresponding to the optimized fitness data within the first preset number of evolutionary generations and within the preset fitness change range.
[0012] In the second aspect of this application, a device for optimizing HiveSQL script parameters is provided, including:
[0013] A first data acquisition module, configured to acquire the HiveSQL script of the target system, as well as the execution plan and execution time corresponding to the HiveSQL script;
[0014] A target encoding parameter module, configured to acquire target encoding parameters, where the target encoding parameters are obtained by encoding the first target parameter using a preset first encoding method, and the first target parameter is obtained from the HiveSQL script;
[0015] A target training sample data module, configured to acquire target training sample data, where the target training sample data is obtained by encoding the first target parameter, the execution plan, and the execution time of the HiveSQL script using a preset second encoding method;
[0016] A target fitness calculation model module, configured to acquire a target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data, the preset fitness calculation model is constructed using a random forest algorithm, the preset fitness calculation model receives the target training sample data, and outputs the fitness data corresponding to the target training sample data;
[0017] An optimized fitness data module, configured to input the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model, where the optimized target HiveSQL script is obtained by optimizing the second target parameter of the target HiveSQL script using a preset evolutionary algorithm;
[0018] A parameter optimization result module is used to obtain the parameter optimization result of the target HiveSQL script, where the parameter optimization result is the second target parameter corresponding to the optimized fitness data within the first preset evolution algebra and with a change within the preset fitness change range.
[0019] In a third aspect of the present application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned HiveSQL script parameter optimization method are implemented.
[0020] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned HiveSQL script parameter optimization method are implemented.
[0021] For the above-mentioned HiveSQL script parameter optimization method, device, computer device, and storage medium, by obtaining the HiveSQL script of the target system, as well as the execution plan and execution time of the HiveSQL script, and encoding the parameters used for optimization in the HiveSQL script to obtain target encoded parameters, then encoding the target encoded parameters, the execution plan, and the execution time into target training sample data, and using the target training sample data to train a preset fitness calculation model to obtain a target fitness calculation model. During the process of optimizing the parameters to be optimized of the target HiveSQL script using an evolutionary algorithm, the target fitness calculation model is used to calculate the fitness, and the optimized parameters of the final target HiveSQL script are obtained according to the preset fitness change rule. Using the target fitness calculation model to calculate the fitness not only significantly reduces the consumption of system CPU resources but also avoids the execution time-consuming of the HiveSQL script, improving the parameter optimization efficiency of the HiveSQL script. At the same time, using an evolutionary algorithm to optimize the parameters of the HiveSQL script also avoids the instability of optimization relying on artificial experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic diagram of an application environment of the HiveSQL script parameter optimization method in an embodiment of the present application;
[0024] Figure 2 is a flowchart of a method for optimizing HiveSQL script parameters in an embodiment of the present application;
[0025] Figure 3 is a schematic structural diagram of a device for optimizing HiveSQL script parameters in an embodiment of the present application;
[0026] Figure 4 is a schematic diagram of a computer device in an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0028] The method for optimizing HiveSQL script parameters provided by the present application can be applied in an application environment such as Figure 1 . Among them, the computer device can be, but is not limited to, various personal computers and laptop computers. The computer device can also be a server, which can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. It can be understood that Figure 1 the number of computer devices in
[0029] is only illustrative and can be expanded to any number according to actual needs. Figure 2 In one embodiment, as shown in Figure 1 , a method for optimizing HiveSQL script parameters is provided. Taking the method applied to the computer device in
[0030] as an example, the method includes the following steps S101 to S106:
[0031] S101. Obtain the HiveSQL script of the target system, as well as the execution plan and execution time corresponding to the HiveSQL script.
[0031] Among them, the HiveSQL script contains optimization parameters to be optimized that can be used to optimize and improve the execution efficiency of the HiveSQL script. The goal of this embodiment is to improve the execution efficiency of the HiveSQL script by optimizing the parameters to be optimized, and at the same time improve the execution efficiency of the optimization process and the stability of the optimization results. The execution plan corresponding to the HiveSQL script includes the detailed execution process of the HiveSQL and is encoded and represented by preset syntax rules. The execution plan includes the abstract syntax book corresponding to the HiveSQL script, the dependency relationship between different stages of the execution plan, and the detailed description of each stage.
[0032] S102. Obtain target encoding parameters, where the target encoding parameters are obtained by encoding a first target parameter using a preset first encoding method, and the first target parameter is obtained from the HiveSQL script.
[0033] Specifically, the encoding of the first target parameter using the preset first encoding method includes: First, obtain the first target parameter in the HiveSQL script, and the type of the first target parameter includes a boolean type and a non-boolean type. Then, convert the boolean-type first target parameter to the numerical value 0 or the numerical value 1, where true corresponds to the numerical value 1 and false corresponds to the numerical value 0. Finally, convert the non-boolean-type first target parameter into parameter values with a preset number of parameters, where the parameter values are within a preset target parameter range. For example, convert the true or false value corresponding to the parameter named "hive.exec.parallel" into the corresponding 1 or 0 value. Encoding the first target parameter obtained from the HiveSQL script using the preset first encoding method will further reduce the computational difficulty of the first target parameter and improve the running efficiency of the HiveSQL script parameter optimization method.
[0034] S103. Obtain target training sample data, where the target training sample data is obtained by encoding the first target parameter, the execution plan, and the execution time of the HiveSQL script using a preset second encoding method.
[0035] Specifically, first, convert the execution plan corresponding to the HiveSQL script into a corresponding sample execution vector. Then, convert the first target parameter in the HiveSQL script into a sample coding parameter. Secondly, obtain the execution time reduction rate of the HiveSQL script, where the execution time reduction rate is the reduction rate of the execution time of the HiveSQL script relative to the time-consuming of the HiveSQL script executed without using the optimization parameters provided by the Hive data warehouse tool. Finally, use the sample execution vector and the sample coding parameter as the dimensional data in the target training sample data, and use the execution time reduction rate as the label data in the target training sample data.
[0036] S104. Obtain a target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data. The preset fitness calculation model is constructed using the random forest algorithm. The preset fitness calculation model receives the target training sample data and outputs the fitness data corresponding to the target training sample data.
[0037] Among them, the preset fitness calculation model includes a fitness calculation formula, and the fitness calculation formula is as follows:
[0038] f n = 1 - (T n / T0)
[0039] Among them, f n is the fitness, T0 represents the execution time-consuming of the HiveSQL script to be calculated without using the optimization parameters provided by the Hive data warehouse tool, and T n represents the execution time-consuming of the HiveSQL script to be calculated after parameter optimization.
[0040] For example, for a certain HiveSQL script to be calculated, T0 = 1200000ms, and T n = 9000000ms. Then, the fitness of the HiveSQL script to be calculated is 25%, which also means that the execution time of the HiveSQL script to be calculated is optimized by 25%. Further, when the parameters in the HiveSQL script are optimized and cause the HiveSQL script to be unable to execute, assign -1 to the fitness f n to handle the situation where the HiveSQL script cannot be executed during the parameter optimization process of the HiveSQL script using the evolutionary algorithm subsequently.
[0041] S105. Input the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model. Herein, the optimized target HiveSQL script is obtained by optimizing the second target parameter of the target HiveSQL script using a preset evolutionary algorithm.
[0042] Further, the evolutionary algorithm uses a genetic algorithm. The genetic algorithm is a search algorithm for obtaining the optimal solution, which converts the problem-solving process into processes such as replication, crossover, and mutation of chromosome genes in biological evolution. When solving relatively complex combinatorial optimization problems, compared with conventional optimization algorithms, it can obtain better optimization results more quickly. The genetic algorithm has been widely applied in fields such as combinatorial optimization, machine learning, signal processing, adaptive control, and artificial life. The genetic operation operators of the genetic algorithm include a replication operator, a crossover operator, and a mutation operator. Among them, the replication operator, the crossover operator, and the mutation operator are used to perform corresponding replication operations, and / or crossover operations, and / or mutation operations on the second target parameter of the target HiveSQL script to obtain the optimized second target parameter. Replace the original second target parameter in the target HiveSQL script with the optimized second target parameter, and recalculate the fitness data of the target HiveSQL script through the target fitness calculation model to obtain the optimized fitness data.
[0043] S106. Obtain the parameter optimization result of the target HiveSQL script. Herein, the parameter optimization result is the second target parameter corresponding to the optimized fitness data within the first preset evolutionary generation and within the preset fitness change range.
[0044] Specifically, first determine whether the evolutionary generation of the evolutionary algorithm has reached the first preset evolutionary generation. If the evolutionary generation of the evolutionary algorithm has reached the first preset evolutionary generation, then further determine whether the change in the optimized fitness data within the first preset evolutionary generation from the current evolutionary generation forward is within the preset fitness change range. If the change in the optimized fitness data is within the preset fitness change range, output the corresponding second target parameter as the parameter optimization result, and the HiveSQL script parameter optimization method ends. If the change in the optimized fitness data is not within the preset fitness change range, then determine whether the evolutionary generation has reached the second preset evolutionary generation. If the evolutionary generation reaches the second preset evolutionary generation, set the evolutionary generation to 0, and return to execute the step of inputting the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model.
[0045] For example, obtain example target parameters for parameter optimization from an example HiveSQL script, and use the genetic algorithm to optimize the example target parameters. During the optimization process of the example target parameters, first determine whether the number of genetic generations for optimizing the example target parameters using the genetic algorithm has reached 50 generations. If the number of genetic generations has reached 50 generations, obtain 50 different fitness data corresponding to the example target parameters during the genetic process from the 1st generation to the 50th generation of the genetic algorithm. If the variation ranges of these 50 different fitness data are all within ±0.05, then use the example target parameters corresponding to the genetic result of the 50th generation of the genetic algorithm as the parameter optimization result. If there is at least one generation of fitness data among these 50 different fitness data whose variation range is not within ±0.05, then the genetic algorithm continues to inherit and optimize the example target parameters, and further determine whether the number of genetic generations of the genetic algorithm has reached the 1000th generation. If the number of genetic generations of the genetic algorithm has not reached the 1000th generation, continue to use the genetic algorithm to optimize the example target parameters in the example HiveSQL script. If the number of genetic generations of the genetic algorithm has reached the 1000th generation, set the number of evolution generations of the genetic algorithm to 0 and restart the technology, and continue to use the genetic algorithm to optimize the example target parameters in the example HiveSQL script.
[0046] Furthermore, the numerical value of the first evolution generation is adjusted according to Figure 1 the system resource situation of the computer device in. During the execution of the HiveSQL script parameter optimization method, when the numerical value of the first evolution generation is the first value, Figure 1 the system resources of the computer device in are severely consumed, then before the next execution of the HiveSQL script parameter optimization method, further reduce the numerical value of the first evolution generation on the basis of the first value, so that the HiveSQL script parameter optimization method consumes Figure 1The system resource occupancy of the computer device therein is further reduced, and at the same time, the number of evolution generations required for the genetic algorithm to evolve to obtain the parameter optimization result can also be reduced. However, the problem with this execution step is that the stability of the parameter optimization result obtained by the genetic algorithm through evolution will decrease because the number of genetic generations for the genetic algorithm to perform stability tests on the parameters to be optimized has been reduced. The executor of this embodiment needs to find a more balanced implementation method between consuming the system resources of the computer device and the stability of the parameter optimization result. On the contrary, if, during the execution of the HiveSQL script parameter optimization method, the numerical value of the first evolution generation is the first value, Figure 1 there is a situation where too much system resource of the computer device therein remains unused. Then, before the next execution of the HiveSQL script parameter optimization method, the numerical value of the first evolution generation is further increased on the basis of the first value, so that the system resource utilization rate of the computer device therein during the execution of the HiveSQL script parameter optimization method is further improved, and at the same time, the number of evolution generations required for the genetic algorithm to evolve to obtain the parameter optimization result is increased, thereby ensuring the stability of the parameter optimization result because the genetic algorithm has performed stability tests on the parameter optimization result through more genetic generations. Figure 1 Furthermore, after obtaining the parameter optimization result of the target HiveSQL script, the following steps are also included: First, obtain the execution plan and execution time of the target HiveSQL script in the target system, where the parameters of the target HiveSQL script use the parameter optimization result. Then, use the target HiveSQL script, the execution plan corresponding to the target HiveSQL script, and the execution time corresponding to the target HiveSQL script to optimize the target fitness calculation model, and then continuously optimize the target fitness calculation model, so that the fitness corresponding to the HiveSQL script output by the target fitness calculation model is more accurate. Further, record the first optimized fitness data corresponding to the parameter optimization result of the target HiveSQL script, and then record the actual fitness data of the target HiveSQL script after using the parameter optimization result. The actual fitness data is obtained according to the foregoing fitness calculation formula, and then optimize the target fitness calculation model according to the deviation value between the actual fitness data and the first optimized fitness data, so that the fitness data output by the target fitness calculation model is more accurate.
[0047]
[0048] Figure 1 Furthermore, according to Figure 1The remaining current system resources on the computer device are obtained, and the corresponding number of parameter optimization threads are established according to the preset number of threads rule and the number of HiveSQL scripts in the target system. The parameter optimization tasks of all HiveSQL scripts in the target system are assigned to the corresponding number of parameter optimization threads, so that all HiveSQL scripts in the target system can be optimized simultaneously, further improving the optimization efficiency of all HiveSQL scripts in the target system.
[0049] The HiveSQL script parameter optimization method provided in this embodiment obtains the HiveSQL script of the target system, as well as the execution plan and execution time of the HiveSQL script, and encodes the parameters used for optimization in the HiveSQL script to obtain the target encoded parameters. Then, the target encoded parameters, the execution plan, and the execution time are re-encoded into target training sample data, and the preset fitness calculation model is trained using the target training sample data to obtain the target fitness calculation model. During the process of optimizing the parameters to be optimized of the target HiveSQL script using the evolutionary algorithm, the fitness is calculated using the target fitness calculation model, and the optimized parameters of the final target HiveSQL script are obtained according to the preset fitness change rule. Calculating the fitness using the target fitness calculation model not only significantly reduces the consumption of system CPU resources, but also avoids the execution time-consuming of the HiveSQL script, improving the parameter optimization efficiency of the HiveSQL script. At the same time, using the evolutionary algorithm to optimize the parameters of the HiveSQL script also avoids the instability of manual experience-based optimization.
[0050] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0051] In one embodiment, a HiveSQL script parameter optimization device 100 is provided, and the HiveSQL script parameter optimization device 100 corresponds one-to-one to the HiveSQL script parameter optimization method in the above embodiment. As Figure 3 shown, the HiveSQL script parameter optimization device 100 includes a first data acquisition module 11, a target encoded parameter module 12, a target training sample data module 13, a target fitness calculation model module 14, an optimized fitness data module 15, and a parameter optimization result module 16. The detailed descriptions of each functional module are as follows:
[0052] The first data acquisition module 11 is used to obtain the HiveSQL script of the target system, as well as the execution plan and execution time corresponding to the HiveSQL script;
[0053] A target encoding parameter module 12, configured to obtain target encoding parameters, where the target encoding parameters are obtained by encoding first target parameters using a preset first encoding method, and the first target parameters are obtained from the HiveSQL script;
[0054] A target training sample data module 13, configured to obtain target training sample data, where the target training sample data is obtained by encoding the first target parameters, the execution plan, and the execution time of the HiveSQL script using a preset second encoding method;
[0055] A target fitness calculation model module 14, configured to obtain a target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data, the preset fitness calculation model is constructed using a random forest algorithm, the preset fitness calculation model receives the target training sample data, and outputs fitness data corresponding to the target training sample data;
[0056] An optimized fitness data module 15, configured to input an optimized target HiveSQL script into the target fitness calculation model, and obtain optimized fitness data of the target HiveSQL script output by the target fitness calculation model, where the optimized target HiveSQL script is obtained by optimizing second target parameters of the target HiveSQL script using a preset evolutionary algorithm;
[0057] A parameter optimization result module 16, configured to obtain a parameter optimization result of the target HiveSQL script, where the parameter optimization result is the second target parameter corresponding to the optimized fitness data within a first preset number of evolutionary generations and within a preset fitness change range.
[0058] Further, the target fitness calculation model module 14 further includes:
[0059] A fitness calculation sub-module, where the preset fitness calculation model includes a fitness calculation formula, and the fitness calculation formula is as follows:
[0060] f n = 1 - (T n / T0)
[0061] where f n is the fitness, T0 represents the execution time of the HiveSQL script to be calculated without using the optimization parameters provided by the Hive data warehouse tool, and T n represents the execution time of the HiveSQL script to be calculated after parameter optimization.
[0062] Further, the target encoding parameter module 12 further includes:
[0063] A first target parameter sub-module, configured to obtain the first target parameter in the HiveSQL script, where the type of the first target parameter includes a boolean type and a non-boolean type;
[0064] A boolean type conversion sub-module, configured to convert the boolean type first target parameter into 0 or 1, where true corresponds to 1 and false corresponds to 0;
[0065] A non-boolean type conversion sub-module, configured to convert the non-boolean type first target parameter into a parameter value with a preset number of parameters, where the parameter value is within a preset target parameter range.
[0066] Further, the target training sample data module 13 further includes:
[0067] A sample execution vector sub-module, configured to convert the execution plan corresponding to the HiveSQL script into a corresponding sample execution vector;
[0068] A sample encoding parameter sub-module, configured to convert the first target parameter in the HiveSQL script into a sample encoding parameter;
[0069] An execution time reduction rate sub-module, configured to obtain the execution time reduction rate of the HiveSQL script, where the execution time reduction rate is the reduction rate of the execution time of the HiveSQL script relative to the time-consuming of the HiveSQL script executed without using the optimization parameters provided by the Hive data warehouse tool;
[0070] A sample training data sub-module, configured to use the sample execution vector and the sample encoding parameter as the dimensional data in the target training sample data, and use the execution time reduction rate as the label data in the target training sample data.
[0071] Further, the parameter optimization result module 16 further includes:
[0072] A first judgment sub-module, configured to judge whether the evolution generation of the evolutionary algorithm reaches the first preset evolution generation;
[0073] A second judgment sub-module, configured to, if so, judge whether the change of the optimization fitness data within the first preset evolution generation from the current evolution generation forward is within the preset fitness change range;
[0074] A result output sub-module, configured to, if so, output the corresponding second target parameter as the parameter optimization result, and end the HiveSQL script parameter optimization method;
[0075] A third judgment sub-module, configured to, if not, judge whether the number of evolution generations reaches a second preset number of evolution generations;
[0076] A first return execution sub-module, configured to, if so, set the number of evolution generations to 0, and return to execute the step of inputting the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model.
[0077] Further, the optimized fitness data module 15 further includes:
[0078] A genetic algorithm sub-module, configured to use a genetic algorithm for the evolutionary algorithm, and the genetic operation operators of the genetic algorithm include a replication operator, a crossover operator, and a mutation operator.
[0079] Further, the parameter optimization result module 16 further includes:
[0080] A second data acquisition sub-module, configured to acquire the execution plan and execution time of the target HiveSQL script in the target system, where the parameters of the target HiveSQL script use the parameter optimization result;
[0081] A second model optimization sub-module, configured to optimize the target fitness calculation model by using the target HiveSQL script, the execution plan corresponding to the target HiveSQL script, and the execution time corresponding to the target HiveSQL script.
[0082] The meanings of "first" and "second" in the above-mentioned modules / units are only used to distinguish different modules / units, and are not used to limit which module / unit has a higher priority or other limiting meanings. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices. The division of modules in this application is only a logical division, and there may be other division methods in actual implementation.
[0083] For the specific limitations of the HiveSQL script parameter optimization device 100, reference can be made to the limitations of the HiveSQL script parameter optimization method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned HiveSQL script parameter optimization device 100 can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0084] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data involved in the HiveSQL script parameter optimization method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a HiveSQL script parameter optimization method.
[0085] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the HiveSQL script parameter optimization method in the above-mentioned embodiment, such as Figure 2 the steps S101 to S106 shown and the extension of other extended and related steps of the method. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit of the HiveSQL script parameter optimization device 100 in the above-mentioned embodiment, such as Figure 3 the functions of module 11 to module 16 shown. To avoid repetition, it will not be elaborated here.
[0086] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and connects various parts of the entire computer device using various interfaces and lines.
[0087] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, video data, etc.).
[0088] The memory can be integrated in the processor or can be separately provided from the processor.
[0089] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps of the HiveSQL script parameter optimization method in the above embodiment, such as Figure 2 the steps S101 to S106 shown and the extensions and related steps of the method. Or, when the computer program is executed by a processor, it realizes the functions of each module / unit of the HiveSQL script parameter optimization device 100 in the above embodiment, such as Figure 3 the functions of the modules 11 to 16 shown. To avoid repetition, it will not be elaborated here.
[0090] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0091] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0092] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for optimizing HiveSQL script parameters, characterized in that, Including: Obtain the Hive SQL script of the target system, as well as the execution plan and execution time corresponding to the Hive SQL script; Obtain the target encoding parameter, where the target encoding parameter is obtained by encoding the first target parameter using a preset first encoding method, and the first target parameter is obtained from the Hive SQL script; Obtain the target training sample data, where the target training sample data is obtained by encoding the target encoding parameter, the execution plan, and the execution time of the Hive SQL script using a preset second encoding method; Obtain the target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data. The preset fitness calculation model is constructed using the random forest algorithm. The preset fitness calculation model receives the target training sample data and outputs the fitness data corresponding to the target training sample data; Input the optimized target Hive SQL script into the target fitness calculation model to obtain the optimized fitness data of the target Hive SQL script output by the target fitness calculation model, where the optimized target Hive SQL script is obtained by optimizing the second target parameter of the target Hive SQL script using a preset evolutionary algorithm; Obtain the parameter optimization result of the target Hive SQL script, where the parameter optimization result is the second target parameter corresponding to the optimized fitness data within the first preset evolutionary generation and within the preset fitness change range; 2. The HiveSQL script parameter optimization method according to claim 1, wherein The preset fitness calculation model includes a fitness calculation formula, and the fitness calculation formula is as follows: fn = 1 - (T n / T0) Among them, f n is the fitness, T0 represents the execution time taken when the HiveSQL script to be calculated does not use the optimization parameters provided by the Hive data warehouse tool, and T n represents the execution time taken when the HiveSQL script to be calculated is optimized with parameters.
3. The HiveSQL script parameter optimization method according to claim 1, wherein The encoding of the first target parameter using the preset first encoding method includes: Obtain the first target parameter in the Hive SQL script, and the type of the first target parameter includes boolean type and non-boolean type; Convert the boolean-type first target parameter to 0 or 1, where true corresponds to 1 and false corresponds to 0; Convert the non-boolean-type first target parameter into parameter values with a preset number of parameters, where the parameter values are within a preset target parameter range; 4. The HiveSQL script parameter optimization method according to claim 1, characterized in that The encoding of the target encoding parameter, the execution plan, and the execution time of the Hive SQL script using the preset second encoding method includes: Convert the execution plan corresponding to the Hive SQL script into a corresponding sample execution vector; Convert the target encoding parameter in the Hive SQL script into a sample encoding parameter; Obtain the execution time reduction rate of the Hive SQL script, where the execution time reduction rate is the reduction rate of the execution time of the Hive SQL script relative to the time-consuming of the Hive SQL script executed without using the optimization parameters provided by the Hive data warehouse tool; Use the sample execution vector and the sample encoding parameter as the dimensional data in the target training sample data, and use the execution time reduction rate as the label data in the target training sample data; Encode the dimension data and the label data into the target training sample data using the preset second encoding method.
5. The HiveSQL script parameter optimization method according to claim 1, characterized in that The obtaining of the parameter optimization result of the target HiveSQL script includes: Determine whether the evolution generation of the evolutionary algorithm reaches the first preset evolution generation; If so, determine whether the change in the optimization fitness data from the current evolution generation to the previous first preset evolution generation is within the preset fitness change range; If so, output the corresponding second target parameter as the parameter optimization result, and the HiveSQL script parameter optimization method ends; If not, determine whether the evolution generation reaches the second preset evolution generation; If so, set the evolution generation to 0, and return to execute the step of inputting the optimized target HiveSQL script into the target fitness calculation model to obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model.
6. The HiveSQL script parameter optimization method according to claim 1, characterized in that The evolutionary algorithm uses a genetic algorithm, and the genetic operation operators of the genetic algorithm include a replication operator, a crossover operator, and a mutation operator.
7. The HiveSQL script parameter optimization method according to claim 1, wherein After obtaining the parameter optimization result of the target HiveSQL script, it further includes: Obtain the execution plan and execution time of the target HiveSQL script in the target system, where the parameters of the target HiveSQL script use the parameter optimization result; Optimize the target fitness calculation model using the target HiveSQL script, the execution plan corresponding to the target HiveSQL script, and the execution time corresponding to the target HiveSQL script.
8. An apparatus for optimizing HiveSQL script parameters, characterized in that, It includes: A first data acquisition module for acquiring the HiveSQL script of the target system, as well as the execution plan and execution time corresponding to the HiveSQL script; A target encoding parameter module for acquiring target encoding parameters, where the target encoding parameters are obtained by encoding the first target parameters using a preset first encoding method, and the first target parameters are obtained from the HiveSQL script; A target training sample data module for acquiring target training sample data, where the target training sample data is obtained by encoding the target encoding parameters, the execution plan, and the execution time of the HiveSQL script using a preset second encoding method; A target fitness calculation model module for acquiring a target fitness calculation model, where the target fitness calculation model is obtained by training a preset fitness calculation model using the target training sample data, the preset fitness calculation model is constructed using a random forest algorithm, the preset fitness calculation model receives the target training sample data, and outputs the fitness data corresponding to the target training sample data; An optimized fitness data module, configured to input the optimized target HiveSQL script into the target fitness calculation model, and obtain the optimized fitness data of the target HiveSQL script output by the target fitness calculation model, where the optimized target HiveSQL script is obtained by optimizing the second target parameter of the target HiveSQL script using a preset evolutionary algorithm; A parameter optimization result module, configured to obtain the parameter optimization result of the target HiveSQL script, where the parameter optimization result is the second target parameter corresponding to the optimized fitness data within the first preset number of evolutionary generations and within a preset fitness change range.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the HiveSQL script parameter optimization method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the HiveSQL script parameter optimization method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-join query optimization method for database based on improved SDD-1 (System for Distributed Database) algorithm
CN102110158A
Hadoop optimal parameter evaluation method and device
CN111858003A