Hive script running exception handling method and device, electronic equipment and storage medium

CN118467498BActive Publication Date: 2026-10-09CHINA TELECOM CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410501670.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2026-10-09
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

又由于Hive程序运行比较耗时,当长脚本或重要脚本凌晨出错时导致时间和人工成本较大,经常耽误业务报表的准时发布

Benefits of technology

[0064] The Hive script execution exception handling method, apparatus, electronic device, and storage medium provided in this application embodiment obtain error script information and runtime environment information when a Hive script execution error occurs. The error script information and runtime environment information are parsed to determine the attribute values ​​of decision attributes. Based on the attribute values ​​of the decision attributes, a decision tree is used to determine whether the Hive script should be rescheduled. When it is determined that the Hive script should be rescheduled, a new temporary script is generated based on the breakpoint statement identifiers in the error script information, and the new temporary script is scheduled for execution. Because the decision to reschedule the erroneous Hive script can be based on the error script information and runtime environment information, instead of rerunning the complete code or blindly re-tuning, the waste of cluster system resources can be avoided. Moreover, by re-tuning when it is determined that re-tuning is possible, the efficiency of fault handling can be improved, and the labor costs of system operation can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118467498B_ABST
    Figure CN118467498B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a Hive script running exception processing method and device, electronic equipment and a storage medium. The method comprises: when an error occurs in Hive script running, obtaining error script information and running environment information; analyzing the error script information and the running environment information to determine attribute values of decision attributes, the decision attributes being attributes used to determine whether to reschedule the Hive script, the decision attributes comprising error types, script information attributes and environment information attributes of the Hive script; determining whether to reschedule the Hive script by means of a decision tree according to the attribute values of the decision attributes; and when it is determined that the Hive script is to be rescheduled, generating a new temporary script according to a breakpoint statement identifier in the error script information, and scheduling and executing the new temporary script. The embodiments of the present application can avoid waste of cluster system resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for handling Hive script execution exceptions. Background Technology

[0002] Hive is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading. It is a mechanism for storing, querying, and analyzing large-scale data stored in Hadoop. Hive can map structured data files to a database table and provide SQL query functionality, transforming SQL statements into MapReduce tasks for execution.

[0003] Using a Hadoop-based data warehouse can essentially achieve the integration of a big data platform. However, due to factors such as the instability of the big data platform environment and data resource updates, network failures, and task scheduling issues, Hive scripts may occasionally experience unexpected interruptions. Furthermore, since Hive programs are relatively time-consuming to run, errors in long or important scripts in the early morning can lead to significant time and labor costs, frequently delaying the timely release of business reports.

[0004] Script task exceptions caused by certain reasons can be resolved by re-invoking the statements at the point of error, allowing the script to continue running. However, rerunning the entire code or blindly re-invoking the erroneous code will result in a significant waste of cluster system resources. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for handling Hive script execution exceptions, which helps to save system resources.

[0006] To address the aforementioned problems, in a first aspect, embodiments of this application provide a method for handling Hive script execution exceptions, including:

[0007] When an error occurs during Hive script execution, obtain the error script information and runtime environment information;

[0008] The error script information and the runtime environment information are parsed to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type, script information attribute, and environment information attribute of the Hive script.

[0009] Based on the attribute values ​​of the decision attributes, a decision tree is used to determine whether to reschedule the Hive script.

[0010] When it is determined that the Hive script needs to be rescheduled, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

[0011] Optionally, the error reporting script information includes: script priority, execution start time, error time, error SQL statement identifier, and script error message;

[0012] The runtime environment information includes: the current environment information of the Hive library tenant and the current task scheduling environment information.

[0013] Optionally, the current environment information of the Hive library tenant includes: memory information and CPU usage information;

[0014] The current task scheduling environment information includes: the number of concurrent script scheduling requests.

[0015] Optionally, the step of parsing the error script information and the runtime environment information to determine the attribute values ​​of the decision attributes includes:

[0016] Obtain keywords from the script error message and determine the error type of the Hive script based on the keywords;

[0017] The error script information is parsed to determine the attribute values ​​corresponding to the script information attributes;

[0018] The runtime environment information is parsed to determine the attribute values ​​corresponding to the environment information attributes.

[0019] Optionally, based on the attribute value of the decision attribute, a decision tree is used to determine whether to reschedule the Hive script, including:

[0020] Based on each decision attribute in the decision tree, the attribute values ​​of the decision attributes are sequentially substituted into the branches corresponding to the respective decision attributes in the decision tree to obtain the classification result of whether or not to reschedule the Hive script.

[0021] Optionally, generating a new temporary script based on the breakpoint statement identifier in the error script information and scheduling the execution of the new temporary script includes:

[0022] Get the number of times the Hive script can resume execution from breakpoints;

[0023] When the number of times the breakpoint can be resumed is less than the threshold, after a preset delay, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

[0024] Optionally, the process of constructing the decision tree includes:

[0025] Obtain the training sample set and the test sample set;

[0026] Based on the training sample set, determine the information gain and information gain ratio of each decision attribute;

[0027] The decision attribute to be used as the root node is determined based on the information gain and information gain ratio of each decision attribute.

[0028] Based on the attribute value of the decision attribute of the root node, the training sample set is divided into multiple subsets, and the decision attribute of the child node of the root node is determined according to each subset. The decision attribute of the child node of the root node is determined recursively until the initial decision tree is constructed.

[0029] Based on the test sample set, the initial decision tree is pruned according to the sample error rate to obtain the final decision tree.

[0030] Optionally, the step of pruning the initial decision tree according to the sample error rate based on the test sample set to obtain the decision tree includes:

[0031] The initial decision tree is pruned from bottom to top to determine the category corresponding to making the current node a new leaf node. The current node is a non-leaf node in the initial decision tree, and the category includes rescheduling or not rescheduling.

[0032] Based on the test sample set, determine the first sample error rate when the current node is used as the new leaf node, and determine the second sample error rate when the current node is not used as the new leaf node;

[0033] When the error rate of the first sample is less than the error rate of the second sample, the subtree rooted at the current node is deleted, and the current node is used as the new leaf node.

[0034] Secondly, embodiments of this application provide a Hive script execution exception handling device, including:

[0035] The script error and environment information collection module is used to obtain error information and runtime environment information when errors occur during Hive script execution;

[0036] The script error and environment information parsing module is used to parse the error script information and the runtime environment information to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type, script information attribute, and environment information attribute of the Hive script.

[0037] The scheduling decision module is used to determine whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree;

[0038] The breakpoint task scheduling module is used to generate a new temporary script based on the breakpoint statement identifier in the error script information when it is determined that the Hive script needs to be rescheduled, and to schedule the execution of the new temporary script.

[0039] Optionally, the error reporting script information includes: script priority, execution start time, error time, error SQL statement identifier, and script error message;

[0040] The runtime environment information includes: the current environment information of the Hive library tenant and the current task scheduling environment information.

[0041] Optionally, the current environment information of the Hive library tenant includes: memory information and CPU usage information;

[0042] The current task scheduling environment information includes: the number of concurrent script scheduling requests.

[0043] Optionally, the script error and environment information parsing module includes:

[0044] An error type determination unit is used to obtain keywords from the script error message and determine the error type of the Hive script based on the keywords.

[0045] The error script information parsing unit is used to parse the error script information and determine the attribute value corresponding to the script information attribute.

[0046] The runtime environment information parsing unit is used to parse the runtime environment information and determine the attribute values ​​corresponding to the environment information attributes.

[0047] Optionally, the scheduling decision module is specifically used for:

[0048] Based on each decision attribute in the decision tree, the attribute values ​​of the decision attributes are sequentially substituted into the branches corresponding to the respective decision attributes in the decision tree to obtain the classification result of whether or not to reschedule the Hive script.

[0049] Optionally, the breakpoint task scheduling module is specifically used for:

[0050] Get the number of times the Hive script can resume execution from breakpoints;

[0051] When the number of times the breakpoint can be resumed is less than the threshold, after a preset delay, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

[0052] Optionally, the apparatus further includes a decision tree construction module, the decision tree construction module comprising:

[0053] The sample acquisition unit is used to acquire the training sample set and the test sample set;

[0054] The information gain determination unit is used to determine the information gain and information gain ratio of each decision attribute based on the training sample set.

[0055] The root node determination unit is used to determine the decision attribute as the root node based on the information gain and information gain ratio of each decision attribute.

[0056] The recursive construction unit is used to divide the training sample set into multiple subsets based on the attribute value of the decision attribute of the root node, and to determine the decision attribute of the child node of the root node based on each subset, and recursively execute the determination of the decision attribute of the child node of the next layer until the initial decision tree is constructed.

[0057] The pruning unit is used to prune the initial decision tree according to the sample error rate based on the test sample set to obtain the decision tree.

[0058] Optionally, the pruning unit is specifically used for:

[0059] The initial decision tree is pruned from bottom to top to determine the category corresponding to making the current node a new leaf node. The current node is a non-leaf node in the initial decision tree, and the category includes rescheduling or not rescheduling.

[0060] Based on the test sample set, determine the first sample error rate when the current node is used as the new leaf node, and determine the second sample error rate when the current node is not used as the new leaf node;

[0061] When the error rate of the first sample is less than the error rate of the second sample, the subtree rooted at the current node is deleted, and the current node is used as the new leaf node.

[0062] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the Hive script execution exception handling method described in embodiments of this application.

[0063] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the Hive script execution exception handling method disclosed in embodiments of this application.

[0064] The Hive script execution exception handling method, apparatus, electronic device, and storage medium provided in this application embodiment obtain error script information and runtime environment information when a Hive script execution error occurs. The error script information and runtime environment information are parsed to determine the attribute values ​​of decision attributes. Based on the attribute values ​​of the decision attributes, a decision tree is used to determine whether the Hive script should be rescheduled. When it is determined that the Hive script should be rescheduled, a new temporary script is generated based on the breakpoint statement identifiers in the error script information, and the new temporary script is scheduled for execution. Because the decision to reschedule the erroneous Hive script can be based on the error script information and runtime environment information, instead of rerunning the complete code or blindly re-tuning, the waste of cluster system resources can be avoided. Moreover, by re-tuning when it is determined that re-tuning is possible, the efficiency of fault handling can be improved, and the labor costs of system operation can be reduced. Attached Figure Description

[0065] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 This is a flowchart of a method for handling Hive script execution exceptions provided in an embodiment of this application;

[0067] Figure 2 This is a dependency graph between modules in the embodiments of this application;

[0068] Figure 3 This is a flowchart of the decision tree construction process in the embodiments of this application;

[0069] Figure 4 This is a schematic diagram of the training data structure in the training sample set in the embodiments of this application;

[0070] Figure 5 This is a flowchart of a method for handling Hive script execution exceptions provided in an embodiment of this application;

[0071] Figure 6 This is a schematic diagram of the structure of a Hive script execution exception handling device provided in an embodiment of this application;

[0072] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0073] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0074] Figure 1 This is a flowchart of a Hive script execution exception handling method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes steps 110 to 140.

[0075] Step 110: When an error occurs during the execution of a Hive script, obtain the error script information and runtime environment information.

[0076] The error-reporting script information may include: script priority, start time, error time, error SQL statement identifier, and script error message. The runtime environment information may include: current environment information of the Hive library tenant and current task scheduling environment information. The current environment information of the Hive library tenant may include: memory information and CPU usage information, thereby determining whether memory is sufficient and whether CPU usage is normal. The current task scheduling environment information may include: script scheduling concurrency, thereby determining whether the script scheduling concurrency is sufficient and whether the error occurred at night.

[0077] The execution of Hive scripts can be monitored through a script error and environment information collection module. When errors occur during Hive script execution, it collects error script information and runtime environment information. The collected error script information may include: script priority, start time, error time, error SQL statement identifier, and script error message. Runtime environment information may include the current environment information of the Hive database tenant and the current task scheduling environment information. The current environment information of the Hive database tenant may include: memory information, CPU usage information, etc.; the current task scheduling environment information may include: script scheduling concurrency, etc. The script scheduling concurrency refers to the number of scripts simultaneously sent to the Hive database for execution.

[0078] The script start time and error time in the script error message can be used to indicate whether the script was running at night, a factor that will affect the operation of subsequent decision-making mechanisms. The erroneous SQL statement identifier, i.e., the erroneous SQL statement number, is used to mark the statements in the script that caused errors, facilitating the handling of the script by the breakpoint task scheduling module. The current environment information of the Hive library tenant and the current task scheduling environment information can be used to determine whether a breakpoint resume environment is available. Breakpoint resume refers to the process where, during the execution of a Hive script, if the script is abnormally interrupted due to reasons such as missing source data, syntax errors, or temporary cluster failures, it resumes execution at the point of interruption, according to certain rules, for scripts that are suitable for continuing execution below the interruption point. A script interruption point can also be called a breakpoint.

[0079] Step 120: Parse the error script information and the runtime environment information to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type of the Hive script, script information attribute, and environment information attribute.

[0080] The Hive script error type is used to characterize the type or cause of an error in the Hive script. Examples include syntax errors, data source anomalies, brief network failures, cluster anomalies, runtime memory overflows, source table anomalies, sub-script failures, and table compression errors. Script information attributes can be used to characterize different properties of the script. Environment information attributes can be used to determine whether a breakpoint-resume-enabled environment is available. Environment information attributes can include the current environment information of the Hive library tenant and the current task scheduling environment information. For example, script information attributes can include script priority; the current environment information of the Hive library tenant can include whether the Hive cluster memory is sufficient and whether CPU usage is normal; the current task scheduling environment information can include whether the scheduling concurrency is idle, whether the Hive environment is idle, and whether it is running at night.

[0081] The script error and environment information parsing module can be used to parse error script information and runtime environment information, identify the type of script error and the current runtime environment, and obtain the error type of the Hive script, the attribute values ​​corresponding to the script information attributes, and the attribute values ​​corresponding to the environment information attributes. For example, the script error and environment information parsing module can analyze error script information to determine the error type, such as runtime memory overflow, source table exception, subscript failure, table compression error, etc. Runtime memory overflow refers to Hive statements exceeding the tenant's memory limit; source table exception refers to problems such as tables not existing or not being updated; subscripts are other subroutines nested within a script, such as Python scripts; table compression error refers to errors during Hive table compression.

[0082] In one embodiment of this application, the step of parsing the error script information and the runtime environment information to determine the attribute value of the decision attribute includes: obtaining keywords in the script error message information and determining the error type of the Hive script based on the keywords; parsing the error script information to determine the attribute value corresponding to the script information attribute; and parsing the runtime environment information to determine the attribute value corresponding to the environment information attribute.

[0083] The format of error messages when Hive database scripts encounter errors is relatively fixed. Keywords can be filtered from these messages to determine the error type, such as syntax errors or data source anomalies. The runtime environment information is primarily numerical. Whether the values ​​of these values ​​meet scheduling requirements can help determine actual production needs. For example, runtime environment information can be compared to corresponding thresholds to determine the attribute values. For instance, if CPU usage exceeds the threshold, the attribute "CPU usage is normal?" is determined to be abnormal. Similarly, if the script scheduling concurrency exceeds the threshold, the attribute "Script scheduling concurrency is sufficient?" is determined to be insufficient, or the attribute "Script scheduling concurrency is idle?" is determined to be busy (not idle).

[0084] By analyzing keywords in script error messages, the error type of a Hive script can be accurately determined. Furthermore, by parsing the error script information and runtime environment information, the attribute values ​​corresponding to the script information attributes and environment information attributes can be accurately determined, thus providing a data analysis basis for deciding whether to reschedule the Hive script.

[0085] Step 130: Based on the attribute value of the decision attribute, determine whether to reschedule the Hive script using a decision tree.

[0086] The scheduling decision module can determine whether to reschedule a Hive script. Based on the attribute values ​​of the decision attributes corresponding to the Hive script's execution error, the scheduling decision module traverses the decision tree, classifies the Hive script using the decision tree, and obtains a classification result. The classification result can be either to reschedule the Hive script or not to reschedule it.

[0087] Decision trees can be pre-constructed based on training samples of known categories, and can be built using the C4.5 algorithm. The C4.5 algorithm is a decision tree algorithm that is an improvement on the ID3 algorithm. It has a fast classification speed, high accuracy, and easy-to-understand classification rules. It uses information gain ratio instead of information gain to select attributes, overcoming the defect of the ID3 algorithm that favors attributes with more values ​​when using information gain to select branch attributes.

[0088] In one embodiment of this application, determining whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree includes: substituting the attribute value of the decision attribute into the branch corresponding to the decision attribute in the decision tree according to each decision attribute in the decision tree, to obtain a classification result of whether to reschedule the Hive script.

[0089] In a decision tree, each decision attribute is a node, and different attribute values ​​of the decision attribute are different branches of that node. According to the decision attribute represented by each node in the decision tree, the attribute value of the decision attribute corresponding to the Hive script error is substituted sequentially from the root node of the decision tree into the branch corresponding to the corresponding decision attribute (node) in the decision tree, until the leaf node is reached, to obtain the classification result of whether the Hive script should be rescheduled.

[0090] For example, when the error type is a brief network failure, cluster anomaly, or data source query anomaly, and the runtime environment has the ability to resume execution from breakpoints, the classification result will be the rescheduling of Hive scripts.

[0091] By using a decision tree to determine whether to reschedule a Hive script based on the attribute values ​​of different decision attributes, that is, by making decisions based on different error types and runtime environment information, it is possible to determine whether to reschedule the erroneous script statements, thereby avoiding the waste of resources caused by blindly rescheduling erroneous code.

[0092] Step 140: When it is determined that the Hive script needs to be rescheduled, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

[0093] The breakpoint statement identifier is an identifier for the script statement that malfunctioned, such as an malfunctioning SQL statement identifier.

[0094] When it's determined that a Hive script needs to be rescheduled, the breakpoint task scheduling module can use the breakpoint statement identifier in the error script information to copy the breakpoint statement and the code following it to a new temporary script, creating a new temporary script, and then schedule its execution. If the determined classification result is that the Hive script should not be rescheduled, maintenance personnel can be prompted to investigate.

[0095] In one embodiment of this application, generating a new temporary script based on the breakpoint statement identifier in the error script information and scheduling the execution of the new temporary script includes: obtaining the number of breakpoint resume attempts of the Hive script; when the number of breakpoint resume attempts is less than a threshold, after a preset delay, generating a new temporary script based on the breakpoint statement identifier in the error script information and scheduling the execution of the new temporary script.

[0096] The count threshold is a pre-set total number of times a single script can be resumed from a breakpoint (re-scheduled from the breakpoint), for example, 5 times. This threshold ensures the program terminates normally. The preset duration is the delay required for each resume from a breakpoint, for example, 5 minutes.

[0097] If the decision indicates that the Hive script needs to be rescheduled, then the script is rescheduled at the breakpoint based on the breakpoint statement identifier in the error script information. The rescheduling at the script breakpoint specifically includes the following: 1) Obtaining the number of times the script has resumed execution from the breakpoint; 2) Obtaining the temporary script that previously malfunctioned for the Hive tenant corresponding to the malfunctioning SQL statement; 3) Generating a new temporary script for the previously malfunctioning temporary script, i.e., copying the previously malfunctioning SQL statement and the code following that SQL statement into the new temporary script, and then rescheduling the new temporary script.

[0098] It should be noted that the above rescheduling can be performed when the Hive script type is a non-mandatory script, but not when the Hive script type is a mandatory script.

[0099] By rescheduling the script when the number of times it can resume from a breakpoint is less than the threshold, the waste of resources caused by unlimited rescheduling can be avoided. Furthermore, rescheduling after a preset delay can avoid errors caused by temporary network failures.

[0100] Figure 2 This is a dependency graph between modules in the embodiments of this application, such as... Figure 2As shown, the Hive script execution exception handling method provided in this application embodiment can be implemented through a script error and environment information collection module, a script error and environment information parsing module, a scheduling decision module, and a breakpoint task scheduling module. This method uses the script error and environment information collection module to collect information related to the erroneous script; the script error and environment information parsing module analyzes and judges the causes and types of various errors, such as runtime memory overflow, source table anomalies, sub-script failures, table compression errors, etc.; the scheduling decision module makes decisions based on different error types and the current runtime environment information, determining whether the erroneous script statements should be re-executed; and the breakpoint task scheduling module uses scheduling tools to re-execute statements at breakpoints in the erroneous script.

[0101] The Hive script execution exception handling method provided in this application obtains error script information and runtime environment information when a Hive script encounters an error, parses the error script information and runtime environment information, determines the attribute value of a decision attribute, and determines whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree. When it is determined that the Hive script should be rescheduled, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled for execution. Since the decision to reschedule the erroneous Hive script can be based on the error script information and runtime environment information, instead of rerunning the complete code or blindly re-tuning, the waste of cluster system resources can be avoided. Moreover, by re-tuning when it is determined that it can be re-tuned, the efficiency of fault handling can be improved and the labor cost of system operation can be reduced.

[0102] Figure 3 This is a flowchart of the decision tree construction process in the embodiments of this application, such as... Figure 3 As shown, the process of constructing the decision tree includes the following steps:

[0103] Step 310: Obtain the training sample set and the test sample set.

[0104] The script error and environment information collection module collects script error information and runtime environment information, including the current environment information of the Hive library tenant and the current task scheduling environment information. Based on actual production needs, the module obtains category information (whether rescheduling is required) for labeling script error information and runtime environment information, forming training and testing data for the C4.5 algorithm, resulting in training and testing sample sets.

[0105] Figure 4 This is a schematic diagram of the training data structure in the training sample set in the embodiments of this application, such as... Figure 4As shown, each data point includes attribute values ​​corresponding to each decision attribute and labeled category information. Decision attributes may include: whether it runs at night, priority level, whether the scheduling concurrency is idle, whether the Hive environment is idle, script error type, etc. Labeled category information may include rescheduling or not rescheduling.

[0106] Step 320: Determine the information gain and information gain ratio of each decision attribute based on the training sample set.

[0107] Based on the training sample set, the information entropy of the training sample set and the information entropy of each decision attribute are determined. Then, based on the information entropy of the training sample set and the information entropy of each decision attribute, the information gain and information gain ratio of each decision attribute are determined. Information entropy is used to measure the degree of uncertainty or disorder in the data.

[0108] The information entropy of the training sample set can be determined using the following formula:

[0109]

[0110] Where Entropy(S) represents the information entropy of the training sample set, |S| represents the total number of samples in the training sample set S, and m is the number of categories in the training sample set. In this embodiment, the category information is either rescheduled or not rescheduled, so m is 2, |C i | represents the number of samples in the i-th category.

[0111] The information entropy of each decision attribute is determined using the following formula:

[0112]

[0113] Where Entropy(S,A) represents the information entropy of decision attribute A in the training sample set S, v represents the v-th attribute value corresponding to decision attribute A, and S v Entropy(S) represents a subset of the training sample set S in which the decision attribute A has a value of the vth attribute. v ) represents a subset S containing the v-th attribute value. v Information entropy.

[0114] The subset S containing the v-th attribute value is determined using the following formula. v Information entropy:

[0115]

[0116] Among them, |S v | represents a subset S v The number of samples, |C i v | represents a subset Sv The number of samples in the i-th category.

[0117] Based on the information entropy of the training sample set and the information entropy of each decision attribute, the information gain of each decision attribute is determined according to the following formula:

[0118] Gain(S,A)=Entropy(S)-Entropy(S,A)

[0119] Where Gain(S,A) represents the information gain of decision attribute A, Entropy(S) represents the information entropy of training sample set S, and Entropy(S,A) represents the information entropy of decision attribute A.

[0120] Based on the information gain of each decision attribute, the information gain ratio of each decision attribute is determined according to the following formula:

[0121]

[0122] Wherein, GainRatio(S,A) represents the information gain ratio of decision attribute A, Gain(S,A) represents the information gain of decision attribute A, SplitInformation(S,A) represents the split information of decision attribute A, and the information gain ratio uses the split information to normalize the information gain.

[0123] The formula for calculating the splitting information of decision attribute A is as follows:

[0124]

[0125] Where |S| represents the total number of samples in the training sample set S, |S v | represents a subset S v The number of samples.

[0126] Step 330: Determine the decision attribute as the root node based on the information gain and information gain ratio of each decision attribute.

[0127] Based on the information gain and information gain ratio of each decision attribute, the optimal decision attribute is selected from all decision attributes as the decision attribute of the root node. For example, the average information gain of all decision attributes can be determined based on their information gain, and the decision attribute with an information gain greater than the average information gain can be selected as a candidate attribute. From all candidate attributes, the candidate attribute with the largest information gain ratio is selected as the decision attribute of the root node.

[0128] Step 340: Based on the attribute value of the decision attribute of the root node, the training sample set is divided into multiple subsets, and the decision attribute of the child node of the root node is determined according to each subset. The decision attribute of the child node of the root node is determined recursively until the initial decision tree is constructed.

[0129] Based on the attribute value of the root node's decision attribute, the training sample set is divided into multiple subsets. For each subset, the decision attribute of the child node designated as the root node is recursively determined. This process is repeated for each child node in the next level until leaf nodes are obtained, thus completing the initial decision tree construction. The recursive determination of the decision attribute for each level of child node can be performed in the same way as determining the decision attribute of the root node's child nodes.

[0130] Step 350: Based on the test sample set, prune the initial decision tree according to the sample error rate to obtain the decision tree.

[0131] The initial decision tree may overfit. To avoid this, the initial decision tree can be pruned based on the test sample set and the sample error rate. This removes unnecessary nodes from the initial decision tree, resulting in the final decision tree. The sample error rate refers to the ratio of incorrectly classified samples, specifically the proportion between the number of samples misclassified by a node and the total number of test samples classified by that node.

[0132] In one embodiment of this application, the step of pruning the initial decision tree according to the sample error rate based on the test sample set to obtain the decision tree includes:

[0133] The initial decision tree is pruned from bottom to top to determine the category corresponding to making the current node a new leaf node. The current node is a non-leaf node in the initial decision tree, and the category includes rescheduling or not rescheduling.

[0134] Based on the test sample set, determine the first sample error rate when the current node is used as the new leaf node, and determine the second sample error rate when the current node is not used as the new leaf node;

[0135] When the error rate of the first sample is less than the error rate of the second sample, the subtree rooted at the current node is deleted, and the current node is used as the new leaf node.

[0136] The current node is the node currently performing the pruning judgment.

[0137] This application employs an error rate pruning method to prune the initial decision tree. Error rate pruning (Reduce Error Pruning) is a decision tree pruning method whose main process is as follows: A training set and a test set are partitioned. The training set is used to form the learned decision tree, and the test set is used to evaluate and prune the decision tree. For overfitted decision trees built on the training set, all subtrees are traversed from bottom to top for pruning until the error rate cannot be further reduced for the test set.

[0138] The initial decision tree is pruned from bottom to top, starting from the parent node of each leaf node and proceeding upwards. For example, when deleting a subtree rooted at the current node, the current node is designated as the new leaf node, and its category is determined by majority voting. For instance, to determine the category of the new leaf node, the number of test samples in each category corresponding to the current node can be determined from the test sample set, and the category with the most test samples is chosen as the category corresponding to the new leaf node.

[0139] Compare the sample error rates before and after deleting the subtree rooted at the current node. Based on the test sample set, determine the first sample error rate when the current node is the new leaf node (i.e., the first sample error rate when assuming the subtree rooted at the current node is deleted), and determine the second sample error rate when the current node is not the new leaf node (i.e., the second sample error rate when the subtree rooted at the current node is not deleted). Compare the first sample error rate and the second sample error rate to determine whether to prune the current node. If the first sample error rate is less than the second sample error rate, determine to prune the current node, i.e., delete the subtree rooted at the current node, and make the current node the new leaf node. The category of the new leaf node is determined as the category of the current node.

[0140] The formula for error rate pruning is as follows:

[0141]

[0142] Among them, R test (T) represents the sample error rate, n test (T) represents the number of samples in the test sample set assigned to node T, e test (T) represents the number of samples that node T misclassifies in the test sample set.

[0143] The first-sample error rate and the second-sample error rate can be determined using the formulas described above.

[0144] By pruning the initial decision tree according to the error rate, overfitting of the decision tree can be prevented, and the correctness of the decision tree's decisions can be guaranteed.

[0145] After constructing an initial decision tree based on the training sample set, the initial decision tree is pruned based on the test sample set to obtain a final decision tree. This final decision tree can prevent overfitting and improve the decision accuracy of the decision tree in the application process.

[0146] Figure 5 This is a flowchart of a Hive script execution exception handling method provided in an embodiment of this application, such as... Figure 5 As shown, the exception handling method for this Hive script execution includes:

[0147] Step 510: Collect script error information and runtime environment information.

[0148] The script error information and runtime environment information are collected according to the information collection method in the above embodiments.

[0149] Step 520: Analyze the script error messages and runtime environment information.

[0150] The script error information and runtime environment information are parsed according to the information parsing method in the above embodiments to obtain the attribute values ​​of each decision attribute.

[0151] Step 530: Construct the training sample set and the test sample set.

[0152] Based on the actual production needs, the category information is labeled to obtain the category information of each error script, that is, the category information corresponding to each attribute value. As a sample, all samples are divided into training sample set and test sample set.

[0153] Step 540: Train the model to obtain a decision tree.

[0154] The model is trained according to the decision tree construction method described in the above embodiment to construct a decision tree.

[0155] Step 550: Obtain the new error script.

[0156] The new error scripts are the actual error scripts monitored, including error script information, runtime environment information, and attribute values ​​of each decision attribute obtained through parsing.

[0157] Step 560: Use the decision tree to make a decision on the new erroneous script and determine whether to resume execution from the breakpoint (rescheduling).

[0158] Step 570: When it is determined that the breakpoint is being resumed, a temporary script is generated and the temporary script is rescheduled.

[0159] The program terminates when it is determined that the process will not continue continuously.

[0160] This application proposes a method and intelligent operation tool for automatic recovery of Hive programs from unexpected interruptions, addressing the frequent occurrence of such interruptions in big data analytics. This approach can be applied in the big data field. This application makes decisions regarding resuming Hive scripts after breakpoints based on error script information and runtime environment information. When the decision is to reschedule, the erroneous script is rescheduled. This effectively eliminates various errors caused by script problems, system anomalies, such as runtime memory overflow, compression errors, missing source tables, and unsuccessful sub-script execution. It enables real-time monitoring of script errors and immediate handling of any errors, avoiding manual monitoring, reducing labor costs, and improving fault handling efficiency. In practice, the rerun success rate exceeds 90%, significantly improving fault handling efficiency while greatly reducing the labor costs of system operation.

[0161] Figure 6 This is a schematic diagram of the structure of a Hive script execution exception handling device provided in an embodiment of this application, as shown below. Figure 6 As shown, the device includes:

[0162] The script error and environment information collection module 610 is used to obtain error script information and runtime environment information when an error occurs during Hive script execution;

[0163] The script error and environment information parsing module 620 is used to parse the error script information and the runtime environment information to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type, script information attribute and environment information attribute of the Hive script.

[0164] The scheduling decision module 630 is used to determine whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree;

[0165] The breakpoint task scheduling module 640 is used to generate a new temporary script based on the breakpoint statement identifier in the error script information when it is determined that the Hive script needs to be rescheduled, and to schedule the execution of the new temporary script.

[0166] Optionally, the error reporting script information includes: script priority, execution start time, error time, error SQL statement identifier, and script error message;

[0167] The runtime environment information includes: the current environment information of the Hive library tenant and the current task scheduling environment information.

[0168] Optionally, the current environment information of the Hive library tenant includes: memory information and CPU usage information;

[0169] The current task scheduling environment information includes: the number of concurrent script scheduling requests.

[0170] Optionally, the script error and environment information parsing module includes:

[0171] An error type determination unit is used to obtain keywords from the script error message and determine the error type of the Hive script based on the keywords.

[0172] The error script information parsing unit is used to parse the error script information and determine the attribute value corresponding to the script information attribute.

[0173] The runtime environment information parsing unit is used to parse the runtime environment information and determine the attribute values ​​corresponding to the environment information attributes.

[0174] Optionally, the scheduling decision module is specifically used for:

[0175] Based on each decision attribute in the decision tree, the attribute values ​​of the decision attributes are sequentially substituted into the branches corresponding to the respective decision attributes in the decision tree to obtain the classification result of whether or not to reschedule the Hive script.

[0176] Optionally, the breakpoint task scheduling module is specifically used for:

[0177] Get the number of times the Hive script can resume execution from breakpoints;

[0178] When the number of times the breakpoint can be resumed is less than the threshold, after a preset delay, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

[0179] Optionally, the apparatus further includes a decision tree construction module, the decision tree construction module comprising:

[0180] The sample acquisition unit is used to acquire the training sample set and the test sample set;

[0181] The information gain determination unit is used to determine the information gain and information gain ratio of each decision attribute based on the training sample set.

[0182] The root node determination unit is used to determine the decision attribute as the root node based on the information gain and information gain ratio of each decision attribute.

[0183] The recursive construction unit is used to divide the training sample set into multiple subsets based on the attribute value of the decision attribute of the root node, and to determine the decision attribute of the child node of the root node based on each subset, and recursively execute the determination of the decision attribute of the child node of the next layer until the initial decision tree is constructed.

[0184] The pruning unit is used to prune the initial decision tree according to the sample error rate based on the test sample set to obtain the decision tree.

[0185] Optionally, the pruning unit is specifically used for:

[0186] The initial decision tree is pruned from bottom to top to determine the category corresponding to making the current node a new leaf node. The current node is a non-leaf node in the initial decision tree, and the category includes rescheduling or not rescheduling.

[0187] Based on the test sample set, determine the first sample error rate when the current node is used as the new leaf node, and determine the second sample error rate when the current node is not used as the new leaf node;

[0188] When the error rate of the first sample is less than the error rate of the second sample, the subtree rooted at the current node is deleted, and the current node is used as the new leaf node.

[0189] The Hive script execution exception handling device provided in this application embodiment is used to implement the steps of the Hive script execution exception handling method described in this application embodiment. The specific implementation of each module of the device is described in the corresponding steps, and will not be repeated here.

[0190] The Hive script execution exception handling device provided in this application embodiment obtains error script information and runtime environment information when a Hive script encounters an error, parses the error script information and runtime environment information, determines the attribute value of a decision attribute, and determines whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree. When it is determined that the Hive script should be rescheduled, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled for execution. Since the decision on whether to reschedule the erroneous Hive script can be based on the error script information and runtime environment information, instead of rerunning the complete code or blindly re-tuning, the waste of cluster system resources can be avoided. Moreover, by re-tuning when it is determined that re-tuning is possible, the efficiency of fault handling can be improved and the labor cost of system operation can be reduced.

[0191] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the electronic device 700 may include one or more processors 710 and one or more memories 720 connected to the processors 710. The electronic device 700 may also include an input interface 730 and an output interface 740 for communicating with another device or system. Program code executed by the processor 710 may be stored in the memory 720.

[0192] The processor 710 in the electronic device 700 calls the program code stored in the memory 720 to execute the Hive script execution exception handling method in the above embodiment.

[0193] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the Hive script execution exception handling method as described in this application.

[0194] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0195] The above provides a detailed description of a Hive script execution exception handling method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

[0196] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

Claims

1. A method for handling exceptions during Hive script execution, characterized in that, include: When an error occurs during Hive script execution, obtain the error script information and runtime environment information; The error script information and the runtime environment information are parsed to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type, script information attribute, and environment information attribute of the Hive script. Based on the attribute values ​​of the decision attributes, a decision tree is used to determine whether to reschedule the Hive script. When it is determined that the Hive script needs to be rescheduled, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed. The error script information includes: script priority, start time, error time, error SQL statement identifier, and script error message; The runtime environment information includes: the current environment information of the Hive library tenant and the current task scheduling environment information; Based on the attribute values ​​of the decision attributes, a decision tree is used to determine whether to reschedule the Hive script, including: According to each decision attribute in the decision tree, the attribute value of the decision attribute is substituted into the branch corresponding to the decision attribute in the decision tree in turn to obtain the classification result of whether the Hive script should be rescheduled; wherein, the classification result is either to reschedule the Hive script or not to reschedule the Hive script.

2. The method according to claim 1, characterized in that, The current environment information of the Hive library tenant includes: memory information and CPU usage information; The current task scheduling environment information includes: the number of concurrent script scheduling requests.

3. The method according to claim 1, characterized in that, The step of parsing the error script information and the runtime environment information to determine the attribute values ​​of the decision attributes includes: Obtain keywords from the script error message and determine the error type of the Hive script based on the keywords; The error script information is parsed to determine the attribute values ​​corresponding to the script information attributes; The runtime environment information is parsed to determine the attribute values ​​corresponding to the environment information attributes.

4. The method according to any one of claims 1-3, characterized in that, The step of generating a new temporary script based on the breakpoint statement identifier in the error script information and scheduling the execution of the new temporary script includes: Get the number of times the Hive script can resume execution from breakpoints; When the number of times the breakpoint can be resumed is less than the threshold, after a preset delay, a new temporary script is generated based on the breakpoint statement identifier in the error script information, and the new temporary script is scheduled to be executed.

5. The method according to any one of claims 1-3, characterized in that, The process of constructing the decision tree includes: Obtain the training sample set and the test sample set; Based on the training sample set, determine the information gain and information gain ratio of each decision attribute; The decision attribute to be used as the root node is determined based on the information gain and information gain ratio of each decision attribute. Based on the attribute value of the decision attribute of the root node, the training sample set is divided into multiple subsets, and the decision attribute of the child node of the root node is determined according to each subset. The decision attribute of the child node of the root node is determined recursively until the initial decision tree is constructed. Based on the test sample set, the initial decision tree is pruned according to the sample error rate to obtain the final decision tree.

6. The method according to claim 5, characterized in that, The step of pruning the initial decision tree according to the sample error rate based on the test sample set to obtain the decision tree includes: The initial decision tree is pruned from bottom to top to determine the category corresponding to making the current node a new leaf node. The current node is a non-leaf node in the initial decision tree, and the category includes rescheduling or not rescheduling. Based on the test sample set, determine the first sample error rate when the current node is used as the new leaf node, and determine the second sample error rate when the current node is not used as the new leaf node; When the error rate of the first sample is less than the error rate of the second sample, the subtree rooted at the current node is deleted, and the current node is used as the new leaf node.

7. A Hive script execution exception handling device, characterized in that, include: The script error and environment information collection module is used to obtain error information and runtime environment information when errors occur during Hive script execution; The script error and environment information parsing module is used to parse the error script information and the runtime environment information to determine the attribute value of the decision attribute. The decision attribute is used to determine whether to reschedule the Hive script. The decision attribute includes the error type, script information attribute, and environment information attribute of the Hive script. The scheduling decision module is used to determine whether to reschedule the Hive script based on the attribute value of the decision attribute through a decision tree; The breakpoint task scheduling module is used to generate a new temporary script based on the breakpoint statement identifier in the error script information when it is determined that the Hive script needs to be rescheduled, and to schedule the execution of the new temporary script. The error script information includes: script priority, start time, error time, error SQL statement identifier, and script error message; The runtime environment information includes: the current environment information of the Hive library tenant and the current task scheduling environment information; The scheduling decision module is specifically used for: According to each decision attribute in the decision tree, the attribute value of the decision attribute is substituted into the branch corresponding to the decision attribute in the decision tree in turn to obtain the classification result of whether the Hive script should be rescheduled; wherein, the classification result is either to reschedule the Hive script or not to reschedule the Hive script.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the Hive script execution exception handling method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the Hive script execution exception handling method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Telecom operator mass data processing method based on Hadoop platform

    CN103425762A

  • Script error correction processing method and device

    CN114328219A

  • Intelligent analysis method and system for software test result

    CN117591395A