Hive table anomaly detection method, device, electronic device and storage medium
By analyzing data modeling configuration documents and using API to detect partitioned data exceptions, the problems of low efficiency and low accuracy of Hive table data inspection in the prior art are solved, and automated abnormal detection of Hive tables is realized, improving efficiency and accuracy.
Patent Information
- Application Number
- CN202210296294.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-24
AI Technical Summary
In the prior art, data inspection of Hive tables mainly relies on manual methods, resulting in low inspection efficiency, high labor cost and low accuracy of inspection results, making it difficult to find historical data problems.
By loading and analyzing the data modeling configuration document, obtaining normal and actual data modeling indicators and partition information, comparing and generating exception information, detecting partitioned data exceptions through the API, and outputting exception information.
Automatic detection of abnormalities of Hive tables is realized, labor costs are saved, inspection efficiency and accuracy are improved, data abnormalities can be detected in a short time, and missed inspections are avoided.
Smart Images

Figure CN114676134B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device, electronic device and storage medium for detecting anomalies in a Hive table. Background Art
[0002] Hive is a data warehouse tool based on Hadoop. Data warehouse is referred to as data warehouse. Most tables in the data warehouse are partitioned tables. Date, week, and month are used as partition fields and data is incrementally aggregated in fixed periods. In the prior art, data inspection is implemented manually based on experience. With the continuous expansion of data scale, the number of data tables in Hive data warehouse increases dramatically. Data inspection requires a lot of manpower costs, and the inspection efficiency is low and the inspection results are inaccurate. Summary of the invention
[0003] The purpose of this application is to provide a method, device, electronic device and storage medium for detecting anomalies in a Hive table. In order to have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important components or describe the scope of protection of these embodiments. Its only purpose is to present some concepts in a simple form as a preface to the detailed description that follows.
[0004] According to one aspect of an embodiment of the present application, a method for detecting anomalies in a Hive table is provided, including:
[0005] Load and parse the data modeling configuration document to obtain the corresponding normal data modeling indicators and normal partition information;
[0006] Obtaining actual data modeling indicators of the Hive table to be tested and actual partition information of the Hive table to be tested;
[0007] Compare respectively whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information, and if they are inconsistent, generate corresponding abnormal information;
[0008] Detect whether the partition data of the Hive table is abnormal, and if abnormal, generate corresponding abnormal information;
[0009] Output all the exception information.
[0010] In some embodiments of the present application, the normal partition information includes all normal partitions of the Hive table and the life cycle of each partition; parsing the data modeling configuration document to obtain the normal partition information includes:
[0011] The data modeling rules of data update frequency, data life cycle and data start time are used to parse the data modeling configuration document to obtain all normal partitions of the Hive table and the life cycle of each partition.
[0012] In some embodiments of the present application, the actual partition information of the Hive table is inconsistent with the normal partition information, including:
[0013] The partition of the Hive table is not created, the partition directory on the HDFS of the Hive table is not created, the partition data on the HDFS of the Hive table does not exist, and the partition data of the Hive table has expired but has not been backed up and migrated.
[0014] In some embodiments of the present application, the detecting whether the partition data of the Hive table is abnormal includes:
[0015] Use the hdfs API of the Hive table to obtain the data size of all partitions, and analyze whether there are partitions with abnormal data based on the data size.
[0016] In some embodiments of the present application, the detecting whether the partition data of the Hive table is abnormal includes:
[0017] For all partitions in the life cycle of the Hive table, sort all the partitions according to the partition time;
[0018] According to the sequence obtained by sorting, the first number of partitions before the detection time point is taken;
[0019] Obtaining a minimum data volume and a maximum data volume from the first number of partitions, calculating a ratio of the minimum data volume to the maximum data volume, and obtaining a data difference ratio;
[0020] Finding the valid partition data volume closest to the detection time point from the first number of partitions;
[0021] The partition data of the Hive table that does not belong to a normal closed interval is determined as abnormal data, the left endpoint of the normal closed interval is the product of the valid partition data volume and the data difference ratio, and the right endpoint is the quotient of the valid partition data volume and the data difference ratio.
[0022] In some embodiments of the present application, the detecting whether the partition data of the Hive table is abnormal includes:
[0023] Using the isolation forest algorithm, find out the data in each partition that differs from the data in the same partition by more than a preset threshold as an outlier;
[0024] Calculate the proportion of the abnormal values and the degree of difference between the abnormal values and the normal values;
[0025] If both the proportion and the degree of difference exceed respective preset thresholds, it is determined that the partition data of the Hive table is abnormal.
[0026] In some embodiments of the present application, outputting all the exception information includes: integrating all the exception information into a json file and outputting it.
[0027] According to another aspect of an embodiment of the present application, a Hive table anomaly detection device is provided, including:
[0028] Loading and parsing module, used to load and parse the data modeling configuration document to obtain the corresponding normal data modeling indicators and normal partition information;
[0029] An acquisition module is used to acquire actual data modeling indicators of a Hive table to be detected and actual partition information of the Hive table to be detected;
[0030] A comparison module, used to compare whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information, and if they are inconsistent, generate corresponding abnormal information;
[0031] A detection module is used to detect whether the partition data of the Hive table is abnormal, and if abnormal, generate corresponding abnormal information;
[0032] The output module is used to output all the abnormal information.
[0033] According to another aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-described methods for detecting anomalies in a Hive table.
[0034] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The program is executed by a processor to implement any of the above-mentioned Hive table anomaly detection methods.
[0035] The technical solution provided by one aspect of the embodiments of the present application may have the following beneficial effects:
[0036] The method for detecting anomalies in a Hive table provided in an embodiment of the present application loads and parses a data modeling configuration document, obtains corresponding normal data modeling indicators and normal partition information, and obtains actual data modeling indicators and actual partition information of a Hive table to be detected; respectively compares whether the actual data modeling indicators are consistent with the normal data modeling indicators and whether the actual partition information is consistent with the normal partition information; if they are inconsistent, generates corresponding anomaly information, detects whether the partition data of the Hive table is abnormal; if so, generates corresponding anomaly information, and outputs all anomaly information, thereby realizing automatic detection of anomalies in the Hive table, saving labor costs, improving inspection efficiency, and achieving high accuracy of inspection results.
[0037] Other features and advantages of the present application will be described in the subsequent description, and some of them will become obvious from the description, or some of them can be inferred or determined unambiguously from the description, or can be understood by implementing the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 A flow chart of an anomaly detection method for a Hive table according to an embodiment of the present application is shown;
[0040] Figure 2 A flowchart of detecting whether partition data of the Hive table is abnormal in some implementations of the present application is shown;
[0041] Figure 3 A flowchart of detecting whether partition data of the Hive table is abnormal in other embodiments of the present application is shown;
[0042] Figure 4 An example diagram of a json file in an example of the present application is shown;
[0043] Figure 5 A structural block diagram of an abnormality detection device for a Hive table according to an embodiment of the present application is shown;
[0044] Figure 6 A structural block diagram of an electronic device according to an embodiment of the present application is shown;
[0045] Figure 7 A schematic diagram of a computer-readable storage medium according to an embodiment of the present application is shown.
[0046] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0048] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0049] Hive is a data warehouse tool based on Hadoop, which is used for data extraction, conversion and loading. It is a mechanism that can store, query and analyze large-scale data stored in Hadoop. The Hive data warehouse tool can map structured data files into a database table, called a Hive table, and provide SQL query functions. The inventor found that the most common application scenario of Hive is historical offline data analysis. The problematic data may not be discovered in the first time. When the abnormality is found, it may be difficult to fill in the historical data, all sub-node data needs to be redone, and the published data is difficult to correct. The inventor also found that the existing partition data inspection method of each partition of the Hive table is mainly to first summarize the number of partition data and common indicators, and then manually distinguish whether the data is abnormal. The reliability of manual inspection is low, and there is a high probability of missed inspections. The accuracy of the inspection results is low, and manual routine inspections can only check the incremental data of the latest partition, and it is difficult to find historical data problems.
[0050] like Figure 1 As shown, an embodiment of the present application provides a method for detecting anomalies in a Hive table. In some implementations, the method includes steps S10 to S50:
[0051] S10. Load and parse the data modeling configuration document to obtain corresponding normal data modeling indicators and normal partition information.
[0052] The data modeling indicators include six indicators: library name, table name, table comment, field, field name, and field comment. The Hive table to be tested can be a Hive table partitioned by date, month, or week. Fixed-period incremental aggregation usually adds new partitions over time. The update frequency can be T+n, where T refers to an instance of a cycle and n represents the number of cycles. For example, the partition is partitioned by day, and the partition is generated based on the date as an instance. T refers to the date of the partition instance, and n refers to the actual date of generating the T partition minus the number of days between the date of the T partition instance. This embodiment can be applied to fields such as traffic flow information statistics.
[0053] An example of a Hive table is: bigdata_warehouse.dwd_vehicle_heatspeed.
[0054] In some implementations, the normal partition information includes all normal partitions of the Hive table and the life cycle of each partition; parsing the data modeling configuration document to obtain the normal partition information includes:
[0055] The data modeling rules of data update frequency, data life cycle and data start time are used to parse the data modeling configuration document to obtain all normal partitions of the Hive table and the life cycle of each partition.
[0056] After loading the data modeling configuration document, use the three data modeling rules of update frequency, data life cycle, and data start time to parse all the partitions that the table should have and the life cycle of each partition, so as to compare and judge with the actual partition information obtained from the Hive table.
[0057] S20: Acquire actual data modeling indicators of the Hive table to be detected and actual partition information of the Hive table to be detected.
[0058] Obtain the actual data modeling indicators of the Hive table to be tested. For example, you can use the show createtable [table] and desc [table] commands of Hivesql to obtain the actual modeling indicators of the Hive table.
[0059] Get the actual partition information of the Hive table to be tested. For example, you can use the Hivesql: show partitions [table] command to get all the actual partitions of the Hive table. You can use the show create table [table] command to get the storage path of the Hive table in hdfs. You can view all the files of all the actual partitions of Hive through the storage path. By analyzing the partition update frequency, data life cycle, and data start time of the table in the document, you can analyze all the partitions that should exist in the table in theory.
[0060] For example: The document states that the table is partitioned by day based on T+1, and the data life cycle is 1 year. The data start date is 2018-01-01, and the run check date is 2021-12-31. Then all the partitions that should theoretically exist are: 2020-12-31, 2021-01-01, 2021-01-02…2021-12-30, and all partitions are calculated by the program.
[0061] S30, respectively comparing whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information. If they are inconsistent, generating corresponding abnormal information.
[0062] Comparing the actual data modeling index with the normal data modeling index to see if they are consistent, and if they are inconsistent, generating corresponding abnormal information, which may include:
[0063] The library name, table name, table comment, field, field name and field comment in the actual data modeling indicators of the Hive table are compared with the corresponding library name, table name, table comment, field, field name and field comment in the normal data modeling indicators. If the indicators are completely consistent, it is determined that there is no abnormality in the indicators. If there is at least one inconsistent indicator, it is determined that the indicator is abnormal and the corresponding abnormal information is generated.
[0064] In a specific example, for vehicle data, the indicators in the normal data document are compared with the actual modeling indicators in the Hive table, and the following abnormal information is obtained:
[0065] 1) The Chinese name of the table (table annotation) is inconsistent:
[0066] In normal data files: Details - Vehicle rapid acceleration and deceleration table
[0067] In the table: ? ? ? ? (garbled string)
[0068] 2) Field types are inconsistent:
[0069] a. In normal data files: first_rapid_quicken = INT
[0070] b. In the table: first_rapid_quicken = string
[0071] 3) Field comments (field Chinese names) are inconsistent:
[0072] a. In normal data files: vid = vehicle ID
[0073] b. In the table: vid = vehicle unique identifier
[0074] Compare the actual partition information with the normal partition information to see if they are consistent, and if they are inconsistent, generate corresponding exception information. In some implementations, the actual partition information of the Hive table is inconsistent with the normal partition information, including:
[0075] The partition of the Hive table is not created, the partition directory on the hdfs of the Hive table is not created, the partition data on the hdfs of the Hive table does not exist, and the partition data of the Hive table has expired but has not been backed up and migrated. If the partition data has expired but has not been backed up and migrated, it will occupy cluster storage resources.
[0076] In a specific example, the following abnormal information can be obtained by comparing the actual partition information with the normal partition information:
[0077] The 20210829 partition is missing; the 20210829 partition file is missing.
[0078] S40: Detect whether the partition data of the Hive table is abnormal. If abnormal, generate corresponding abnormal information.
[0079] In some implementations, the detecting whether the partition data of the Hive table is abnormal includes: using the hdfs API of the Hive table to obtain the data size of all partitions, and analyzing whether there are partitions with abnormal data according to the data size. This implementation has high detection efficiency, low computing resource consumption, and wide table adaptability.
[0080] like Figure 2 As shown, in some implementations, detecting whether the partition data of the Hive table is abnormal includes:
[0081] For all partitions in the life cycle of the Hive table, sort all the partitions according to the partition time;
[0082] According to the sequence obtained by sorting, the first number of partitions before the detection time point is taken;
[0083] Obtaining a minimum data volume and a maximum data volume from the first number of partitions, calculating a ratio of the minimum data volume to the maximum data volume, and obtaining a data difference ratio;
[0084] Finding the valid partition data volume closest to the detection time point from the first number of partitions;
[0085] The partition data of the Hive table that does not belong to a normal closed interval is determined as abnormal data, the left endpoint of the normal closed interval is the product of the valid partition data volume and the data difference ratio, and the right endpoint is the quotient of the valid partition data volume and the data difference ratio.
[0086] In a specific example, we first count and detect the partition data volume, and use the HDFS API to obtain the data size of all partitions to quickly perform a preliminary analysis to obtain all abnormal partitions: the data storage service for the Hive table is HDFS, and the HDFS namenode service caches the data size, path and other information of all data files in HDFS in memory. The program uses the HDFS API to access HDFS to quickly obtain and count the data size of the specified file or folder.
[0087] Take all partitions in the life cycle, sort them according to the partition time, and then take n (partition anomaly detection sample window length) partitions from the detection date forward. If there are less than n partitions, skip them (the default is that the first n partitions have normal data). Eliminate the partitions marked as abnormal. Take the minimum data volume minSize and the maximum data volume maxSize to calculate the data difference ratio a=minSize / maxSize. The valid partition data volume b closest to the detection time point in the first n partitions, b*a as the data starting value of the normal closed interval of the current partition, and b / a as the data ending value of the normal closed interval of the current partition, that is, the normal closed interval is [b*a,b / a]. If it is not in the normal closed interval, it is initially judged as abnormal. The value of n can be set. After comparing the calculation results of the algorithm with the manual detection results, most tables have better results when n is 10 to 20. The default setting of n=14.
[0088] Count and compare the number of identical data. After detecting the partition data volume, some partitions with abnormal detection results are actually due to sudden changes in vehicle activity caused by holidays, natural environmental factors, etc., which causes a sudden increase in the number of data differences and is mistakenly judged as abnormal partitions. Therefore, after detecting the partition data volume, you can take y data from the first n normal partitions, and count the storage space occupied by y data in each partition. The partition to be detected also takes n data, and uses the same calculation method as the statistics and detection of partition data volume to obtain a more accurate abnormal partition. The larger the value of y, the more accurate the result, which needs to be adjusted according to the actual table data volume and computing resources.
[0089] like Figure 3 As shown, in some other implementations, the detecting whether the partition data of the Hive table is abnormal includes:
[0090] Using the isolation forest algorithm, find out the data in each partition that differs from the data in the same partition by more than a preset threshold as an outlier;
[0091] Calculate the proportion of the abnormal values and the degree of difference between the abnormal values and the normal values;
[0092] If both the proportion and the degree of difference exceed respective preset thresholds, it is determined that the partition data of the Hive table is abnormal.
[0093] The Isolation Forest Algorithm is an unsupervised machine learning algorithm that does not require the definition of a mathematical model or sample data training. The algorithm can perform fast anomaly detection based on Ensemble, has linear time complexity and high accuracy, and is suitable for anomaly detection of continuous data. It defines anomalies as outliers that are easily isolated, which can be understood as points that are sparsely distributed and far away from dense groups. The Isolation Forest Algorithm is mainly used to detect the outlier data in which the values of numeric type fields may be abnormal. Most of the data in a partition is normal. The Isolation Forest Algorithm can be used to find the data in a partition that is very different from most of the data. After finding the outliers, the proportion of outliers and the degree of difference between outliers and normal values can be calculated, and a threshold can be set for these two results. If the threshold is exceeded, the partition is considered to be abnormal. If this anomaly detection algorithm is used in the ODS layer, it can also be used to complete the data cleaning process. The detection method based on the Isolation Forest Algorithm is more suitable for checking the ODS layer table.
[0094] For example, the following exception information is obtained through the above steps:
[0095] The partition data from 20200706 to 20200823 is abnormal.
[0096] Check and compare the abnormal partition and the normal partition to find out the cause of the abnormality = the values of the vid field in the abnormal partition are all null.
[0097] When detecting whether the partition data of the Hive table is abnormal based on the isolation forest algorithm, the partition data closest to the running date can be detected. If the table starts with ods, the abnormal information and the abnormal row are recorded, otherwise only the abnormal information is recorded. The abnormal partition is marked as false, and the normal partition is marked as true.
[0098] When detecting whether the partition data of the Hive table is abnormal, one of the above-mentioned implementation methods can be used alone to determine whether it is abnormal, or the detection method based on the isolation forest algorithm and the detection methods of other implementation methods can be used simultaneously to obtain the detection results of each implementation method, and the final result is determined whether it is normal based on the detection results obtained by each method.
[0099] S50: Output all the abnormal information.
[0100] In some implementations, outputting all of the exception information includes: integrating all of the exception information into a json file and outputting it. Figure 4The following is a sample diagram of a json file. The json file is easier to view and analyze.
[0101] In certain implementations, the method of this embodiment further includes generating an alarm signal to remind staff to check abnormal information.
[0102] The Hive table anomaly detection method provided by the embodiment of the present application loads and parses the data modeling configuration document, obtains the corresponding normal data modeling indicators and normal partition information, obtains the actual data modeling indicators and actual partition information of the Hive table to be detected; respectively compares whether the actual data modeling indicators are consistent with the normal data modeling indicators and whether the actual partition information is consistent with the normal partition information. If they are inconsistent, corresponding anomaly information is generated to detect whether the partition data of the Hive table is abnormal. If it is abnormal, corresponding anomaly information is generated and all anomaly information is output, thereby realizing automatic anomaly detection of the Hive table, saving labor costs, improving inspection efficiency, and having high work efficiency. In the case of a relatively short time and using less server resources, data anomaly detection can be performed on all tables and all partitions in the Hive data warehouse. The inspection result has high accuracy, can effectively detect the Hive table partitions with data anomalies, has high reliability, can avoid missed detection, and can detect table partition data in various application scenarios without supervision.
[0103] In addition, the method of the embodiment of the present application can discover problematic data in the first place, thereby facilitating timely replenishment of historical data, solving the problems of the prior art such as difficulty in replenishing historical data, the need to redo all sub-node data, and difficulty in correcting published data.
[0104] like Figure 5 As shown, another embodiment of the present application provides an anomaly detection device for a Hive table, including:
[0105] Loading and parsing module, used to load and parse the data modeling configuration document to obtain the corresponding normal data modeling indicators and normal partition information;
[0106] An acquisition module is used to acquire actual data modeling indicators of a Hive table to be detected and actual partition information of the Hive table to be detected;
[0107] A comparison module, used to compare whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information, and if they are inconsistent, generate corresponding abnormal information;
[0108] A detection module is used to detect whether the partition data of the Hive table is abnormal, and if abnormal, generate corresponding abnormal information;
[0109] The output module is used to output all the abnormal information.
[0110] In some implementations, the normal partition information includes all normal partitions of the Hive table and the life cycle of each partition; parsing the data modeling configuration document to obtain the normal partition information includes:
[0111] The data modeling rules of data update frequency, data life cycle and data start time are used to parse the data modeling configuration document to obtain all normal partitions of the Hive table and the life cycle of each partition.
[0112] In some implementations, the actual partition information of the Hive table is inconsistent with the normal partition information, including:
[0113] The partition of the Hive table is not created, the partition directory on the HDFS of the Hive table is not created, the partition data on the HDFS of the Hive table does not exist, and the partition data of the Hive table has expired but has not been backed up and migrated.
[0114] In some implementations, detecting whether the partition data of the Hive table is abnormal includes:
[0115] Use the hdfs API of the Hive table to obtain the data size of all partitions, and analyze whether there are partitions with abnormal data based on the data size.
[0116] In some implementations, detecting whether the partition data of the Hive table is abnormal includes:
[0117] For all partitions in the life cycle of the Hive table, sort all the partitions according to the partition time;
[0118] According to the sequence obtained by sorting, the first number of partitions before the detection time point is taken;
[0119] Obtaining a minimum data volume and a maximum data volume from the first number of partitions, calculating a ratio of the minimum data volume to the maximum data volume, and obtaining a data difference ratio;
[0120] Finding the valid partition data volume closest to the detection time point from the first number of partitions;
[0121] The partition data of the Hive table that does not belong to a normal closed interval is determined as abnormal data, the left endpoint of the normal closed interval is the product of the valid partition data volume and the data difference ratio, and the right endpoint is the quotient of the valid partition data volume and the data difference ratio.
[0122] In some implementations, detecting whether the partition data of the Hive table is abnormal includes:
[0123] Using the isolation forest algorithm, find out the data in each partition that differs from the data in the same partition by more than a preset threshold as an outlier;
[0124] Calculate the proportion of the abnormal values and the degree of difference between the abnormal values and the normal values;
[0125] If both the proportion and the degree of difference exceed respective preset thresholds, it is determined that the partition data of the Hive table is abnormal.
[0126] In some implementations, outputting all of the exception information includes: integrating all of the exception information into a json file and outputting it.
[0127] The anomaly detection device for a Hive table provided in an embodiment of the present application and the anomaly detection method for a Hive table provided in an embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.
[0128] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the Hive table anomaly detection method of any of the above embodiments.
[0129] like Figure 6 As shown, the electronic device 10 may include: a processor 100, a memory 101, a bus 102 and a communication interface 103, and the processor 100, the communication interface 103 and the memory 101 are connected via the bus 102; the memory 101 stores a computer program that can be run on the processor 100, and the processor 100 executes the method provided in any of the aforementioned embodiments of the present application when running the computer program.
[0130] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0131] The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 101 is used to store programs, and the processor 100 executes the programs after receiving execution instructions. The method disclosed in any implementation of the above-mentioned embodiment of the present application may be applied to the processor 100, or implemented by the processor 100.
[0132] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 100 or an instruction in the form of software. The above processor 100 may be a general-purpose processor, which may include a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and completes the steps of the above method in combination with its hardware.
[0133] The electronic device provided in the embodiment of the present application and the method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, operated or implemented by them.
[0134] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the Hive table anomaly detection method of any of the above embodiments.
[0135] refer to Figure 7 The computer-readable storage medium shown is a CD 20 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the method provided in any of the aforementioned embodiments is executed.
[0136] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0137] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0138] It should be noted that:
[0139] The term "module" is not intended to be limited to a specific physical form. Depending on the specific application, a module can be implemented as hardware, firmware, software, and / or a combination thereof. In addition, different modules can share common components or even be implemented by the same components. There may or may not be clear boundaries between different modules.
[0140] The algorithm and display provided herein are not inherently related to any specific computer, virtual device or other equipment. Various general devices can also be used together with examples based on this. According to the above description, it is obvious to construct the structure required for this type of device. In addition, the application is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the application described herein, and the description made to specific languages above is for the purpose of disclosing the best mode of implementation of the application.
[0141] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0142] The above-mentioned embodiments only express the implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for detecting anomalies in a Hive table, characterized in that: include: Load and parse the data modeling configuration document to obtain the corresponding normal data modeling indicators and normal partition information; Obtaining actual data modeling indicators of the Hive table to be tested and actual partition information of the Hive table to be tested; Compare respectively whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information, and if they are inconsistent, generate corresponding abnormal information; Detect whether the partition data of the Hive table is abnormal, and if abnormal, generate corresponding abnormal information; Output all the abnormal information; The detecting whether the partition data of the Hive table is abnormal includes: Use the hdfs API of the Hive table to obtain the data size of all partitions, and analyze whether there are partitions with abnormal data based on the data size.
2. The method according to claim 1, characterized in that The normal partition information includes all normal partitions of the Hive table and the life cycle of each partition; parsing the data modeling configuration document to obtain normal partition information includes: The data modeling rules of data update frequency, data life cycle and data start time are used to parse the data modeling configuration document to obtain all normal partitions of the Hive table and the life cycle of each partition.
3. The method according to claim 1, characterized in that The actual partition information of the Hive table is inconsistent with the normal partition information, including: The partition of the Hive table is not created, the partition directory on the HDFS of the Hive table is not created, the partition data on the HDFS of the Hive table does not exist, and the partition data of the Hive table has expired but has not been backed up and migrated.
4. The method according to claim 1, characterized in that: The detecting whether the partition data of the Hive table is abnormal includes: For all partitions in the life cycle of the Hive table, sort all the partitions according to the partition time; According to the sequence obtained by sorting, the first number of partitions before the detection time point is taken; Obtaining a minimum data volume and a maximum data volume from the first number of partitions, calculating a ratio of the minimum data volume to the maximum data volume, and obtaining a data difference ratio; Finding the valid partition data volume closest to the detection time point from the first number of partitions; The partition data of the Hive table that does not belong to a normal closed interval is determined as abnormal data, the left endpoint of the normal closed interval is the product of the valid partition data volume and the data difference ratio, and the right endpoint is the quotient of the valid partition data volume and the data difference ratio.
5. The method according to claim 1, characterized in that The detecting whether the partition data of the Hive table is abnormal includes: Using the isolation forest algorithm, find out the data in each partition that differs from the data in the same partition by more than a preset threshold as an outlier; Calculate the proportion of the abnormal values and the degree of difference between the abnormal values and normal values; If both the proportion and the degree of difference exceed respective preset thresholds, it is determined that the partition data of the Hive table is abnormal.
6. The method according to claim 1, characterized in that The outputting all the exception information includes: integrating all the exception information into a json file and outputting it.
7. A Hive table anomaly detection device, characterized in that: include: Loading and parsing module, used to load and parse the data modeling configuration document to obtain the corresponding normal data modeling indicators and normal partition information; An acquisition module is used to acquire actual data modeling indicators of a Hive table to be detected and actual partition information of the Hive table to be detected; A comparison module, used to compare whether the actual data modeling index is consistent with the normal data modeling index and whether the actual partition information is consistent with the normal partition information, and if they are inconsistent, generate corresponding abnormal information; A detection module is used to detect whether the partition data of the Hive table is abnormal, and if abnormal, generate corresponding abnormal information; An output module, used for outputting all the abnormal information; The detecting whether the partition data of the Hive table is abnormal includes: Use the hdfs API of the Hive table to obtain the data size of all partitions, and analyze whether there are partitions with abnormal data based on the data size.
8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Internet of things log processing method and device based on Hadoop platform
CN105608203A
Method and apparatus for migrating data in HIVE, and terminal device
CN107301214A
An implementation method for supporting data life cycle management of a multi-database engine
CN109815219A
A loading system supporting HIVE automatic partitioning and an implementation method thereof
CN109902126A
Abnormal data detection method and device, equipment and storage medium
CN111931860A