Data operation optimization processing method and device
By obtaining the data to be detected and determining the number of hot spot data based on the outlier detection algorithm, and planning the task operation is carried out in combination with the available resources in the task operation period, the planning lag and resource waste caused by hot spot data processing in the existing technology is solved, and efficient data processing and user experience improvement are achieved.
Patent Information
- Application Number
- CN202111086782.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-09-16
AI Technical Summary
The prior art has problems such as planning lag caused by data tilt, inability to take into account both timeliness and stability, and wasted computing resources when processing hotspot data.
By obtaining the data to be detected, the hot spot data is determined based on the outlier value detection algorithm, the number of buckets is determined based on the indicator count value of the hot spot data, and the task operation plan is carried out based on the number of buckets and the available resources in the task operation period.
Automatic and reasonable planning of hot spot data processing is realized, avoiding delays and resource waste caused by data tilt and improving user experience.
Smart Images

Figure CN113806047B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for optimizing data operation processing. Background Art
[0002] With the popularization of major online platforms, various popular products, promotional activities, etc. are likely to cause hot data in multiple links such as exposure, browsing, clicking, adding to cart, and orders, resulting in data skew during the operation of tasks in the background when performing operations such as association, grouping, and aggregation on the hot data, and then data processing delay, and delay and lag occur on the user side.
[0003] Currently, for this kind of phenomenon, most of them are to manually adjust the SQL calculation logic setting to fix and scatter the hot data by several times after the data skew occurs, and the scattering multiple cannot be adaptively regulated. When performing parallel calculations subsequently, only the number of tasks for parallel calculations can be fixed, and the number of tasks running simultaneously cannot be adaptively adjusted according to the resource usage.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the existing manual intervention and processing to adjust the operation plan after data skew:
[0005] (1) The problem of planning lag. Since the hot spots cannot be predicted in advance, the effective allocation of task resources cannot be carried out in advance.
[0006] (2) The problem of being unable to balance timeliness and stability. Since manually adjusting the SQL calculation logic setting to fix and scatter the hot data by several times is already after the data skew occurs, there will not only be delays and instability on the user side, but also poor user experience due to the need for real-time adjustment.
[0007] (3) The problem of wasting computing resources. When performing parallel calculations after scattering the hot data, only the maximum task parallelism can be fixed for parallel calculations. If the number of task links is less than the fixed maximum task parallelism at this time, computing resources will be wasted. Summary of the Invention
[0008] In view of this, an embodiment of the present invention provides a method and device for optimizing data operation processing, which can well solve at least one of the problems of planning lag, being unable to balance timeliness and stability, and wasting computing resources existing in the existing processing of hot data resulting in data skew, and improve the user experience.
[0009] To achieve the above object, according to the first aspect of the present invention, a method for optimizing data operation processing is provided.
[0010] The data operation optimization and processing method of the present invention includes: obtaining the data to be detected, determining the hot data in the data to be detected based on the outlier detection algorithm, determining the number of buckets for the hot data according to the index count value of the hot data, and performing task operation planning on the hot data according to the number of buckets of the hot data and the available resource situation during the task operation period.
[0011] Optionally, the obtaining the data to be detected includes: obtaining real-time stream data corresponding to the acquisition configuration information based on a real-time stream processing engine, and using the real-time stream data as the data to be detected; and / or querying an offline data table according to the acquisition configuration information to obtain the corresponding offline data, and using the queried offline data as the data to be detected. Wherein, the acquisition configuration information includes: the fields of interest, and the hot data evaluation indicators corresponding to the fields of interest.
[0012] Optionally, the outlier detection algorithm includes: the normal distribution three-standard deviation algorithm, or the quartile box plot algorithm.
[0013] Optionally, when the outlier detection algorithm is the normal distribution three-standard deviation algorithm, the determining the hot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hot data evaluation indicator to obtain the aggregated index count value; determining the three-standard deviation of the normal distribution corresponding to the data to be detected, and using it as the first statistic; screening out the data in the data to be detected whose aggregated index count value exceeds the first statistic as sample data; determining the three-standard deviation of the normal distribution corresponding to the sample data, and using it as the second statistic; screening out the data in the sample data whose index count value exceeds the second statistic as hot data.
[0014] Optionally, when the outlier detection algorithm is the quartile box plot algorithm, the determining the hot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hot data evaluation indicator to obtain the aggregated index count value; arranging the data to be detected in ascending order according to the aggregated index count value, and calculating the difference between the third quartile and the first quartile of the data to be detected as the third statistic; calculating the sum of the third quartile and 1.5 times the third statistic as the first upper limit; screening out the data in the data to be detected whose index count value is higher than the first upper limit as sample data; arranging the sample data in ascending order according to the index count value, and calculating the difference between the third quartile and the first quartile of the sample data as the fourth statistic; calculating the sum of the third quartile and 1.5 times the fourth statistic as the second upper limit; screening out the data in the sample data whose index count value is higher than the second upper limit as hot data.
[0015] Optionally, determining the number of buckets for the hotspot data according to the metric count value of the hotspot data includes: determining the average value of the metric count values of the sample data; determining the average value of the metric count values of the hotspot data; and determining the number of buckets for the hotspot data according to the average value of the metric count values of the hotspot data and the average value of the metric count values of the sample data.
[0016] Optionally, performing task operation planning on the hotspot data according to the number of buckets for the hotspot data and the available resource situation corresponding to the task operation period includes: determining the number of task links required to process the hotspot data according to the number of buckets for the hotspot data and the maximum number of task links; determining the task parallelism of the task operation period according to the available resource situation of the task operation period; and determining the operation plan of the task links required to process the hotspot data according to the task parallelism required to process the hotspot data and the processing duration of a single task link.
[0017] Optionally, determining the number of task links required to process the hotspot data according to the number of buckets for the hotspot data includes: comparing the number of buckets with the maximum number of task links. If the number of buckets is less than or equal to the maximum number of task links, then using the number of buckets as the number of task links for the hotspot data. If the number of buckets is greater than the maximum number of task links, then using the maximum number of task links as the number of task links for the hotspot data.
[0018] To achieve the above object, according to the second aspect of the present invention, a data operation optimization processing device is provided.
[0019] The data operation optimization processing device of the present invention includes: a data acquisition module for acquiring data to be detected; a hotspot data detection module for determining hotspot data in the data to be detected based on an outlier detection algorithm; a bucket number determination module for determining the number of buckets for the hotspot data according to the data volume statistical result of the hotspot data; and an operation optimization module for performing task operation planning on the hotspot data according to the number of buckets for the hotspot data and the available resource situation of the task operation period.
[0020] To achieve the above object, according to the third aspect of the present invention, an electronic device is provided.
[0021] The electronic device of the present invention includes: one or more processors; and a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the data operation optimization processing method of the present invention.
[0022] To achieve the above object, according to the sixth aspect of the present invention, a computer-readable medium is provided.
[0023] A computer-readable medium of the present invention stores a computer program thereon, and when the program is executed by a processor, it implements the data operation optimization processing method of the present invention.
[0024] One embodiment of the above invention has the following advantages or beneficial effects: By acquiring the data to be detected and determining the hot data in the data to be detected based on the outlier detection algorithm, the technical problem of not being able to predict hot data in the early stage is overcome, thus providing a basis for solving subsequent problems. By using the technical means of determining the number of buckets for the hot data according to the index count value of the hot data, the technical problems of lagging planning and inability to balance timeliness and stability existing in real-time data skew processing in the prior art are overcome, and further the technical effect of pre-judging hot data and performing bucket planning processing on it to ensure stable and orderly data operation is achieved. By using the technical means of performing task operation planning on the hot data according to the number of buckets of the hot data and the available resources during the task operation period, the technical problem of wasting computing resources caused by fixing the number of tasks in parallel computing in the prior art is overcome, and further the technical effect of reasonably planning and using computing resources is achieved. Through the technical means adopted by the present invention, finally, the system can automatically and reasonably plan the processing of hot data in advance, thereby improving the user experience.
[0025] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0027] Figure 1 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;
[0028] Figure 2 is a main flowchart of the data operation optimization processing method according to the first embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of an optional hot data detection principle according to an embodiment of the present invention;
[0030] Figure 4 is a main flowchart of the data operation optimization processing method according to the second embodiment of the present invention;
[0031] Figure 5 is a main module diagram of the data operation optimization processing device according to the third embodiment of the present invention;
[0032] Figure 6It is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners
[0033] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0034] It should be noted that, without affecting the implementation of the present invention, the various embodiments in the present invention and the technical features in the embodiments can be combined with each other.
[0035] Figure 1 An exemplary system architecture 100 to which the data operation optimization method or the data operation optimization device of the embodiments of the present invention can be applied is shown.
[0036] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0038] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0039] The server 105 may be a server providing various services, such as a background management server that supports shopping applications browsed by users using the terminal devices 101, 102, 103. For example, the background management server may detect hot data and perform task operation planning on the hot data, etc.
[0040] It should be noted that the data operation optimization method provided by the embodiments of the present invention is generally executed by the server. Correspondingly, the data operation optimization device is generally set in the server.
[0041] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in [[ ]] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0042] The First Embodiment
[0043] Figure 2 is a schematic diagram of the main process of the data operation optimization processing method according to the first embodiment of the present invention. As Figure 2 shown, the data operation optimization processing method of the embodiment of the present invention includes:
[0044] Step S201: Obtain the data to be detected.
[0045] In an optional example, obtaining the data to be detected includes: obtaining real-time stream data corresponding to the acquisition configuration information based on a real-time stream processing engine, and using the real-time stream data as the data to be detected; and / or querying an offline data table according to the acquisition configuration information to obtain the corresponding offline data, and using the queried offline data as the data to be detected; wherein, the acquisition configuration information includes: the fields of interest, and the hot data evaluation indicators corresponding to the fields of interest.
[0046] When judging the hot data in different links, it is necessary to collect different types of original data information. For example, when judging the hot data in the click link, it is necessary to collect the data information of the product identifier (sku_id) of each clicked product. Then the corresponding acquisition configuration information includes: sku_id, count amount. Wherein sku_id indicates that the field of interest is the product identifier, and count amount indicates that the hot data evaluation indicator corresponding to the field of interest is the click volume.
[0047] In an optional implementation manner, when obtaining real-time data, the real-time stream processing engine obtains the real-time stream data corresponding to the acquisition configuration according to the acquisition configuration information, stores it in the local file system as the data to be detected for retrieval during hot data judgment. Real-time data can help predict the data hot situation in advance, give early warnings, and reduce the time-consuming of directly detecting offline data.
[0048] In another alternative embodiment, during offline data acquisition, direct detection is performed using an offline data table. Offline data detection can only be performed on the T+1 day of real-time data. For the offline data table generated the previous day, an offline data query SQL is generated according to the acquisition configuration information to determine the offline data corresponding to the acquisition configuration, and the offline data is stored in the local file system as the data to be detected for retrieval during hotspot judgment. This method is applicable to relational databases or distributed databases, such as mysql (a relational database), hive (a data warehouse tool), hbase (a non-structured data warehouse), etc.
[0049] Step S202: Determine the hotspot data in the data to be detected based on the outlier detection algorithm.
[0050] In an alternative example, determining the hotspot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hotspot data evaluation index to obtain the aggregated index count value. For example, when determining whether each product is a hotspot data in the click link, it is necessary to aggregate the data information of the click volume (count amount) of each product identifier (sku_id). Therefore, when determining the hotspot data in the click link, it is necessary to aggregate the data information of the click volume (count amount) of each product identifier (sku_id) in the data to be detected, and then use the aggregation function to obtain the aggregated index count value.
[0051] In an alternative example, the outlier detection algorithm includes: the normal distribution three-standard deviation algorithm, or the quartile box plot algorithm.
[0052] In an alternative embodiment, when using the normal distribution three-standard deviation algorithm to determine outliers, first, it is necessary to determine the normal distribution of the data to be detected, and then determine the three standard deviations of the normal distribution corresponding to the data to be detected and use it as the first statistic; select the data whose index count value exceeds the first statistic from the data to be detected as the sample data, as shown in the left figure below; for the sample data, determine the normal distribution corresponding to the sample data, and then determine the three standard deviations of the normal distribution corresponding to the sample data and use it as the second statistic; select the data whose index count value exceeds the second statistic from the sample data as the hotspot data, as shown in the right figure below. Figure 3 As shown in the left figure; for the sample data, determine the normal distribution corresponding to the sample data, and then determine the three standard deviations of the normal distribution corresponding to the sample data and use it as the second statistic; select the data whose index count value exceeds the second statistic from the sample data as the hotspot data, as shown in the right figure. Figure 3 As shown in the right figure.
[0053] In another alternative embodiment, the quartile box plot algorithm arranges the data to be detected into four equal parts in ascending order of the aggregated metric count values. The values at the three split points are the quartiles. The 25th percentile of the data to be detected arranged in ascending order of the aggregated metric count values is determined as the first quartile (Q1), also known as the "lower quartile". The 50th percentile of the data to be detected arranged in ascending order of the aggregated metric count values is determined as the second quartile (Q2), also known as the "median". The 75th percentile of the data to be detected arranged in ascending order of the aggregated metric count values is determined as the third quartile (Q3), also known as the "upper quartile". The difference between the third quartile and the first quartile of the data to be detected is calculated as the third statistic, and on this basis, the sum of the third quartile and 1.5 times the third statistic is calculated as the first upper limit; the data higher than the first upper limit is screened out from the data to be detected as sample data; the sample data is arranged in ascending order of the metric count values, and the difference between the third quartile and the first quartile of the sample data is calculated as the fourth statistic; the sum of the third quartile and 1.5 times the fourth statistic is calculated as the second upper limit; the data higher than the second upper limit is screened out from the sample data as hot data.
[0054] It should be noted that the number of rounds of execution of the outlier detection algorithm of the present invention is not limited to a fixed two. Different rounds are carried out to find the data volume that better fits the actual hot data. Specifically, the number of rounds can be adjusted according to the actual situation such as the data volume level and the number of empirical outliers in each industry.
[0055] Step S203: Determine the number of buckets for the hot data according to the metric count values of the hot data.
[0056] In an alternative example, determining the number of buckets for the hot data according to the metric count values of the hot data includes: determining the average value of the metric count values of the sample data; determining the average value of the metric count values of the hot data; and determining the number of buckets for the hot data according to the average value of the metric count values of the hot data and the average value of the metric count values of the sample data.
[0057] For example, assume that the average value of the metric count values of the sample data in the click link is 500, and the average value of the metric count values of the hot data in the click link is 50. Then the number of buckets in the click link is 10 (the calculation process is 500÷50 = 10).
[0058] Step S204: Perform task operation planning on the hot data according to the number of buckets for the hot data and the available resources during the task operation period.
[0059] In this step, for a certain determined task running period, determine the task parallelism of the task running period according to the available resource situation of the task running period. If the number of buckets is greater than the task parallelism of the running period, perform the scattering operation according to the task parallelism of the running period, and perform multiple rounds of bucket scattering operations on the remaining buckets in the same way; if the number of task links is less than the task parallelism of the running period, perform the bucket scattering operation according to the number of buckets.
[0060] For example, the number of buckets is 12, and the task parallelism of the task running period is 8. At this time, the number of buckets 12 is greater than the task parallelism 8 of the task running period. In this case, use 8 as the number of buckets for this round of scattering operation, and this round of planning ends; in the next round, there are 4 remaining buckets. It can be seen that the number of buckets 4 at this time is less than the task parallelism 8 of the task running period. In this case, use 4 as the number of buckets for this round of scattering operation, and the data scattering operation plan ends at this time.
[0061] In the embodiment of the present invention, the data operation optimization processing is realized through the above steps. Compared with the prior art, because the to-be-detected data is obtained and the hot data in the to-be-detected data is determined based on the outlier detection algorithm, the technical problem that hot data cannot be predicted in the early stage is overcome, thereby providing a basis for solving subsequent problems. Because the technical means of determining the number of buckets of the hot data according to the index count value of the hot data is adopted, the technical problems of lagging planning and inability to balance timeliness and stability existing in real-time processing of data skew in the prior art are overcome, and then the technical effect of realizing pre-judging hot data and performing bucket planning processing on it to ensure stable and orderly data operation is achieved. Because the task operation plan for the hot data is carried out according to the number of buckets of the hot data and the available resource situation of the task running period, the technical problem that the number of tasks in fixed parallel computing in the prior art leads to waste of computing resources is overcome, and then the technical effect of reasonably planning and using computing resources is achieved. Through the technical means adopted by the present invention, finally, the system can automatically and reasonably plan the hot data processing in advance, thereby improving the user experience.
[0062] Second Embodiment
[0063] Figure 4 is the main flowchart of the data operation optimization processing method according to the second embodiment of the present invention. As Figure 4 shown, the data operation optimization processing method of the embodiment of the present invention includes:
[0064] Step S401: Obtain the to-be-detected data.
[0065] In an optional example, obtaining the data to be detected includes: obtaining real-time stream data corresponding to the acquisition configuration information based on a real-time stream processing engine, and using the real-time stream data as the data to be detected; and / or querying an offline data table according to the acquisition configuration information to obtain corresponding offline data, and using the queried offline data as the data to be detected; wherein the acquisition configuration information includes: fields of interest, and hot data evaluation indicators corresponding to the fields of interest.
[0066] When determining hot data in different links, different types of raw data information need to be collected. For example, when determining hot data in the click link, data information of the product identifier (sku_id) of each clicked product needs to be collected. Then the corresponding acquisition configuration information includes: sku_id, count amount. Where sku_id indicates that the field of interest is the product identifier, and count amount indicates that the hot data evaluation indicator corresponding to the field of interest is the click volume.
[0067] In an optional implementation manner, when obtaining real-time data, the real-time stream processing engine obtains the real-time stream data corresponding to the acquisition configuration according to the acquisition configuration information, and stores it in the local file system as the data to be detected for retrieval during hot spot judgment. Real-time data can help predict data hot spot situations in advance, give early warnings, and reduce the time-consuming of directly detecting offline data.
[0068] In another optional implementation manner, when obtaining offline data, direct detection is performed using the offline data table. Offline data detection can be performed one day after the real-time data (T+1). For the offline data table generated the previous day, an offline data query sql is generated according to the acquisition configuration information to determine the offline data corresponding to the acquisition configuration, and the offline data is stored in the local file system as the data to be detected for retrieval during hot spot judgment. This method is applicable to relational databases or distributed databases, such as mysql (a relational database), hive (a data warehouse tool), hbase (a non-structured data warehouse), etc.
[0069] Step S402: Determine the hot data in the data to be detected based on an outlier detection algorithm.
[0070] In an alternative example, determining the hot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hot data evaluation index to obtain the aggregated index count value. For example, when determining whether each product is hot data in the click link, it is necessary to aggregate the data information of the click volume (count amount) of each product identifier (sku_id). Therefore, when determining the click hot data, it is necessary to aggregate in the dimension of the data information of the click volume (count amount) of each product identifier (sku_id), and then use the aggregation function to obtain the aggregated index count value.
[0071] In an alternative example, the outlier detection algorithm includes: the normal distribution three - standard - deviation algorithm, or the interquartile box - plot algorithm.
[0072] In an alternative embodiment, when using the normal distribution three - standard - deviation algorithm to determine outliers, first, it is necessary to determine the normal distribution of the data to be detected. Secondly, determine the three - standard - deviation of the normal distribution corresponding to the data to be detected and use it as the first statistic; screen out the data whose index count value exceeds the first statistic from the data to be detected as sample data, as Figure 3 shown in the left - hand figure; for the sample data, determine the normal distribution corresponding to the sample data, and then determine the three - standard - deviation of the normal distribution corresponding to the sample data and use it as the second statistic; screen out the data whose index count value exceeds the second statistic from the sample data as hot data, as Figure 3 shown in the right - hand figure.
[0073] In another alternative embodiment, when using the quartile box plot algorithm to determine outliers, the data to be detected is sorted into four equal parts in ascending order according to the aggregated metric count values. The values at the three splitting points are the quartiles. The 25% number after arranging the data to be detected in ascending order according to the aggregated metric count values is determined as the first quartile point (Q1), also known as the "lower quartile". The 50% number after arranging the data to be detected in ascending order according to the aggregated metric count values is determined as the second quartile point (Q2), also known as the "median". The 75% number after arranging the data to be detected in ascending order according to the aggregated metric count values is determined as the third quartile point (Q3), also known as the "upper quartile". Calculate the difference between the third quartile point and the first quartile point of the data to be detected as the third statistic, and on this basis, calculate the sum of the third quartile point and 1.5 times the third statistic as the first upper limit; screen out the data higher than the first upper limit from the data to be detected as sample data; arrange the sample data in ascending order according to the metric count values, and calculate the difference between the third quartile point and the first quartile point of the sample data as the fourth statistic; calculate the sum of the third quartile point and 1.5 times the fourth statistic as the second upper limit; screen out the data higher than the second upper limit from the sample data as hot data.
[0074] It should be noted that the number of rounds executed by the outlier detection algorithm of the present invention is not limited to a fixed two rounds. Different rounds are carried out to find the data volume that is more in line with the actual hot data. Specifically, the number of rounds can be adjusted according to the actual situations such as the data volume level and the number of empirical outliers in each industry.
[0075] Step S403: Determine the number of buckets for the hot data according to the metric count values of the hot data.
[0076] In an alternative example, determining the number of buckets for the hot data according to the metric count values of the hot data includes: determining the average value of the metric count values of the sample data; determining the average value of the metric count values of the hot data; determining the number of buckets for the hot data according to the average value of the metric count values of the hot data and the average value of the metric count values of the sample data.
[0077] For example, assume that the average value of the metric count values of the sample data in the click link is 500, and the average value of the metric count values of the hot data in the click link is 50. Then the number of buckets in the click link is 10 (500÷50 = 10). In specific implementation, considering that the ratio of the average value of the metric count values of the hot data to the average value of the metric count values of the sample data may be a decimal, in view of this, the ratio can be rounded to ensure that the obtained number of buckets is an integer.
[0078] Step S404: Determine whether the number of buckets is greater than the maximum number of task links. If so, execute Step S405; if not, execute S406.
[0079] In this step, for a certain determined task running period, if the number of buckets is greater than the maximum number of task links in the running period, then execute S406; if the number of buckets is less than or equal to the maximum number of task links in the running period, then execute S407.
[0080] Step S405: Take the maximum number of task links as the number of task links.
[0081] In this step, for the case where the number of buckets is greater than the maximum number of task links, take the maximum number of task links as the number of task links required to process this hot data.
[0082] Step S406: Take the number of buckets as the number of task links.
[0083] In this step, for the case where the number of buckets determined in Step S403 is less than or equal to the maximum number of task links, take the number of buckets as the number of task links required to process this hot data.
[0084] Exemplarily, if the number of buckets is 12 and the maximum number of task links in the task running period is 8, at this time the number of buckets 12 is greater than the maximum number of task links 8 in the task running period. In this case, select 8 as the number of task links required to process this hot data; if the number of buckets is 4 and the maximum number of task links in the task running period is 8, at this time the number of buckets 4 is less than the maximum number of task links 8 in the task running period. In this case, select 4 as the number of task links required to process this hot data.
[0085] Step S407: Determine the task parallelism of the task running period according to the available resource situation of the task running period.
[0086] In an optional example, the task parallelism is determined according to the resource idle situation in each period of the historical data. Further, in this optional example, the resource idle situation in each period can be determined by the memory usage situation of the queue. The task parallelism can be stored in the configuration table after being determined for use when determining the operation plan later.
[0087] It should be noted that the historical data of the present invention can be the past half hour, or one hour, or one day, etc. The specific selected time period can be determined according to the specific implementation situation.
[0088] Exemplarily, the task parallelism for each time period is determined based on the memory usage of a certain queue as follows: 0:00 - 1:00: the task parallelism is 10; 1:00 - 2:00: the task parallelism is 8; 2:00 - 3:00: the task parallelism is 6; 3:00 - 4:00: the task parallelism is 4; 4:00 - 5:00: the task parallelism is 6; 5:00 - 6:00: the task parallelism is 4; 6:00 - 7:00: the task parallelism is 8; 7:00 - 8:00: the task parallelism is 10; 8:00 - 9:00: the task parallelism is 10; 9:00 - 10:00: the task parallelism is 4.
[0089] Step S408: Determine the operation plan of the task links required to process the hot data according to the task parallelism of the task running period and the processing duration of a single task link.
[0090] The processing duration of a single task link is determined according to the computing power of the specific system. For example, the processing duration of a single task is 10 minutes, 5 minutes or other values.
[0091] Exemplarily, step S408 includes one or more rounds of operation plans. When it is necessary to execute the second round of this step, it is necessary to determine whether it is the same time period as the previous round of operation. If so, compare the remaining number of task links with the same task parallelism in the previous round. If not, compare the remaining number of task links with the task parallelism corresponding to the time period of this round.
[0092] It should be noted that during the task running stage, each bucket is run and calculated by a concurrent task instance, and at the same time, the task instance can correspond one by one with the data sharding bucket number according to the configured parameters.
[0093] Exemplarily, assume that the number of buckets is 25, the maximum number of task links is 23, the processing duration of a single task link is 10 minutes, and the task starts running at 3:40. Comparing the number of buckets 25 with the maximum number of task links 23, it is known that 23 is selected as the number of task links required for processing hotspots. By querying, the task parallelism from 3:00 to 4:00 is 4. Comparing the number of task links with the task parallelism at this time, it can be seen that the number of task links at this time is greater than the task parallelism. That is, the task parallelism of 4 is selected for scattering operation. That is, at 3:40 - 3:50, 4 task links are started for operation; the task parallelism from 3:50 to 4:00 is 4. Comparing the remaining number of task links (19) with the task parallelism, it can be seen that the remaining number of task links at this time is greater than the task parallelism. That is, the task parallelism of 4 is selected for scattering operation. That is, at 3:50 - 4:00, 4 task links are started for operation; from 4:00 to 5:00, the task parallelism is 6. Comparing the remaining number of task links (15) with the task parallelism, it can be seen that the remaining number of task links at this time is greater than the task parallelism. That is, the task parallelism of 6 is selected for scattering operation. That is, at 4:00 - 4:10, 4 task links are started for operation; the task parallelism from 4:10 to 4:20 is 6. Comparing the remaining number of task links (9) with the task parallelism, it can be seen that the remaining number of task links at this time is greater than the task parallelism. That is, the task parallelism of 6 is selected for scattering operation. That is, at 4:10 - 4:20, 6 task links are started for operation; the task parallelism from 4:20 to 4:30 is 6. Comparing the remaining number of task links (3) with the task parallelism, it can be seen that the remaining number of task links at this time is less than the task parallelism. That is, the remaining number of task links 3 is selected for scattering operation. That is, at 4:20 - 4:30, 3 task links are started for operation, and the planning is completed. The specific plan is shown in Table 1 below. In the case of directly scattering task concurrency, the obtained task operation plan is: 10 task links run from 3:40 to 3:50; 10 task links run from 3:50 to 4:00; 3 task links run from 4:00 to 4:10; and the task operation plan determined according to step S408 of the embodiment of the present invention is: 4 task links run from 3:40 to 3:50; 4 task links run from 3:50 to 4:00; 6 task links run from 4:00 to 4:10; 6 task links run from 4:10 to 4:20; 3 task links run from 4:20 to 4:30. It can be seen that through step S408 of the present invention, the technical problem in the prior art that only the number of tasks for fixed parallel computing can be overcome, and the technical effect of reasonably planning and using computing resources to avoid wasting computing resources is produced.
[0094] Table 1
[0095]
[0096] In the embodiments of the present invention, the above steps are used to achieve optimized processing of data operation. Compared with the prior art, by acquiring the data to be detected and determining the hot data in the data to be detected based on the outlier detection algorithm, the technical problem that hot data cannot be predicted in the early stage is overcome, thus providing a basis for solving subsequent problems. By using the technical means of determining the number of buckets for the hot data according to the index count value of the hot data, the technical problems of lagging planning and inability to balance timeliness and stability existing in real-time processing of data skew in the prior art are overcome, and further the technical effect of pre-judging hot data and performing bucket planning processing on it to ensure stable and orderly data operation is achieved. By using the technical means of performing task operation planning on the hot data according to the number of buckets of the hot data and the available resources during the task operation period, the technical problem that the fixed number of tasks for parallel computing in the prior art leads to waste of computing resources is overcome, and further the technical effect of reasonably planning and using computing resources is achieved. Through the technical means adopted in the present invention, finally, the system can automatically and reasonably plan the processing of hot data in advance, thereby improving the user experience.
[0097] Third Embodiment
[0098] Figure 5 It is a schematic diagram of the main modules of the data operation optimization device according to the third embodiment of the present invention. As Figure 5 shown, the data operation optimization device 500 in the embodiments of the present invention includes: a data acquisition module 501, a hot data detection module 502, a bucket number determination module 503, and an operation optimization module 504.
[0099] The data acquisition module 501 is used to acquire the data to be detected.
[0100] Exemplarily, acquiring the data to be detected includes: acquiring real-time stream data corresponding to the acquisition configuration information based on a real-time stream processing engine, and using the real-time stream data as the data to be detected; and / or, querying an offline data table according to the acquisition configuration information to obtain the corresponding offline data, and using the queried offline data as the data to be detected; wherein, the acquisition configuration information includes: fields of interest, hot data evaluation indicators corresponding to the fields of interest.
[0101] When judging hot data in different links, different types of original data information need to be collected. For example, when judging hot data in the click link, data information of the product identifier (sku_id) of each clicked product needs to be collected. Then the corresponding acquisition configuration information includes: sku_id, count amount. Where sku_id represents that the field of interest is the product identifier, and count amount represents that the hot data evaluation indicator corresponding to the field of interest is the click volume.
[0102] In an alternative embodiment, when acquiring real-time data, the real-time stream processing engine obtains the real-time stream data corresponding to the acquisition configuration according to the acquisition configuration information, stores it in the local file system as data to be detected, so as to be retrieved during hotspot judgment. Real-time data can help predict data hotspot situations in advance, issue early warnings, and reduce the time consumption of directly detecting offline data.
[0103] In another alternative embodiment, when acquiring offline data, direct detection is performed using an offline data table. Offline data detection can be performed one day after the real-time data (T+1). For the offline data table generated the previous day, an offline data query sql is generated according to the acquisition configuration information to determine the offline data corresponding to the acquisition configuration, and the offline data is stored in the local file system as data to be detected, so as to be retrieved during hotspot judgment. This method is applicable to relational databases or distributed databases, such as mysql (a relational database), hive (a data warehouse tool), hbase (a non-structured data warehouse), etc.
[0104] The hotspot data detection module 502 is used to determine the hotspot data in the data to be detected based on an outlier detection algorithm.
[0105] Exemplarily, determining the hotspot data in the data to be detected based on an outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hotspot data evaluation index to obtain the aggregated index count value. For example, when determining whether each commodity is a hotspot data in the click link, it is necessary to aggregate the data information of the click volume (count amount) of each commodity identifier (sku_id). Therefore, when judging the click hotspot data, it is necessary to aggregate in the dimension of the data information of the click volume (count amount) of each commodity identifier (sku_id), and then use an aggregation function to obtain the aggregated index count value.
[0106] In an alternative example, the outlier detection algorithm includes: the normal distribution three-standard-deviation algorithm, or, the quartile box plot algorithm.
[0107] In an alternative embodiment, when using the normal distribution three-standard-deviation algorithm to determine outliers, first, it is necessary to determine the normal distribution of the data to be detected, and secondly, determine the three standard deviations of the normal distribution corresponding to the data to be detected, and use it as the first statistic; select the data whose index count value exceeds the first statistic from the data to be detected as sample data, such as Figure 3As shown in the left figure; for the sample data, determine the normal distribution corresponding to the sample data, and then determine the three - standard - deviation of the normal distribution corresponding to the sample data, and use it as the second statistic; screen out the data with the index count value exceeding the second statistic from the sample data as the hotspot data, such as Figure 3 As shown in the right figure.
[0108] In another alternative embodiment, when using the inter - quartile box - plot algorithm to determine outliers, the data to be detected is sorted in ascending order of the aggregated index count value and divided into four equal parts. The values at the three split - point positions are the quartiles. The 25% number of the data to be detected sorted in ascending order of the aggregated index count value is determined as the first quartile point (Q1), also known as the "lower quartile". The 50% number of the data to be detected sorted in ascending order of the aggregated index count value is determined as the second quartile point (Q2), also known as the "median". The 75% number of the data to be detected sorted in ascending order of the aggregated index count value is determined as the third quartile point (Q3), also known as the "upper quartile". Calculate the difference between the third quartile point and the first quartile point of the data to be detected as the third statistic. On this basis, calculate the sum of the third quartile point and 1.5 times the third statistic as the first upper limit; screen out the data higher than the first upper limit from the data to be detected as the sample data; sort the sample data in ascending order of the index count value, calculate the difference between the third quartile point and the first quartile point of the sample data as the fourth statistic; calculate the sum of the third quartile point and 1.5 times the fourth statistic as the second upper limit; screen out the data higher than the second upper limit from the sample data as the hotspot data.
[0109] It should be noted that the number of rounds executed by the outlier detection algorithm of the present invention is not fixed at two. Different rounds are carried out to find the data volume that is more in line with the actual hotspot data. Specifically, the number of rounds depends on the actual situation such as the data volume of each industry and the number of empirical outliers.
[0110] The bin - number determination module 503 is used to determine the number of bins of the hotspot data according to the index count value of the hotspot data.
[0111] Exemplarily, determining the number of bins of the hotspot data according to the index count value of the hotspot data includes: determining the average value of the index count value of the sample data; determining the average value of the index count value of the hotspot data; determining the number of bins of the hotspot data according to the average value of the index count value of the hotspot data and the average value of the index count value of the sample data.
[0112] For example, assume that the average value of the metric count of the sample data in the click session is 500, and the average value of the metric count of the hot data in the click session is 50. Then, the number of buckets in the click session is 10 (500 ÷ 50 = 10).
[0113] The operation optimization module 504 is configured to perform task operation planning on the hot data according to the number of buckets of the hot data and the available resources in the task operation period.
[0114] Exemplarily, in this step, for a certain determined task operation period, the task parallelism of the task operation period is determined according to the available resources in the task operation period. If the number of buckets is greater than the task parallelism of the operation period, the scatter operation is performed according to the task parallelism of the operation period, and the remaining buckets are subjected to multiple rounds of scatter operations on the buckets in the same method; if the number of task links is less than the task parallelism of the operation period, the scatter operation on the buckets is performed according to the number of buckets. For example, the number of buckets is 12, and the task parallelism of the task operation period is 8. At this time, the number of buckets 12 is greater than the task parallelism 8 of the task operation period. In this case, 8 is used as the number of buckets in this round for the scatter operation, and the planning of this round ends; in the next round, there are 4 remaining buckets. It can be seen that the number of buckets 4 at this time is less than the task parallelism 8 of the task operation period. In this case, 4 is used as the number of buckets in this round for the scatter operation, and the data scatter operation planning ends at this time.
[0115] In the embodiment of the present invention, the data operation optimization processing is realized through the above device. Compared with the prior art, because the to-be-detected data is acquired and the hot data in the to-be-detected data is determined based on the outlier detection algorithm, the technical problem that the hot data cannot be predicted in the early stage is overcome, thus providing a basis for solving subsequent problems. Because the technical means of determining the number of buckets of the hot data according to the metric count value of the hot data is adopted, the technical problems of lagging planning and inability to balance timeliness and stability existing in real-time data skew processing in the prior art are overcome, and further the technical effect of pre-judging hot data in advance and performing bucket planning processing on it to ensure the stable and orderly operation of data is achieved. Because the task operation planning is performed on the hot data according to the number of buckets of the hot data and the available resources in the task operation period, the technical problem of wasting computing resources caused by fixing the number of tasks in parallel computing in the prior art is overcome, and further the technical effect of reasonably planning and using computing resources is achieved. Through the technical means adopted by the present invention, finally, the system can automatically and reasonably plan the hot data processing in advance, thereby improving the user experience.
[0116] Next, refer to Figure 6, which shows a schematic structural diagram of a computer system 600 of an electronic device suitable for implementing the embodiments of the present invention. Figure 6 The illustrated computer system is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present invention.
[0117] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 508 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0118] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it can be installed into the storage section 608 as required.
[0119] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the system of the present invention are executed.
[0120] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0122] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a data acquisition module, a hot data detection module, a bucket number determination module, and a running optimization module. Among them, the names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the data acquisition module can also be described as "a module for acquiring data to be detected".
[0123] As another aspect, the present invention further provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device performs the following processes: acquiring data to be detected; determining hot data in the data to be detected based on an outlier detection algorithm; determining the number of buckets for the hot data according to the index count value of the hot data; and performing task running planning on the hot data according to the number of buckets of the hot data and the available resources during the task running period.
[0124] According to the technical solution of the embodiments of the present invention, when implementing data running optimization processing, it is possible to avoid at least one of the problems of planning lag, inability to balance timeliness and stability, and waste of computing resources caused by data skew when processing hot data, thereby improving the user experience.
[0125] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing data operation processing, characterized in that, The method includes: Obtaining data to be detected; Determining hot data in the data to be detected based on an outlier detection algorithm; When the outlier detection algorithm is the three - standard - deviation algorithm of normal distribution, the determining hot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hot - data evaluation index to obtain the aggregated index count value; determining the three - standard - deviation of the normal distribution corresponding to the data to be detected and taking it as the first statistic; screening out the data whose aggregated index count value exceeds the first statistic from the data to be detected as sample data; determining the three - standard - deviation of the normal distribution corresponding to the sample data and taking it as the second statistic; screening out the data whose index count value exceeds the second statistic from the sample data as hot data; Determining the number of buckets for the hot data according to the index count value of the hot data, including: determining the average value of the index count value of the sample data; determining the average value of the index count value of the hot data; determining the number of buckets for the hot data according to the average value of the index count value of the hot data and the average value of the index count value of the sample data; Performing task - running planning on the hot data according to the number of buckets for the hot data and the available resource situation during the task - running period.
2. The method according to claim 1, characterized in that, The obtaining data to be detected includes: Obtaining real - time stream data corresponding to the acquisition configuration information based on a real - time stream processing engine and taking the real - time stream data as the data to be detected; And / or, Querying an offline data table according to the acquisition configuration information to obtain the corresponding offline data and taking the queried offline data as the data to be detected; Wherein, the acquisition configuration information includes: fields of interest, hot - data evaluation indexes corresponding to the fields of interest.
3. The method according to claim 1, characterized in that, The outlier detection algorithm includes: the three - standard - deviation algorithm of normal distribution, or, the quartile box - plot algorithm.
4. The method according to claim 3, characterized in that, When the outlier detection algorithm is the quartile box - plot algorithm, the determining hot data in the data to be detected based on the outlier detection algorithm includes: Aggregating the data to be detected in the dimension of the hot - data evaluation index to obtain the aggregated index count value; Sorting the data to be detected in ascending order of the aggregated index count value, calculating the difference between the third quartile and the first quartile of the data to be detected as the third statistic; calculating the sum of the third quartile and 1.5 times the third statistic as the first upper limit; Screening out the data whose index count value is higher than the first upper limit from the data to be detected as sample data; sorting the sample data in ascending order of the index count value, calculating the difference between the third quartile and the first quartile of the sample data as the fourth statistic; calculating the sum of the third quartile and 1.5 times the fourth statistic as the second upper limit; Screening out the data whose index count value is higher than the second upper limit from the sample data as hot data.
5. The method according to claim 1, characterized in that, The performing task - running planning on the hot data according to the number of buckets for the hot data and the available resource situation corresponding to the task - running period includes: Determine the number of task links required to process the hotspot data according to the number of buckets of the hotspot data and the maximum number of task links; determine the task parallelism of the task running period according to the available resource situation during the task running period; determine the operation plan of the task links required to process the hotspot data according to the task parallelism required to process the hotspot data and the processing duration of a single task link.
6. The method according to claim 5, characterized in that, The determining the number of task links required to process the hotspot data according to the number of buckets of the hotspot data includes: Compare the number of buckets with the maximum number of task links. If the number of buckets is less than or equal to the maximum number of task links, use the number of buckets as the number of task links of the hotspot data. If the number of buckets is greater than the maximum number of task links, use the maximum number of task links as the number of task links of the hotspot data.
7. A device for optimizing data operation processing, characterized in that, The device includes: A data acquisition module, configured to acquire data to be detected; A hotspot data detection module, configured to determine the hotspot data in the data to be detected based on an outlier detection algorithm; when the outlier detection algorithm is the three - standard - deviation algorithm of the normal distribution, the determining the hotspot data in the data to be detected based on the outlier detection algorithm includes: aggregating the data to be detected in the dimension of the hotspot data evaluation index to obtain the aggregated index count value; determining the three - standard - deviation of the normal distribution corresponding to the data to be detected and using it as the first statistic; screening out the data with the aggregated index count value exceeding the first statistic from the data to be detected as sample data; determining the three - standard - deviation of the normal distribution corresponding to the sample data and using it as the second statistic; screening out the data with the index count value exceeding the second statistic from the sample data as hotspot data; A bucket number determination module, configured to determine the number of buckets of the hotspot data according to the index count value of the hotspot data, including: determining the average value of the index count value of the sample data; determining the average value of the index count value of the hotspot data; determining the number of buckets of the hotspot data according to the average value of the index count value of the hotspot data and the average value of the index count value of the sample data; An operation optimization module, configured to perform a task operation plan for the hotspot data according to the number of buckets of the hotspot data and the available resource situation during the task running period.
8. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 - 6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 - 6.
Citation Information
Patent Citations
Hot spot data bucket dividing method and system and computer equipment
CN111726266A
Method for automatically discovering hot keywords and hot news
CN112597280A