Data processing method and device, storage medium and program product
By classifying data and dynamically adjusting the sampling rate, the problem of irrational resource allocation in traditional data collection methods is solved, and refined management and efficiency improvement of data sampling are achieved.
Patent Information
- Application Number
- CN202510772146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional data collection methods cannot be dynamically adjusted according to the actual situation of the data, resulting in irrational allocation of data processing resources. This may cause insufficient collection of high-value data, waste resources on collecting low-value data, and result in low sampling efficiency.
Data is classified based on its source and outliers, and the initial sampling rate is determined based on the preset sampling rate value, data density and timeliness indicators. The second sampling rate is dynamically adjusted according to the amount of data in different time periods to achieve refined data management.
It realizes the refined management of data sampling, fully considers the differences in data characteristics, effectively avoids over-sampling or under-sampling, and improves sampling efficiency.
Smart Images

Figure CN120705651A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, storage medium, and program product. Background Art
[0002] In existing data collection systems, it is often necessary to collect and process various types of data generated by different devices. As the number of devices increases and the scale of data expands, the amount of data generated by different devices varies significantly, and the value of this data also varies.
[0003] However, traditional data collection methods cannot be dynamically adjusted according to the actual situation of the data, resulting in irrational allocation of data processing resources. This may cause insufficient collection of some high-value data, waste resources on the collection of low-value data, and result in low sampling efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method, device, storage medium and program product to determine the initial sampling rate based on data density, timeliness index and preset mapping relationship, and dynamically adjust the second sampling rate based on the amount of data obtained in different time periods. It can fully consider the differences in data characteristics, achieve balanced data processing, and improve sampling efficiency.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, comprising: Based on the source of the first data or the request status code of the first data, the first data is classified to obtain the second data of each category; based on the mapping relationship between the sampling rate preset value and the second data of each category, as well as the data density and timeliness index of the second data of each category, the first sampling rate of the second data of each category is determined; for the first sampling rate of the second data of each category, if it is detected that the first sampling rate is within the sampling rate range, the second sampling rate of the second data is determined according to the first sampling rate, the first and second quantities of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period; based on the second sampling rate, the second data is sampled.
[0006] In a second aspect, an embodiment of the present application provides a data processing device, including: a classification module, configured to classify the first data of the first device based on a source of the first data or a request status code of the first data to obtain second data of multiple categories; a determining module, configured to determine a first sampling rate of the second data of each category based on a mapping relationship between a preset sampling rate value and the second data of each category, and a data density and a timeliness index of the second data of each category; The determination module is further configured to, for each category of the second data, determine a second sampling rate of the second data based on the first sampling rate, a first quantity and a second quantity of the second data acquired in a first time period, and a third quantity of the second data acquired in a second time period if it is detected that the first sampling rate is within a sampling rate interval; A sampling module is used to sample the second data based on the second sampling rate.
[0007] In a third aspect, an embodiment of the present application provides a data processing device, the device comprising: A memory, a processor, and a data processing program stored in the memory and executable on the processor, wherein the data processing program is configured to implement part or all of the steps described in any method in the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a data processing program is stored. When the data processing program is executed by a processor, some or all of the steps described in any method in the first aspect are implemented.
[0009] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0010] By implementing the embodiment of the present application, firstly, based on the source of the first data or the request status code of the first data, the first data is classified to obtain the second data of each category; then, based on the mapping relationship between the sampling rate preset value and the second data of each category, as well as the data density and timeliness index of the second data of each category, the first sampling rate of the second data of each category is determined; then, for the first sampling rate of the second data of each category, if it is detected that the first sampling rate is within the sampling rate interval, the second sampling rate of the second data is determined based on the first sampling rate, the first and second quantities of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period; finally, based on the second sampling rate, the second data is sampled. By combining the source of the data with the characteristics of the outliers to classify the data, accurately determining the initial sampling rate based on the data density, timeliness index and the preset mapping relationship, and dynamically adjusting the sampling rate based on the sampling rate interval and the number of data in different time periods to obtain the second sampling rate, not only can the refined management of data sampling be achieved, but also the differences in data characteristics can be fully considered, effectively avoiding over-sampling or under-sampling, achieving balanced data processing, and improving sampling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.
[0012] Figure 1 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of the present application; Figure 2 This is a flow chart of a data processing method provided by an embodiment of the present application; Figure 3 This is a flow chart of determining a first sampling rate of second data of each category provided by an embodiment of the present application; Figure 4 This is a flow chart of determining a second sampling rate of second data provided by an embodiment of the present application; Figure 5 This is a flow chart of determining a first sampling rate adjustment coefficient provided by an embodiment of the present application; Figure 6 This is a flow chart of determining a second sampling rate provided by an embodiment of the present application; Figure 7 is a structural diagram of a data processing device provided in an embodiment of the present application; Figure 8 It is a structural diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work should fall within the scope of protection of the present invention.
[0014] The terms "first," "second," and "third," etc. in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a particular order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0015] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0016] In existing data collection systems, it is often necessary to collect and process various types of data generated by different devices. As the number of devices increases and the scale of data expands, the amount of data generated by different devices varies significantly, and the value of this data also varies.
[0017] However, traditional data collection methods cannot be dynamically adjusted according to the actual situation of the data, resulting in irrational allocation of data processing resources. This may cause insufficient collection of some high-value data, waste resources on the collection of low-value data, and result in low sampling efficiency.
[0018] To address the above problems, embodiments of the present application provide a data processing method, device, storage medium, and program product. First, based on the source of the first data and the outliers of the first data, the first data is classified to obtain the second data of each category. Then, based on the mapping relationship between the preset sampling rate value and the second data of each category, as well as the data density and timeliness index of the second data of each category, the first sampling rate of the second data of each category is determined. Then, for the first sampling rate of the second data of each category, if it is detected that the first sampling rate is within the sampling rate interval, the second sampling rate of the second data is determined based on the first sampling rate, the first and second quantities of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period. Finally, based on the second sampling rate, the second data is sampled. By combining the source of the data and the outlier characteristics to classify the data, accurately determining the initial sampling rate based on the data density, timeliness index, and the preset mapping relationship, and dynamically adjusting the second sampling rate based on the sampling rate interval and the number of data in different time periods, not only can the refined management of data sampling be achieved, but also the differences in data characteristics can be fully considered, effectively avoiding over-sampling or under-sampling, achieving balanced data processing, and improving sampling efficiency.
[0019] The data processing method, device, storage medium and program product provided in the embodiments of the present application can be applied to Figure 1 In the data processing system shown, see Figure 1 , Figure 11 is a schematic diagram of the architecture of a data processing system provided in an embodiment of the present application. Data processing system 100 includes client 101 and server 102. Client 101 can communicate with server 102 via a network. Client 101 refers to a device used by a user, such as a smartphone or computer. In this solution, a user can interact with server 102 via client 101 to report data. At the same time, the user can receive multiple parameters required for sampling rate calculation on client 101. The sampling rate calculation can be performed on client 101 or server 102, without limitation.
[0020] The server 102 refers to a remote computer used to process a large number of computing tasks and store data. In this solution, the server 102 first classifies the first data based on the source of the first data and the abnormal value of the first data; then, based on the mapping relationship between the sampling rate preset value and the second data of each category, as well as the data density and timeliness index of the second data of each category, the first sampling rate of the second data of each category is determined; then, for the first sampling rate of the second data of each category, if it is detected that the first sampling rate is within the sampling rate range, the second sampling rate of the second data is determined based on the first sampling rate, the first and second quantities of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period; finally, based on the second sampling rate, the second data is sampled. In addition, the server 102 will regularly count the total amount of reported data, the number of client devices, etc., and feed back the upper and lower limits obtained based on the statistical results to each client 101.
[0021] Based on this, the present application provides a data processing method, device, storage medium and program product, which are described in detail below with reference to the accompanying drawings.
[0022] See also Figure 2 , Figure 2 This is a flow chart of a data processing method provided in an embodiment of the present application. Figure 2 As shown, the method includes the following steps: S201 , classifying the first data based on a source of the first data or a request status code of the first data to obtain second data of each category.
[0023] Among them, the execution subject of this method can be Figure 1 Data processing system 100 is shown with server 102 .
[0024] Among them, the first data is a set of original data from the first device that needs to be classified and processed. The first data can be all the data to be reported by the first device, or the first data can be a part of the data generated by the first device, such as data within a specific time period, data of a specific type, etc.
[0025] Among them, when classifying the first data based on the source of the first data and the abnormal value of the first data, the first data can be divided into page loading data collected by the front-end monitoring SDK, user click data captured by the event listening program, and request log data collected by the server NGINX / Access log according to the source. At the same time, the first data can be further classified based on the request status code. For example, for data with status code = 200-299, it is classified as a successful request, for data with status code = 400-499, it is classified as a client error, for data with status code = 500-599, it is classified as a server error, and for data that takes more than a certain time and has status code = 0, it is classified as a timed-out request. Among them, types such as timed-out requests, server errors, and client errors can be uniformly classified as abnormal logs.
[0026] Among them, for data in the observable field, it can be classified according to the behavior or process of data generation, such as classifying data into page loading data, user click data and request log data; it can be classified according to the nature and purpose of the data, such as abnormal log data specifically used to record abnormal situations that occur during system operation; it can also be classified according to the integrity of the data, that is, according to whether the data is complete during the collection, transmission or storage process. The standard for classifying the first data is not restricted here.
[0027] For example, when classifying the first data, it can be divided into the following data: page load, user click, request log, exception log, and incomplete data. At the same time, page load and request log can be further classified. Specifically, page load can be classified into slow load, normal load, and abnormal page, and request log can be classified into normal request, slow request, and abnormal request.
[0028] S202 : Determine a first sampling rate of the second data of each category based on a mapping relationship between a preset sampling rate value and the second data of each category, and a data density and a timeliness index of the second data of each category.
[0029] The range of the first sampling rate is 0-1. When the first sampling rate is 1, it indicates that all data needs to be sampled. When the first sampling rate is 0, it indicates that no sampling is required.
[0030] In one possible implementation, see Figure 3, Figure 3 This is a flow chart of determining a first sampling rate of second data of each category provided by an embodiment of the present application, such as Figure 3 As shown, based on the mapping relationship between the preset sampling rate value and the second data of each category, and the data density and timeliness index of the second data of each category, determining the first sampling rate of the second data of each category includes the following steps: S301 : Determine a preset sampling rate value for the second data of each category based on a mapping relationship between the preset sampling rate value and the second data of each category.
[0031] Among them, the sampling rate preset values of multiple categories of second data can be set directly by the user, that is, the mapping relationship between multiple sampling rate preset values and the second data of each category is set in advance. For example, the mapping relationship between multiple categories of second data and the corresponding sampling rate preset values are: slow loading (sampling rate preset value is 0.8), normal loading (sampling rate preset value is 0.6), abnormal page (sampling rate preset value is 1), user click (sampling rate preset value is 0.8), normal request (sampling rate preset value is 0.5), slow request (sampling rate preset value is 0.8), abnormal request (sampling rate preset value is 1), abnormal log (sampling rate preset value is 1), incomplete data (sampling rate preset value is 0).
[0032] Among them, the sampling rate preset values of multiple categories of second data can also be obtained through a rule model written by a rule engine or other means, and the priorities of different devices or data can be set through rules. Specifically, the user can set the sampling rate priority of the same category of data for different client devices. For example, when collecting page performance logs, the user focuses on the performance issues of model A mobile phone, and the sampling rate of this model can be increased accordingly. For example, the sampling rate priority can be set to: level: 1, model: A, sampling rate preset value 0.8; level: 2, model: all, sampling rate preset value 0.3. In this way, when determining the sampling rate preset value of the page performance log, it is necessary to first detect whether the model of the mobile phone connected to the client device is A. If it is A, the sampling rate preset value is 0.8, otherwise the sampling rate preset value is 0.3.
[0033] For example, after obtaining a preset sampling rate using a rule model written using a rule engine, the user may request to retain all abnormal data. Therefore, abnormal requests are given the highest value, while normal requests are only used for reference. In this case, rules can be written for the request log sampling rate to adjust the sampling rate. Specifically, the rules may be as follows: Rule 1: If (the request status code is 4XX or 5XX), then the value is 1; otherwise, the sampling rate is 0.5. Rule 2: If (the request takes longer than or equal to 300ms), then the sampling rate is 0.8.
[0034] Among them, the preset values of the sampling rates of multiple categories of second data can also be obtained through a data value assessment model. The data value assessment model is a model that calculates the value of data based on rules and specific evaluation dimensions, including an evaluation dimension layer, a rule engine layer, and an application output layer. The data value assessment model can be a multilayer perceptron (MLP). The multilayer perceptron is a feedforward artificial neural network composed of multiple neurons (neural nodes). These neurons are arranged in a hierarchical structure, including an input layer, a hidden layer, and an output layer. The neurons between layers are connected by weights, and information is propagated forward from the input layer to the output layer in sequence without feedback connection.
[0035] Specifically, the evaluation dimension layer extracts value assessment factors from multiple dimensions, such as business importance and data quality (completeness, accuracy, and timeliness), to construct a basic indicator system for quantifying data value. The rule engine layer performs weighted calculations on each dimension's indicators using pre-set static weights or dynamically adjusted rules (such as anomaly triggering mechanisms), achieving standardized quantitative scoring of data value. The application output layer assigns data a high, medium, or low value rating based on the value rating results. This solution uses a data value assessment model to identify dimensions such as data importance to users, data density, and data completeness and accuracy. Corresponding rules, such as assigning different weights or scoring criteria to different data types, are then developed to quantify data value and convert this value into a sampling rate. Specifically, when calculating the first sampling rate based on the data value assessment model, factors influencing the data sampling rate include, but are not limited to: 1. Data importance to users: The importance of different types of data (e.g., abnormal data and normal data) to users is determined, with abnormal data being given higher importance and normal data being used for reference only. 2. Data density reporting: Based on data statistics, scoring ranges are set based on the quantity and ranking of each type of data, reducing the value of data with large volumes but low user-labeled value. 3. Data integrity and accuracy: Determine whether the data metric value is empty, erroneous, or historical data (for example, the data request time exceeds 7 days). If the metric value is empty, erroneous, or the value is historical data from 1 month ago, it can be directly set to 0. In this case, the first sampling rate of the corresponding data is also 0.
[0036] S302 : Determine a sampling rate adjustment value of the second data of each category based on the data density and timeliness index of the second data of each category.
[0037] The data density of the second data refers to the amount of data generated per unit time. In practice, a large amount of data does not necessarily mean high data value. If a certain type of data is generated in large quantities per unit time, but the user does not need too much, the sampling rate adjustment value needs to be determined based on the data density, thereby adjusting the sampling rate.
[0038] Among them, some devices may generate a large amount of log data in a short period of time. These log data only record some routine operation information. Although the data volume per unit time is large, it may not be of great value to the user. By counting the data volume of each category of data and ranking them, different scoring intervals are set to determine the sampling rate adjustment value of the second data of each category. Specifically, the user can set the sampling rate adjustment value of the second data of the category with the highest data volume to obtain a lower sampling rate.
[0039] For example, if the amount of data in a normal request is very large and the preset sampling rate is 0.5, the sampling rate for this type of data needs to be further reduced due to the high density of this type of data. The rule can be set as follows: if the preset sampling rate is less than 0.5 and the data volume accounts for more than 50%, the sampling rate adjustment value is the preset sampling rate * the data volume percentage. When the first sampling rate is equal to the preset sampling rate minus the sampling rate adjustment value, the first sampling rate = the preset sampling rate * (1 - data volume percentage). This formula can be adjusted according to actual conditions and does not restrict the sampling rate adjustment value or the calculation formula for the first sampling rate.
[0040] Among them, timeliness indicators are used to evaluate the validity of data or requests in the time dimension, including but not limited to time proximity and response speed. Time proximity refers to the closeness of the timestamp of the data or request to the current time, measuring whether the data is newly generated or updated. Response speed refers to the time from issuing a data request to receiving the response result, measuring the efficiency of the system in processing requests. Data with abnormal response speed and low time freshness can be considered invalid data. For example, to avoid sampling invalid data, you can write the following rules: If (response speed (i.e. request time) <= 0 or time freshness (i.e. request time) is 7 days ago), then the sampling rate adjustment value is set so that the first sampling rate of the data that meets the rule is 0.
[0041] S303: Determine a first sampling rate of the second data of each category based on the preset sampling rate value and the sampling rate adjustment value.
[0042] Among them, there can be multiple ways to combine the sampling rate preset value and the sampling rate adjustment value to obtain the first sampling rate of the second data of each category. For example, the sampling rate preset value and the sampling rate adjustment value can be given different weights respectively through the weighted average method, and then the weighted average value is calculated as the first sampling rate of this type of data; other suitable methods can also be used, such as weighted summation or combined judgment based on different scoring levels, etc., and the method of determining the first sampling rate of the second data of each category is not restricted here.
[0043] It can be seen that in this example, the basic sampling rate is determined through a preset mapping relationship, and the adjustment value is dynamically calculated based on data density (reflecting the distribution of data volume) and timeliness index (measuring time effectiveness), and finally the first sampling rate of the data is comprehensively obtained. This can not only comprehensively consider the value characteristics of data at different levels and avoid the one-sidedness of a single standard, but also more accurately quantify the first sampling rate of the data through a refined evaluation process.
[0044] S203, for the first sampling rate of the second data of each category, if it is detected that the first sampling rate is within the sampling rate range, determine the second sampling rate of the second data based on the first sampling rate, the first and second quantities of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period.
[0045] Among them, the sampling rate interval is a numerical interval preset by the user. Preferably, since the sampling rate range is 0-1, the range of the first interval is set to be greater than 0 and less than 1. When the first sampling rate of the second data is greater than 0 and less than 1, the sampling rate of the second data is adjusted.
[0046] Among them, the first time period and the second time period are time intervals with a fixed length. The lengths of the first time period and the second time period can be the same or different. The lengths of the first time period and the second time period can be 8 hours, 24 hours or 72 hours. There is no restriction on the lengths of the first time period and the second time period.
[0047] Preferably, the first time period and the second time period are both 24 hours in length, and the first time period is the previous time period of the second time period.
[0048] Among them, the first quantity of the second data obtained in the first time period refers to the total amount of second data reported by all client devices 101 devices obtained by the server 102 of the data processing system 100 within the first time period, the second quantity of the second data obtained in the first time period refers to the total amount of second data reported by specific client devices 101 devices obtained by the server 102 of the data processing system 100 within the first time period, and the third quantity of the second data obtained in the second time period refers to the total amount of second data reported by specific client devices 101 devices obtained by the server 102 of the data processing system 100 within the second time period.
[0049] In one possible implementation, see Figure 4 , Figure 4 This is a flow chart of determining a second sampling rate of second data provided by an embodiment of the present application, such as Figure 4As shown, determining a second sampling rate of the second data according to the first sampling rate, the first quantity and the second quantity of the second data acquired in the first time period, and the third quantity of the second data acquired in the second time period includes the following steps: S401, determining a first sampling rate adjustment coefficient according to the first number, the second number, and the third number; S402, determining a second sampling rate adjustment coefficient according to the second number and the third number; S403: Determine the second sampling rate according to the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient.
[0050] Among them, the first quantity refers to the total amount of second data reported by all client devices 101 devices obtained by the server 102 of the data processing system 100 within the first time period; the second quantity refers to the total amount of second data reported by specific client devices 101 devices obtained by the server 102 of the data processing system 100 within the first time period; the third quantity refers to the total amount of second data reported by specific client devices 101 devices obtained by the server 102 of the data processing system 100 within the second time period.
[0051] The first sampling rate adjustment coefficient and the second sampling rate adjustment coefficient are both used to adjust the first sampling rate to obtain the second sampling rate.
[0052] When determining the second sampling rate of the second data in the specific device based on the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient, a weighted sum or other mathematical operation can be performed on the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient to obtain a final value, which is then matched with a preset sampling rate range to determine an appropriate sampling rate. The sampling rate range can be directly preset by the user or calculated based on historical data.
[0053] It can be seen that in this example, by hierarchically calculating the adjustment coefficient based on the total amount of client data (first quantity) obtained by the server in the first time period, the amount of specific client data (second quantity) and the amount of specific client data (third quantity) in the second time period, the global data collection efficiency is combined with the data characteristics of specific devices to achieve dynamic optimization of the sampling rate.
[0054] In one possible implementation, see Figure 5 , Figure 5 This is a flow chart of determining a first sampling rate adjustment coefficient provided by an embodiment of the present application, such as Figure 5 As shown, determining a first sampling rate adjustment coefficient according to the first number, the second number, and the third number includes the following steps: S501: predicting a predicted value of the quantity of the second data acquired in the second time period based on the first quantity; S502: Calculate the average of the number of the acquired second data according to the second number and the third number; S503: Determine the first sampling rate adjustment coefficient according to the predicted value of the quantity of the second data and the average value of the quantity of the second data.
[0055] The predicted value of the second data quantity is the reference data quantity when a single client 101 device in communication with the server 102 reports the second data. The mean value of the second data quantity refers to the average value of the data quantity of this category reported by a specific device over a period of time.
[0056] For example, there are five client devices 101, and the amount of second data reported is 10, 12, 8, 15, and 13, respectively. The total amount of data reported by all client devices 101 is: 10 + 12 + 8 + 15 + 13 = 58. At this point, there are five client devices 101, so the predicted amount of second data = 58 ÷ 5 = 11.6, meaning the reference amount of second data reported by a single client device 101 is 11.6. This number can be rounded up, meaning the reference amount of second data reported by a single client device 101 is 12.
[0057] Optionally, in addition to directly dividing the first quantity by the number of devices to obtain the predicted value of the quantity of the second data, the predicted value of the quantity of the second data can also be obtained by obtaining the first quantile of the first quantity.
[0058] The quantile is the value used to divide a set of data into equal parts after arranging it in ascending order. The first quantile can be set or modified by the user. For example, if the first quantile is 75%, it means that 75% of the data is less than or equal to this value, and 25% of the data is greater than this value.
[0059] Specifically, first, the quantity of second data reported by all client 101 devices is sorted, and then the corresponding value is determined based on the ratio of the quantity of data and the first quantile, that is, the predicted value of the quantity of second data, that is, the reference data volume of a single client 101 device when reporting the second data.
[0060] For example, let's assume the first quantile is the 75th quantile. Eight client devices 101 report the number of second data items, 10, 12, 15, 18, 20, 22, 25, and 30, respectively. First, sort these numbers from smallest to largest: 10, 12, 15, 18, 20, 22, 25, and 30. Calculate the position of the 75th quantile: 8 × 0.75 = 6, meaning it lies between the sixth and seventh numbers. The sixth number is 22, and the seventh number is 25. Therefore, the 75th quantile can be calculated using linear interpolation: 22 + (25 − 22) × (6 − 5) = 22 + 3 × 1 = 25. Therefore, the predicted number of second data items, or the reference amount of second data items reported by a single client device 101, is 25.
[0061] The first sampling rate adjustment coefficient is determined by performing a specific mathematical calculation (e.g., calculating a ratio or difference between the predicted value of the second data quantity and the average value of the second data quantity) on the predicted value and the average value of the second data quantity. The first sampling rate adjustment coefficient is used to measure the weight of the data reported by a specific device under the current circumstances relative to its historical performance and overall expectations.
[0062] Preferably, the first sampling rate adjustment coefficient is obtained by calculating the quotient of the predicted value of the amount of the second data and the mean value of the amount of the second data.
[0063] It can be seen that in this example, the first sampling rate adjustment coefficient is calculated by combining the predicted value of the data reported by a single client device and the average historical data of a specific device. This can not only estimate the future data volume based on the current data trend, but also refer to the historical stable performance of the device to ensure the stability of data sampling, avoid sampling rate imbalance caused by sudden fluctuations or historical deviations, and achieve dynamic optimization of data sampling.
[0064] In a possible implementation, determining a second sampling rate adjustment coefficient according to the second number and the third number includes: calculating a quotient of the second number and the third number to obtain the second sampling rate adjustment coefficient.
[0065] The second sampling rate adjustment coefficient may also be calculated by subtracting 1 from the quotient of the second quantity and the third quantity. The second sampling rate adjustment coefficient reflects a changing trend of the second data reporting amount of a specific device.
[0066] It can be seen that in this example, the second sampling rate adjustment coefficient is determined by calculating the quotient of the second number and the third number, directly quantifying the change ratio of the data volume of the specific device in the previous and next time periods, realizing sampling rate adjustment and improving the efficiency of sampling rate optimization.
[0067] In one possible implementation, see Figure 6 , Figure 6This is a flow chart of determining the second sampling rate provided by an embodiment of the present application, such as Figure 6 As shown, determining the second sampling rate according to the first sampling rate, the first sampling rate adjustment coefficient and the second sampling rate adjustment coefficient includes the following steps: S601: Calculate a first product of a first weight and the first sampling rate.
[0068] Among them, the first sampling rate is represented by the letter P, the first weight can be set or changed by the user, and the first weight is represented by α. The first product at this time is expressed as .
[0069] S602: Calculate a second product of a second weight and the first sampling rate adjustment coefficient.
[0070] The first sampling rate adjustment coefficient is represented by the letter C, and the second weight can be set or changed by the user. The second weight is represented by β. The second product is expressed as .
[0071] S603: Calculate a third product of a third weight and the second sampling rate adjustment coefficient.
[0072] The second sampling rate adjustment coefficient is represented by the letter H, and the third weight can be set or changed by the user. The third weight is represented by γ. The third product is expressed as .
[0073] S604: Calculate the difference between the first product and the second product.
[0074] S605: Calculate the sum of the difference and the third product to obtain the second sampling rate.
[0075] The calculation formula of the second sampling rate is as follows: ; Wherein, f is the sampling rate to be calculated, that is, the second sampling rate.
[0076] Among them, when calculating the sampling rate f, the following factors need to be considered: 1. Define constraints: Define the constraint range of the sampling rate f based on the total amount of resources, data classification and other constraints. For example: 0≤f≤1.
[0077] 2. Optimization algorithm selection: The genetic algorithm is selected as the optimization algorithm. The genetic algorithm gradually optimizes the sampling rate by simulating the process of natural selection.
[0078] 3. Optimization process: Initialize the population: Randomly generate a set of sampling rate values as the initial population.
[0079] Fitness function: Calculate the fitness value of each sampling rate according to the objective function.
[0080] Selection operation: Select an excellent sampling rate value based on the fitness value and enter the next generation.
[0081] Crossover operation: Perform a crossover operation on the selected sampling rate value to generate a new sampling rate value.
[0082] Mutation operation: Perform mutation operation on the new sampling rate value to increase the diversity of the population.
[0083] Termination condition: Repeat the selection, crossover, and mutation operations until the preset number of iterations is reached or the fitness value converges.
[0084] 4. Get the result application: The solution of the objective function is used as the optimized sampling rate, which can be applied to the actual data collection process to dynamically adjust the frequency or quantity of data collection.
[0085] It can be seen that in this example, the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient are integrated by constructing a weighted combination formula (second sampling rate = α×P - β×C + γ×H). This not only retains the basic strategy of the first sampling rate, but also flexibly controls the influence of each adjustment coefficient through weight parameters to achieve precise adjustment of the sampling rate.
[0086] In one possible implementation, the method further includes: If it is detected that the first sampling rate is not within the sampling rate interval, the second data is sampled based on the first sampling rate.
[0087] Specifically, when determining the sampling rate of the second data of the first category in the first device, if the first sampling rate is not within the sampling rate range, the second data is directly sampled using the first sampling rate. For example, the sampling rate range is greater than 0 and less than 1. At this time, the first sampling rate of the abnormal log data is 1, which is not within the first interval. At this time, it is directly determined that the sampling rate of the abnormal log data of the first device is equal to the first sampling rate of the abnormal log data, that is, the sampling rate is equal to 1.
[0088] It can be seen that in this example, when the first sampling rate is not within the sampling rate range, directly sampling the second data based on the sampling rate can avoid resource consumption due to invalid adjustment and ensure the efficiency of data sampling.
[0089] S204: Sample the second data based on the second sampling rate.
[0090] The second sampling rate is the sampling frequency of the second data, and the second data is obtained from a single client 101 device according to the sampling frequency. The collected data will be further processed according to the system design and may be stored in a database for subsequent query and analysis; it may also be directly transmitted to other modules for real-time processing, such as for controlling the operating status of related devices or for data visualization, so that users can intuitively understand the changing trends of the data.
[0091] It can be seen that in this example, the data is classified in combination with the data source and the outlier characteristics, the initial sampling rate is accurately determined based on the data density, timeliness index and the preset mapping relationship, and the second sampling rate is dynamically adjusted based on the sampling rate interval and the amount of data in different time periods, thereby realizing refined management of data sampling, fully considering the differences in data characteristics, effectively avoiding over-sampling or under-sampling, achieving balanced data processing, and improving sampling efficiency.
[0092] See also Figure 7 , Figure 7 is a structural diagram of a data processing device provided in an embodiment of the present application, such as Figure 7 As shown, the data processing device 700 includes: A classification module 701 is configured to classify the first data based on a source of the first data or a request status code of the first data to obtain second data of each category; A determination module 702 is configured to determine a first sampling rate for the second data of each category based on a mapping relationship between a preset sampling rate value and the second data of each category, and a data density and a timeliness index of the second data of each category; and a first sampling rate for the second data of each category, and if it is detected that the first sampling rate is within the sampling rate interval, determining a second sampling rate for the second data based on the first sampling rate, a first amount and a second amount of the second data acquired in a first time period, and a third amount of the second data acquired in a second time period; The sampling module 703 is configured to sample the second data based on the second sampling rate.
[0093] In one possible implementation, in terms of determining the first sampling rate of the second data of each category based on the mapping relationship between multiple sampling rate preset values and the second data of each category, and the data density and timeliness index of the second data of each category, the determination module 702 is specifically used to: determine the sampling rate preset value of the second data of each category based on the mapping relationship between the sampling rate preset value and the second data of each category; determine the sampling rate adjustment value of the second data of each category based on the data density and timeliness index of the second data of each category; and determine the first sampling rate of the second data of each category based on the sampling rate preset value and the sampling rate adjustment value.
[0094] In one possible implementation, in terms of determining the second sampling rate of the second data based on the first sampling rate, the first quantity and second quantity of the second data obtained in the first time period, and the third quantity of the second data obtained in the second time period, the determination module 702 is specifically used to: determine a first sampling rate adjustment coefficient based on the first quantity, the second quantity, and the third quantity; determine a second sampling rate adjustment coefficient based on the second quantity and the third quantity; determine the second sampling rate based on the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient.
[0095] In one possible implementation, in determining the first sampling rate adjustment coefficient based on the first quantity, the second quantity and the third quantity, the determination module 702 is specifically used to: predict the predicted value of the quantity of the second data obtained in the second time period based on the first quantity; calculate the average of the quantity of the second data obtained based on the second quantity and the third quantity; and determine the first sampling rate adjustment coefficient based on the predicted value of the quantity of the second data and the average of the quantity of the second data.
[0096] In a possible implementation, in determining the second sampling rate adjustment coefficient according to the second number and the third number, the determination module 702 is specifically configured to calculate a quotient of the second number and the third number to obtain the second sampling rate adjustment coefficient.
[0097] In one possible implementation, in determining the second sampling rate based on the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient, the determination module 702 is specifically used to: calculate a first product of a first weight and the first sampling rate; calculate a second product of a second weight and the first sampling rate adjustment value; calculate a third product of a third weight and the second sampling rate adjustment value; calculate a difference between the first product and the second product; and calculate the sum of the difference and the third product to obtain the second sampling rate.
[0098] In a possible implementation, the sampling module 703 is further configured to: if it is detected that the first sampling rate is not within the sampling rate interval, sample the second data based on the first sampling rate.
[0099] It is worth noting that the specific functional implementation of the data processing device 700 is shown in the above Figure 2 In the description of the data processing method shown in FIG. 1 , for example, the classification module 701 is used to implement the relevant content of executing S201, the determination module 702 is used to implement the relevant content of executing S202-S203, and the sampling module 703 is used to implement the relevant content of executing S204. Each unit or module in the data processing device 700 can be individually or entirely combined into one or more other units or modules to form a structure, or one (or more) of the units or modules can be further divided into multiple functionally smaller units or modules to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units or modules are divided according to logical functions. In actual applications, the functions of one unit (or module) are implemented by multiple units (or modules), or the functions of multiple units (or modules) are implemented by one unit (or module).
[0100] According to the description of the above method embodiment and related device embodiment, please refer to Figure 8 , Figure 8 is a structural diagram of a data processing device provided in an embodiment of the present application, Figure 8 The data processing device 800 shown includes a processor 801 , a memory 802 , a communication interface 803 , and a bus 804 . The processor 801 , the memory 802 , and the communication interface 803 are communicatively connected to each other via the bus 804 .
[0101] Optionally, the memory 802 is a ROM, a static storage device, a dynamic storage device or a RAM.
[0102] The memory 802 can store executable program codes. When the executable program codes stored in the memory 802 are executed by the processor 801, the processor 801 and the communication interface 803 are used to execute the program codes. Figure 2 The various steps of the data processing method of the embodiment are shown.
[0103] The processor 801 adopts a general-purpose CPU, a microprocessor, an application-specific integrated circuit ASIC, a GPU or one or more integrated circuits to execute relevant programs to perform the data processing method of the embodiment of the method of the present application.
[0104] Processor 801 can also be an integrated circuit chip with signal processing capabilities. During implementation, each step of the data processing method of the present application can be completed by hardware integrated logic circuits in processor 801 or by software instructions. Optionally, processor 801 is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor is a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The optional software module is located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media well-known in the art. The storage medium is located in memory 802. Processor 801 reads information in memory 802 and, in conjunction with its hardware, completes the functions required to be performed by the modules included in a data processing device 700 of the embodiments of the present application, or executes the data processing method of the method embodiments of the present application.
[0105] The communication interface 803 uses, for example but not limited to, a transceiver or other transceiver-related device.
[0106] The bus 804 may include a path for transmitting information between various components of the data processing device 800 (eg, the memory 802 , the processor 801 , and the communication interface 803 ).
[0107] It should be noted that although Figure 8 The data processing device 800 shown only shows a memory, a processor, and a communication interface. However, in the specific implementation process, those skilled in the art should understand that the data processing device 800 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the data processing device 800 may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the data processing device 800 may also include only the devices necessary to implement the embodiments of the present application, and does not necessarily include Figure 8 All devices shown in .
[0108] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program for electronic data exchange. The computer program includes execution instructions, and the execution instructions are used to execute part or all of the steps of any one of the data processing methods described in the above-mentioned data processing method embodiments. The above-mentioned computer includes an electronic client device.
[0109] An embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and the computer program is operable to enable a computer to perform part or all of the steps of any data processing method recorded in the above method embodiments. The computer program product can be a software installation package.
[0110] It should be noted that for the sake of simplicity, the embodiments of any of the aforementioned data processing methods are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0111] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of a data processing method, device, storage medium, and program product of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, based on the idea of a data processing method, device, storage medium, and program product of the present application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting the present application.
[0112] The present application is described with reference to the flowcharts and / or block diagrams of the methods, hardware products, and computer program products of the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0113] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The memory may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0114] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps. The fact that certain measures are recited in different dependent claims does not mean that these measures cannot be combined to produce good results.
[0115] Those skilled in the art will understand that all or part of the steps in the various methods of the method embodiments of any of the above-mentioned data processing methods can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0116] It can be understood that any product that is controlled or configured to execute the processing method of the flowchart described in an embodiment of a data processing method of the present application, such as the device and computer program product of the above flowchart, falls within the scope of the related products described in the present application.
[0117] Obviously, those skilled in the art may make various modifications and variations to the data processing method, device, storage medium, and program product provided herein without departing from the spirit and scope of the present application. Thus, if such modifications and variations fall within the scope of the claims of the present application and their equivalents, the present application is intended to encompass such modifications and variations.
Claims
1. A data processing method, characterized in that: The method comprises: Classify the first data based on a source of the first data or a request status code of the first data to obtain second data of each category; Determining a first sampling rate of the second data of each category based on a mapping relationship between a preset sampling rate value and the second data of each category, and a data density and a timeliness index of the second data of each category; For each category of the first sampling rate of the second data, if it is detected that the first sampling rate is within the sampling rate interval, determining a second sampling rate of the second data according to the first sampling rate, a first quantity and a second quantity of the second data acquired in the first time period, and a third quantity of the second data acquired in the second time period; The second data is sampled based on the second sampling rate.
2. The method according to claim 1, wherein The determining, based on a mapping relationship between a preset sampling rate value and the second data of each category, and a data density and a timeliness index of the second data of each category, of the first sampling rate of the second data of each category includes: determining a preset sampling rate value for the second data of each category based on a mapping relationship between the preset sampling rate value and the second data of each category; determining a sampling rate adjustment value of the second data of each category based on a data density and a timeliness index of the second data of each category; The first sampling rate of the second data of each category is determined based on the preset sampling rate value and the sampling rate adjustment value.
3. The method according to claim 1, wherein Determining a second sampling rate of the second data according to the first sampling rate, the first quantity and the second quantity of the second data acquired in the first time period, and the third quantity of the second data acquired in the second time period includes: determining a first sampling rate adjustment coefficient according to the first number, the second number, and the third number; determining a second sampling rate adjustment coefficient according to the second number and the third number; The second sampling rate is determined according to the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient.
4. The method according to claim 3, wherein The determining a first sampling rate adjustment coefficient according to the first number, the second number, and the third number includes: Predicting a predicted value of the quantity of the second data acquired in the second time period based on the first quantity; Calculating an average of the quantities of the second data obtained based on the second quantity and the third quantity; The first sampling rate adjustment coefficient is determined according to the predicted value of the quantity of the second data and the average value of the quantity of the second data.
5. The method according to claim 3, wherein The determining a second sampling rate adjustment coefficient according to the second number and the third number includes: The quotient of the second number and the third number is calculated to obtain the second sampling rate adjustment coefficient.
6. The method according to claim 3, wherein The determining the second sampling rate according to the first sampling rate, the first sampling rate adjustment coefficient, and the second sampling rate adjustment coefficient includes: Calculating a first product of a first weight and the first sampling rate; Calculating a second product of a second weight and the first sampling rate adjustment coefficient; calculating a third product of a third weight and the second sampling rate adjustment coefficient; calculating a difference between the first product and the second product; The sum of the difference and the third product is calculated to obtain the second sampling rate.
7. The method according to any one of claims 1 to 6, wherein: The method further comprises: If it is detected that the first sampling rate is not within the sampling rate interval, the second data is sampled based on the first sampling rate.
8. A data processing device, characterized in that: The device comprises: A memory, a processor, and an executable program code stored in the memory and capable of running on the processor, wherein the processor executes the steps of the data processing method according to any one of claims 1 to 7 when executing the executable program code.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores executable program code, which includes execution instructions for executing the steps of the data processing method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product includes a computer program, and the computer program is used to enable a computer to execute the steps of the data processing method according to any one of claims 1 to 7.