Data processing method and device, computer equipment and storage medium
By dynamically segmenting the data flow and calculating indicators in stages, the calculation delay and resource waste problems in the processing of massive data in the existing technology are solved, the flexible and changeable business needs are met, and the computing efficiency and accuracy of results are improved.
Patent Information
- Application Number
- CN202510238054.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
Existing real-time indicator computing technology faces the problems of computing delays, waste of resources and the inability to meet flexible and changeable business needs when processing massive data.
The data flow is dynamically divided by configuring configurable time window parameters, adjust the calculation strategy according to the actual situation of the data flow, and does not rely on a large amount of cache storage data. The intermediate value of the indicator is first calculated and then the result value of the indicator is then calculated in stages.
It avoids the problem of computing delays and resource waste during peak data flow and meets flexible and changeable business needs, reduces the pressure of processing massive data, reduces the computational complexity and memory usage, and ensures the accuracy and traceability of the results.
Smart Images

Figure CN120179680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of stream computing, and in particular, to a data processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the rapid development of Internet digital technology, the amount of data has shown explosive growth. How to calculate and extract valuable indicator information from massive data in real time and accurately has become a key challenge in the field of big data processing. The traditional batch processing computing mode can no longer meet the real-time requirements. Therefore, real-time computing technology has emerged. Real-time computing technology can perform instant analysis and processing on data streams, providing users with real-time data analysis results, which is crucial for business scenarios that require quick responses.
[0003] However, the existing real-time indicator calculation technologies face many challenges when processing massive data. First, the traditional indicator calculation method using a fixed time dimension cannot flexibly adapt to the changes in data streams, and it is difficult to dynamically adjust the calculation strategy, resulting in calculation delays during data stream peaks, resource waste during valleys, and inability to meet the flexible and changeable business needs; while the dynamic time window calculation method based on cache components is limited by the cache capacity and is difficult to support the indicator storage of massive data and ultra-long time windows, and cannot meet the indicator calculation requirements in large data volume scenarios; if an indicator calculation system is built based on a data warehouse, the data preprocessing method and the way of aggregating large time windows will lead to a decline in system performance, and the micro-batch computing mode cannot meet the rate requirements of data stream updates in terms of performance. Summary of the Invention
[0004] Based on this, it is necessary to provide a data processing method, apparatus, computer device, and storage medium for the above technical problems to solve at least one of the problems existing in the above prior art.
[0005] In a first aspect, a data processing method is provided, including:
[0006] Configuring indicator parameters, where the indicator parameters include time window parameters and indicator calculation rules;
[0007] Collecting a data stream and splitting the data stream according to the time window parameters to obtain a plurality of sub-data streams;
[0008] Calculating the sub-data streams corresponding to different time windows based on the indicator calculation rules to obtain indicator intermediate values that meet preset conditions under different time windows;
[0009] Calculating the indicator intermediate values based on the indicator calculation rules to obtain an indicator result value.
[0010] In one embodiment, splitting the data stream according to the time window parameter to obtain a plurality of sub-data streams includes:
[0011] Extracting a plurality of identical data streams to be split from the data stream, the number of the data streams to be split being the same as the number of the time windows;
[0012] Splitting the data stream to be split according to the time window parameter to obtain a plurality of sub-data streams.
[0013] In one embodiment, calculating the data stream to be split according to the index calculation rule in different time windows to obtain the index intermediate value corresponding to each time window includes:
[0014] In each time window, grouping the corresponding sub-data streams according to a preset field respectively;
[0015] Aggregating each group of sub-data streams;
[0016] Statistically analyzing the aggregated data according to the index calculation rule to obtain the real-time value of each index in the current time window as the index intermediate value.
[0017] In one embodiment, before calculating the data stream to be split according to the index calculation rule in different time windows to obtain the index intermediate value corresponding to each time window, it further includes:
[0018] Determining whether there are missing values or null values in the sub-data streams corresponding to different time windows respectively;
[0019] If so, supplementing the missing values or null values.
[0020] In one embodiment, calculating the sub-data streams corresponding to different time windows according to the index calculation rule to obtain the index intermediate value meeting the preset conditions in different time windows includes:
[0021] Obtaining the index intermediate value through ordinary statistics, normal number, suspected number or risk number calculation methods in different time windows, the index intermediate value including the unique value of the combination of the index unique identifier and the user account, the total amount of characteristic data obtained by statistics, and the end time identifier of the time window; or
[0022] Obtaining the index intermediate value through correlation calculation or historical correlation calculation methods in different time windows, the index intermediate value including the unique value of the combination of the index unique identifier, the user account, the main attribute and the subordinate attribute, the set of calculated field values, and the end time identifier of the time window; or
[0023] The intermediate value of the indicator is obtained through the normal rate, suspected rate or risk rate calculation method within different time windows. The intermediate value of the indicator includes a unique value composed of an indicator unique identifier and a user account, a total value, a characteristic data volume, and an end time identifier of the time window.
[0024] In one embodiment, calculating the indicator result value based on the indicator calculation rule for the intermediate value of the indicator includes:
[0025] Constructing a retrieval condition;
[0026] Based on the retrieval condition, retrieving the intermediate value of the indicator to obtain complete intermediate result information;
[0027] Statistically calculating the intermediate result information to obtain the indicator result value.
[0028] In one embodiment, the indicator calculation rule includes any one or any combination of ordinary statistics, correlation calculation, historical correlation calculation, normal number, suspected number, risk number, normal rate, suspected rate, and risk rate.
[0029] In a second aspect, a data processing device is provided, including:
[0030] An indicator parameter configuration unit for configuring indicator parameters, where the indicator parameters include time window parameters and an indicator calculation rule;
[0031] A data stream segmentation unit for collecting a data stream and segmenting the data stream according to the time window parameters to obtain a plurality of sub-data streams;
[0032] An intermediate value determination unit of the indicator for calculating the corresponding sub-data stream within different time windows based on the indicator calculation rule to obtain an intermediate value of the indicator that meets the preset conditions under different time windows;
[0033] An indicator result value determination unit for calculating the intermediate value of the indicator based on the indicator calculation rule to obtain an indicator result value.
[0034] In a third aspect, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the data processing method as described above are implemented.
[0035] In a fourth aspect, a readable storage medium is provided. The readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the data processing method as described above are implemented.
[0036] The above data processing method, apparatus, computer device, and storage medium. The implementation of the method includes: configuring metric parameters, where the metric parameters include time window parameters and metric calculation rules; collecting data streams, and segmenting the data streams according to the time window parameters to obtain a number of sub-data streams; calculating the sub-data streams corresponding to different time windows based on the metric calculation rules to obtain metric intermediate values that meet preset conditions for different time windows; and calculating the metric result value based on the metric intermediate values. In the embodiments of the present application, by dynamically segmenting the data stream through configurable time window parameters, the calculation strategy can be adjusted according to the actual situation of the data stream, without relying on a large amount of cache to store data, avoiding the problems of calculation delay during data stream peaks and resource waste during valleys, and meeting the flexible and changeable business needs. Moreover, the phased calculation method of first calculating the metric intermediate value and then calculating the metric result value based on the intermediate value reduces the pressure of processing a large amount of data at one time, reduces the calculation complexity and memory occupancy, and ensures the accuracy and traceability of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0038] Figure 1 is a flowchart of a data processing method in an embodiment of the present invention;
[0039] Figure 2 is a structural diagram of a data processing apparatus in an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0042] In one embodiment, as Figure 1 shown, a data processing method is provided, including the following steps:
[0043] In step S110, configure the metric parameters, where the metric parameters include time window parameters and metric calculation rules;
[0044] Optionally, the operator can fill in or select the metric parameters as required in the metric configuration and definition module according to business needs to define specific calculation metrics. The operator needs to ensure that the filled parameters are accurate so that the system can automatically perform data collection, processing, and calculation based on these parameters, and finally generate the expected metric results. Among them, the metric parameters specifically include, but are not limited to, metric name, metric unique identifier, main attribute, subordinate attribute, calculation field, calculation method, time dimension, metric code, etc., providing clear metric calculation rules for subsequent data stream collection and time window calculation.
[0045] It should be noted that the metric configuration and definition module can also provide a metric adjustment interface, enabling the operator to easily perform parameter settings and modifications. Through such a design, the flexibility of metric configuration can be ensured to meet the business needs in different business scenarios.
[0046] It can be understood that the metric unique identifier can be used to accurately distinguish and manage each metric. For example, information such as business type, metric name, and time dimension is concatenated to obtain the unique identifier; or after concatenating the strings of this information, the hash value is calculated through a hash function and used as the unique identifier. The calculation field is the original data source for metric calculation. For example, in the case of an e-commerce business, when calculating the metric of product sales profit, fields such as sales amount, cost, and freight are the key data required for the calculation. The main attribute refers to the core object around which the metric revolves. For example, in social media data, the user ID can be the main attribute. The subordinate attribute is a supplement to the main attribute. For example, in social media data, the user's age, gender, region, etc. can be used as subordinate attributes. The metric name and metric code are used to uniquely identify each metric in the system, facilitating management and retrieval.
[0047] Optionally, the metric calculation rules can include any one or any combination of ordinary statistics, correlation calculation, historical correlation calculation, normal number, suspected number, risk number, normal rate, suspected rate, and risk rate.
[0048] Optionally, the time window parameter is a pre-defined time window, specifically a rolling window, which can include time dimensions such as hours, days, months, natural days, natural weeks, natural months, and natural years.
[0049] In step S120, collect the data stream and segment the data stream according to the time window parameter to obtain a number of sub-data streams;
[0050] Optionally, the data stream collection module is responsible for continuously monitoring and collecting data stream information and performing necessary data preprocessing, including operations such as data parsing, filtering, and conversion, to prepare the data for subsequent metric calculations. These data streams may include log information, sensor data, transaction records, etc. The module will segment the real-time data stream according to the time window parameters set in the metric configuration and definition module, such as a sliding window of 1 hour in size, to divide it into multiple sub-data streams. This ensures the timeliness and relevance of the data.
[0051] Among them, the data stream may include information such as user account, IP, device fingerprint, time, data unique identifier, risk level, risk type, etc.
[0052] In step S130, based on the metric calculation rules, calculate the corresponding sub-data streams within different time windows to obtain the metric intermediate values that meet the preset conditions under different time windows;
[0053] Optionally, the metric calculation rules may include any one or any combination of ordinary statistics, correlation calculation, historical correlation calculation, normal number, suspected number, risk number, normal rate, suspected rate, risk rate. The time window may include 1 minute, 1 hour, 1 day, etc.
[0054] Optionally, after obtaining the data stream, multiple identical data streams can be extracted. The number of data streams is the same as the number of time windows. The data streams can be segmented respectively through the corresponding time windows to obtain multiple sub-data streams. Then, based on the metric calculation rules set in the metric configuration and definition module, calculate the corresponding sub-data streams in each time window according to the corresponding metric calculation rules to obtain the metric intermediate values that meet the preset conditions under different time windows. Operations such as aggregation, deduplication, application of mathematical functions, and execution of statistical analysis can also be performed on the data set within the time window to obtain the real-time value of each metric in the current window. These real-time values are used as intermediate results and temporarily stored in the distributed storage database for further calculation to obtain the final metric value. In addition, this module also needs to ensure high throughput and low latency of data collection to meet the requirements of real-time calculation.
[0055] It should be noted that the time window needs to be highly flexible and scalable to adapt to different calculation requirements and changing business scenarios. In addition, it also needs to optimize the use of computing resources to achieve the optimization of computing efficiency and cost-effectiveness. The system also has a fault tolerance mechanism to handle situations of data loss or calculation anomalies to ensure the continuity and stability of the metric results.
[0056] In step S140, based on the metric calculation rules, calculate the metric intermediate values to obtain the metric result values.
[0057] Optionally, the metric result value may include a target unique value, which can be obtained by combining the metric unique identifier, user account, main attribute, and subordinate attribute. For example, the unique value can be obtained by concatenating the metric unique identifier, user account, main attribute, and subordinate attribute, or by calculating the hash value after concatenating the metric unique identifier, user account, main attribute, and subordinate attribute as a string, and using the calculated hash value as the target unique value.
[0058] Then, based on the unique value and the metric time dimension, the required metric intermediate value can be retrieved from the metric intermediate values calculated in all time windows corresponding to all metric time dimensions through the metric retrieval module. Then, data aggregation, statistical analysis, function calculation, etc. can be performed across multiple time windows to obtain the metric result value.
[0059] In an embodiment of the present application, the resource usage of the time window calculation module and the metric retrieval module, as well as whether the program application is running normally, can also be monitored and analyzed in real time. When it is detected that the resource usage exceeds the preset normal range or an abnormal pattern appears, the alarm mechanism is automatically triggered to notify the system administrator or relevant personnel for further inspection and handling. This reduces the workload of manual monitoring, improves the response speed to potential problems, and ensures the stability and reliability of the system.
[0060] In an embodiment of the present application, the resource allocation and calculation efficiency in the metric calculation process can also be optimized. Specifically, through means such as algorithm optimization, resource scheduling, and load balancing, the calculation performance can be improved and the calculation delay can be reduced. Especially when dealing with large-scale data streams and complex calculation tasks, it is ensured that the device can operate efficiently and stably. In addition, the calculation resources can be dynamically adjusted according to the actual calculation requirements and system load to achieve the optimization of cost-effectiveness.
[0061] In an embodiment of the present application, a data processing method is provided. The implementation of the method includes: configuring metric parameters, where the metric parameters include time window parameters and metric calculation rules; collecting a data stream and splitting the data stream according to the time window parameters to obtain a number of sub-data streams; calculating, based on the metric calculation rules, the sub-data streams corresponding to different time windows to obtain metric intermediate values that meet preset conditions for different time windows; and calculating, based on the metric calculation rules, the metric intermediate values to obtain a metric result value. In an embodiment of the present application, by dynamically splitting the data stream through configurable time window parameters, the calculation strategy can be adjusted according to the actual situation of the data stream, without relying on a large amount of cache to store data, avoiding the problems of calculation delay during peak data stream periods and resource waste during low periods, and meeting the flexible and variable business requirements. Moreover, the phased calculation method of first calculating the metric intermediate values and then calculating the metric result values based on the intermediate values reduces the pressure of processing a large amount of data at one time, reduces the calculation complexity and memory occupancy, and ensures the accuracy and traceability of the results.
[0062] In an embodiment of the present application, the splitting the data stream according to the time window parameters to obtain a number of sub-data streams includes:
[0063] extracting a number of identical data streams to be split from the data stream, where the number of data streams to be split is the same as the number of time windows;
[0064] splitting the data streams to be split according to the time window parameters to obtain a number of sub-data streams.
[0065] Optionally, the data stream can be collected in real time, and a number of identical data streams to be split are extracted from the data stream, where the number of data streams to be split is the same as the number of time windows. For example, if the time window parameters include 1 minute, 1 hour, and 1 day, then 3 identical data streams to be split can be extracted from the data stream. Then, each data stream to be split is split based on each time window parameter. Taking a 1-hour sliding window as an example, the overall time range of the data stream to be split can be determined first, and a starting time point is selected as the starting position of the sliding window. For example, if the time range of the data stream to be split is from 0:00 to 24:00 on a certain day, 0:00 can be selected as the starting time, and the first 1-hour sliding window starts from 0:00. The data is intercepted in sequence. That is, the first window intercepts the data from 0:00 to 1:00, the second window intercepts the data from 1:00 to 2:00, and so on, until the entire time range of the data stream to be split is covered. During the splitting process, it is carried out strictly in chronological order and with a fixed duration of 1 hour to ensure that each segmented data segment contains complete 1-hour data.
[0066] It can be understood that, in order to ensure the continuity of data and the accuracy of analysis, a certain overlap can be set between windows. For example, if a 10-minute overlap is set, then the second window is the data from 0:50 to 1:50, which can avoid information loss caused by window switching and more comprehensively reflect data changes.
[0067] In an embodiment of the present application, before calculating the to-be-segmented data stream according to the index calculation rule in different time windows to obtain the index intermediate values corresponding to different time windows, it includes:
[0068] In different time windows, calculate the to-be-segmented data stream according to the index calculation rule respectively to obtain the target data corresponding to different time windows;
[0069] In each time window, group the corresponding target data according to the preset fields respectively;
[0070] Based on business requirements, calculate and statistically analyze the grouped target data to obtain the real-time value of each index in the current time window as the index intermediate value.
[0071] Optionally, within each time window, calculations can be performed on the data for any one or any combination of general statistics, correlation calculation, historical correlation calculation, normal numbers, suspected numbers, risk numbers, normal rates, suspected rates, and risk rates to obtain intermediate results. Then, operations such as aggregation, deduplication, application of mathematical functions, and execution of statistical analysis can be performed on the intermediate results obtained within the time window to obtain the real-time value of each index in the current window, and these real-time values are used as intermediate results. For example, check whether there are duplicate records in the index intermediate values, and remove them if any. Then group the data within each time window according to specific fields, such as sales amount, profit, user accounts, etc., and summarize each group of data. Finally, various mathematical functions can be applied to the data, such as logarithmic functions, exponential functions, etc., for example, calculating the average consumption amount, the standard deviation of the consumption amount, etc., to meet different analysis needs. For example, perform a logarithmic transformation on the user consumption amount to better analyze the consumption distribution. More in-depth statistical analysis can also be performed, such as calculating the standard deviation, variance, etc., to understand the dispersion degree and distribution characteristics of the data. Thus, the real-time value of each index in the current window can be obtained, and this real-time value can be temporarily stored in the distributed storage database as an intermediate result. In order to quickly obtain these intermediate results when calculating the final index value in the subsequent calculation, the calculation efficiency can be improved, and at the same time, it is also convenient for data management and maintenance.
[0072] In an embodiment of the present application, before calculating the to-be-segmented data stream according to the index calculation rule in different time windows to obtain the index intermediate values corresponding to different time windows, it further includes:
[0073] Determine whether there are missing values or null values in the sub-data streams corresponding to different time windows respectively;
[0074] If so, supplement the missing values or null values.
[0075] Optionally, when calculating the intermediate value of an indicator for different time windows, each field of the sub-data stream corresponding to each time window can be traversed to determine whether there are null values or invalid values. If so, the missing values or null values can be supplemented based on the data type. For example, for numerical data, if the sales amount of a certain commodity in a certain time window is missing, it can be supplemented with the average sales amount of the commodity in other time windows. For special data, such as user behavior data, if the source channel of a certain user is missing, it can be marked as "unknown source" to ensure the integrity of the data and the feasibility of subsequent analysis.
[0076] In an embodiment of the present application, calculating the sub-data streams corresponding to different time windows based on the indicator calculation rule to obtain the intermediate value of the indicator that meets the preset conditions under different time windows includes:
[0077] Obtaining the intermediate value of the indicator through ordinary statistics, normal numbers, suspected numbers or risk number calculation methods within different time windows, where the intermediate value of the indicator includes the unique value of the combination of the indicator unique identifier and the user account, the total amount of characteristic data obtained by statistics, and the end time identifier of the time window; or
[0078] Obtaining the intermediate value of the indicator through association seeking or historical association seeking calculation methods within different time windows, where the intermediate value of the indicator includes the unique value of the combination of the indicator unique identifier, the user account, the main attribute and the subordinate attribute, the set of calculated field values, and the end time identifier of the time window; or
[0079] Obtaining the intermediate value of the indicator through normal rate, suspected rate or risk rate calculation methods within different time windows, where the intermediate value of the indicator includes the unique value composed of the indicator unique identifier and the user account, the total value and the amount of characteristic data, and the end time identifier of the time window.
[0080] It can be understood that the general statistics refer to the statistics of data by means of summation, averaging, counting, finding the maximum and minimum values, etc. For example, to calculate the average sales volume of product A within time t, the total sales volume of product A within time t is statistically calculated and then divided by time t to obtain the average sales volume. Seeking correlation is to analyze the correlation between two or more variables and judge whether there is a certain connection between them. For example, by analyzing the browsing history and purchase behavior of users, the products that are often browsed or purchased together are found. Historical seeking correlation refers to analyzing the relationship between variables by combining historical data. For example, obtaining the purchase records of users in the past 3 months to predict the product data related to the users. Normal numbers refer to the data that meet specific standard thresholds under normal business conditions. For example, the number of products that meet the production standards within a certain time range. Suspected numbers refer to the number of data that may have problems but have not been determined yet. For example, abnormal logins or multiple off-site logins, etc. Risk numbers refer to the data that clearly have risks. For example, the number of overdue loan payments. The normal rate is the ratio of normal numbers to the total data volume, which is used to measure the proportion of normal data in the overall data. The suspected rate is the ratio of suspected numbers to the total data volume, which reflects the proportion of data that may have problems in the overall data. The risk rate is the ratio of risk numbers to the total data volume, which is used to evaluate the risk level faced by the business.
[0081] Optionally, according to the pre-configured index calculation rules, the calculation method of this index calculation can be determined, and then the index calculation is performed through the corresponding index calculation method in different windows. For example, the index intermediate value can be obtained through general statistics, normal numbers, suspected numbers or risk number calculation methods in different time windows, or the index intermediate value obtained through seeking correlation or historical seeking correlation calculation methods, or the index intermediate value obtained through normal rate, suspected rate or risk rate calculation methods. The index intermediate value may include a unique value, a total value, a characteristic data volume, and an end time identifier of the time window. It can be understood that the total value refers to all the data statistically calculated in this time window, and the characteristic data volume refers to the data volume that meets specific requirements statistically calculated in this time window. The unique value can be calculated through a hash function, or can also be obtained by splicing each component.
[0082] In an embodiment of the present application, calculating the index result value based on the index calculation rules includes:
[0083] Construct a retrieval condition;
[0084] Based on the retrieval condition, retrieve the index intermediate value to obtain complete intermediate result information;
[0085] Statistically calculate the intermediate result information to obtain the index result value.
[0086] Optionally, retrieval conditions can be constructed based on the unique value calculated above and the metric time dimension. Through the metric retrieval module, data that conforms to the metric time dimension and the unique value can be retrieved from the stored intermediate metric values, including all intermediate metric values calculated for all windows of all metric time dimensions. Then, the retrieved intermediate metric values can be calculated again according to the preset metric calculation rules to obtain the metric result value.
[0087] It should be noted that the metric retrieval module is used to provide a high-performance and fast retrieval function for the calculated metrics. Users can retrieve the metric calculation results through this module. The metric retrieval module needs to support efficient data indexing and query optimization to ensure quick positioning and retrieval of metric data in massive data.
[0088] In the embodiments of the present application, by dynamically segmenting the data stream through configurable time window parameters, the calculation strategy can be adjusted according to the actual situation of the data stream, without relying on a large amount of cache to store data, avoiding the problems of calculation delay during data stream peaks and resource waste during valleys, and meeting the flexible and changeable business needs. Moreover, the phased calculation method of first calculating the intermediate metric value and then calculating the metric result value based on the intermediate value reduces the pressure of processing massive data at one time, reduces the calculation complexity and memory occupancy, and ensures the accuracy and traceability of the results.
[0089] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0090] In one embodiment, a data processing device is provided, which corresponds one-to-one with the data processing method in the above embodiment. As Figure 2 shown, the data processing device includes a metric parameter configuration unit 10, a data stream segmentation unit 20, an intermediate metric value determination unit 30, and a metric result value determination unit 40. The detailed description of each functional module is as follows:
[0091] The metric parameter configuration unit 10 is used to configure metric parameters, and the metric parameters include time window parameters and metric calculation rules;
[0092] The data stream segmentation unit 20 is used to collect the data stream and segment the data stream according to the time window parameters to obtain a number of sub-data streams;
[0093] The intermediate metric value determination unit 30 is used to calculate the corresponding sub-data streams within different time windows based on the metric calculation rules to obtain intermediate metric values that meet the preset conditions under different time windows;
[0094] The index result value determination unit 40 is configured to calculate the index result value based on the index calculation rule for the index intermediate value.
[0095] In an embodiment of the present application, the data stream splitting unit 20 is further configured to:
[0096] Extract a plurality of identical data streams to be split from the data stream, where the number of the data streams to be split is the same as the number of the time windows;
[0097] Split the data streams to be split according to the time window parameters to obtain a plurality of sub-data streams.
[0098] In an embodiment of the present application, the data stream splitting unit 20 is further configured to:
[0099] In each time window, group the corresponding sub-data streams respectively according to a preset field;
[0100] Aggregate each group of sub-data streams;
[0101] Statistically analyze the aggregated data according to the index calculation rule to obtain the real-time value of each index in the current time window as the index intermediate value.
[0102] In an embodiment of the present application, the index intermediate value determination unit 30 is further configured to:
[0103] Determine respectively whether there are missing values or null values in the sub-data streams corresponding to different time windows;
[0104] If so, supplement the missing values or null values.
[0105] In an embodiment of the present application, the index intermediate value determination unit 30 is further configured to:
[0106] Obtain the index intermediate value through ordinary statistics, normal number, suspected number or risk number calculation methods within different time windows, where the index intermediate value includes the unique value of the combination of the index unique identifier and the user account, the total amount of feature data statistically obtained, and the end time identifier of the time window; or
[0107] Obtain the index intermediate value through association calculation or historical association calculation methods within different time windows, where the index intermediate value includes the unique value of the combination of the index unique identifier, the user account, the main attribute and the subordinate attribute, the set of calculated field values, and the end time identifier of the time window; or
[0108] Obtain the index intermediate value through normal rate, suspected rate or risk rate calculation methods within different time windows, where the index intermediate value includes the unique value composed of the index unique identifier and the user account, the total amount value and the feature data amount, and the end time identifier of the time window.
[0109] In an embodiment of the present application, the index result value determination unit 40 is further configured to:
[0110] Construct a retrieval condition;
[0111] Based on the retrieval condition, retrieve the index intermediate value to obtain complete intermediate result information;
[0112] Perform statistics and calculations on the intermediate result information to obtain the index result value.
[0113] In an embodiment of the present application, the index calculation rule includes any one or any combination of ordinary statistics, correlation calculation, historical correlation calculation, normal constant, suspected number, risk number, normal rate, suspected rate, and risk rate.
[0114] In the embodiments of the present application, by dynamically segmenting the data stream through configurable time window parameters, the calculation strategy can be adjusted according to the actual situation of the data stream, without relying on a large amount of cache to store data, avoiding the problems of calculation delay during the data stream peak and resource waste during the trough, and meeting the flexible business needs. Moreover, the phased calculation method of first calculating the index intermediate value and then calculating the index result value based on the intermediate value reduces the pressure of processing a large amount of data at one time, reduces the calculation complexity and memory occupancy, and ensures the accuracy and traceability of the results.
[0115] For the specific limitations on the data processing device, reference can be made to the limitations on the data processing method in the above text, which will not be elaborated here. Each module in the above data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0116] In one embodiment, a computer device is provided. The computer device can be a terminal device, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a data processing method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0117] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the data processing method as described above are implemented.
[0118] In an embodiment of the present application, a readable storage medium is provided. The readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the data processing method as described above are implemented.
[0119] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0120] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0121] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Configure indicator parameters, which include time window parameters and indicator calculation rules; Collecting a data stream, and dividing the data stream according to the time window parameter to obtain a plurality of sub-data streams; Based on the indicator calculation rules, the corresponding sub-data streams in different time windows are calculated to obtain the intermediate values of the indicators that meet the preset conditions in different time windows; Based on the indicator calculation rule, the indicator intermediate value is calculated to obtain the indicator result value.
2. The data processing method according to claim 1, characterized in that: The data stream is divided according to the time window parameter to obtain a plurality of sub-data streams, including: Extracting a plurality of identical data streams to be segmented from the data stream, the number of the data streams to be segmented being the same as the number of the time windows; The data stream to be segmented is segmented according to the time window parameter to obtain a plurality of sub-data streams.
3. The data processing method according to claim 2, characterized in that: The step of calculating the data streams to be segmented in different time windows according to the indicator calculation rule to obtain intermediate values of indicators corresponding to different time windows includes: In each time window, the corresponding sub-data streams are grouped according to preset fields; Aggregate each group of sub-data streams; The aggregated data is statistically analyzed according to the indicator calculation rules to obtain the real-time value of each indicator in the current time window as the intermediate value of the indicator.
4. The data processing method according to claim 2, characterized in that: Before respectively calculating the data streams to be segmented in different time windows according to the indicator calculation rule to obtain the intermediate values of the indicators corresponding to the different time windows, the method further includes: Determine whether there are missing values or null values in the sub-data streams corresponding to different time windows; If so, fill in the missing value or empty value.
5. The data processing method according to claim 1, characterized in that: The step of calculating the corresponding sub-data streams in different time windows based on the indicator calculation rule to obtain the intermediate values of the indicators that meet the preset conditions in different time windows includes: The intermediate value of the indicator is obtained by ordinary statistics, normal numbers, suspected numbers or risk number calculation methods in different time windows, and the intermediate value of the indicator includes the unique value of the combination of the indicator unique identifier and the user account, the total amount of characteristic data obtained by statistics, and the end time identifier of the time window; or The intermediate value of the indicator obtained by association or historical association calculation in different time windows, the intermediate value of the indicator includes the indicator unique identifier and the unique value of the combination of the user account and the primary attribute and the secondary attribute, the set of calculated field values, and the end time identifier of the time window; or The intermediate value of the indicator is obtained by calculating the normal rate, suspected rate or risk rate in different time windows. The intermediate value of the indicator includes the unique value of the indicator unique identifier and the user account, the total value and the characteristic data volume, and the end time identifier of the time window.
6. The data processing method according to claim 1, characterized in that: The step of calculating the intermediate value of the indicator based on the indicator calculation rule to obtain the indicator result value includes: Construct search criteria; Based on the search condition, the intermediate value of the indicator is searched to obtain complete intermediate result information; The intermediate result information is counted and calculated to obtain the indicator result value.
7. The data processing method according to any one of claims 1 to 6, characterized in that: The indicator calculation rules include any one of common statistics, correlation seeking, historical correlation seeking, normal number, suspected number, risk number, normal rate, suspected rate, risk rate or any combination thereof.
8. A data processing device, characterized in that: The device comprises: An indicator parameter configuration unit, used to configure indicator parameters, wherein the indicator parameters include time window parameters and indicator calculation rules; A data stream segmentation unit is used to collect a data stream and segment the data stream according to the time window parameter to obtain a plurality of sub-data streams. An indicator intermediate value determination unit, used to calculate the corresponding sub-data streams in different time windows based on the indicator calculation rule to obtain the indicator intermediate values that meet the preset conditions in different time windows; The indicator result value determination unit is used to calculate the indicator intermediate value based on the indicator calculation rule to obtain the indicator result value.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that: When the processor executes the computer-readable instructions, the steps of the data processing method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.