Data processing method and device

By combining Zset and HLL data structures and sliding window solutions, a real-time deduplication accumulation method of automatic switching is realized, which solves the problems of high storage costs and node storage skew in the prior art, and improves the efficiency and stability of the real-time system.

CN120066414APending Publication Date: 2025-05-30SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510159827.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing real-time deduplication accumulation method has high storage costs, resulting in skewed node storage, affecting the efficiency and stability of the real-time system, and cannot support large-scale real-time user behavior analysis.

Method used

The two data structures of Zset and HLL are combined with the sliding window scheme, and the data structure is automatically switched according to the scale of event characteristics to ensure that the data storage occupies at any scale.

Benefits of technology

Real-time deduplication accumulation at different scales is realized, storage costs are reduced, the overall efficiency and stability of the real-time system are improved, and large-scale real-time user behavior analysis is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066414A_ABST
    Figure CN120066414A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, and the method comprises the steps: obtaining target event data, and the target event data comprises a target event feature; determining a storage type, wherein the storage type comprises a first storage type or a second storage type; writing the target event data into a first data structure or a second data structure according to the storage type; the first data structure is used for accumulating deduplication values of event features of a first scale, the second data structure is used for accumulating deduplication values of event features of a second scale, the first scale is smaller than the second scale, and the first data structure is converted into the second data structure when the accumulated deduplication values exceed a preset conversion threshold value. According to the technical scheme, by combining the first data structure and the second data structure, real-time deduplication accumulation under different scales can be achieved, it is ensured that data storage occupation under any scale is minimum, and the overall efficiency and stability of a real-time system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a data processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] An event is a specific action or activity executed in a platform, application, or system. Through cumulative analysis of event data, in-depth understanding of system performance, stability, resource consumption, etc. can be achieved. Different cumulative methods have differences in aspects such as cumulative timeliness, data scale, and counting accuracy. A real-time system needs to perform feature judgment and deduplication and accumulation on each piece of event data, requiring high timeliness and high concurrency capabilities. However, existing real-time deduplication and accumulation methods currently have a high storage cost, are prone to causing node storage skew, affecting the overall efficiency and stability of the real-time system, and are unable to support large-scale real-time user behavior analysis.

[0003] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0004] Embodiments of the present application provide a data processing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.

[0005] One aspect of embodiments of the present application provides a data processing method, the method including: Obtaining target event data, where the target event data includes target event features; Determining a storage type, where the storage type includes a first storage type or a second storage type; Writing the target event data into a first data structure or a second data structure according to the storage type; wherein, the first data structure is used to accumulate the deduplicated values of event features of a first scale, the second data structure is used to accumulate the deduplicated values of event features of a second scale, the first scale is smaller than the second scale, and the first data structure is converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold.

[0006] Optionally, the first storage type corresponds to the first data structure, and the first data structure includes an ordered set; the second storage type corresponds to the second data structure, and the second data structure includes a cardinality estimation data structure.

[0007] Optionally, event features include grouping features and cumulative features; the target event data further includes a target event time; the first data structure includes a plurality of elements, and each element has a corresponding score; Correspondingly, writing the target event data into the first data structure according to the storage type includes: Determining a grouping feature, an accumulation feature, and a target event time according to the target event data; Performing a hash operation on the grouping feature to obtain a hash tag; Determining a corresponding first data structure according to the hash tag; Writing the accumulation feature and the target event time as new elements and their corresponding scores into the corresponding first data structure.

[0008] Optionally, the event feature includes a grouping feature and an accumulation feature; the target event data further includes a target event time; the second data structure includes a plurality of buckets, and each bucket has a corresponding start time; Correspondingly, writing the target event data into the second data structure according to the storage type includes: Determining a grouping feature, an accumulation feature, and a target event time according to the target event data; Performing a hash operation on the grouping feature to obtain a hash tag; Determining a corresponding second data structure according to the hash tag; Determining the start time of the corresponding bucket according to the target event time to determine the current bucket from the plurality of buckets; Writing the accumulation feature into the current bucket.

[0009] Optionally, determining the start time of the corresponding bucket according to the target event time includes: Performing an integer division operation on the target event time and the bucket period to obtain a bucket index number; wherein, the bucket period is the difference between the start times of two adjacent buckets; Performing a multiplication operation on the bucket index number and the bucket period to obtain the start time of the current bucket.

[0010] Optionally, the data processing method further includes: Determining a current window according to the current bucket, where the current window includes a preset number of buckets whose start times are earlier than the current bucket; Writing the accumulation feature into the current window.

[0011] Optionally, the data processing method further includes: Performing a modulo operation on the hash tag and the bucket period to obtain an offset value; wherein, the bucket period is the difference between the start times of two adjacent buckets; Determine a preloading time interval according to the start time of the current bin, the offset value, and the start time of the next bin of the current bin; Determine whether the target event time is within the preloading time interval; When the target event time is within the preloading time interval, create the next window of the target window; Write the cumulative feature into the next window of the target window.

[0012] Optionally, the target event data further includes a target event time; the second data structure includes a plurality of bins, each bin having a corresponding start time; the first data structure includes a plurality of elements, each element having a corresponding score; wherein, the element is event data and the score is an event time; Correspondingly, the first data structure is converted into the second data structure through the following operations: Determine the start time of the current bin according to the target event time; Determine a plurality of target conversion bins according to the start time of the current bin, a preset number of bins, and an accumulation period, each target conversion bin having a corresponding start time; Obtain corresponding event data from the first data structure according to the start times of the plurality of target conversion bins and write it into the plurality of target conversion bins to obtain the second data structure.

[0013] Optionally, the data processing method further includes: When converting the first data structure into the second data structure, determine a conversion flag; Correspondingly, determine the storage type, including: Obtain the conversion flag; When the conversion flag is not obtained, the storage type is the first storage type; When the conversion flag is obtained, the storage type is the second storage type.

[0014] Optionally, the data processing method further includes: When the conversion flag is obtained, reset the expiration time of the conversion flag to the sum of the target event time and the accumulation period; and / or After writing the target event data into the first data structure or the second data structure, write the corresponding storage type into the local cache.

[0015] Optionally, the conversion flag includes the completion time of the most recent conversion; Correspondingly, according to the start time of the current bin, the preset number of bins, and the accumulation period, a plurality of target conversion bins are determined, and each target conversion bin has a corresponding start time, including: Obtain a conversion flag; In the case where the conversion flag is not obtained: According to the start time of the current bin, the preset number of bins, and the accumulation period, determine a plurality of bins within the accumulation period; Determine the plurality of bins within the accumulation period as the plurality of target conversion bins; In the case where the conversion flag is obtained, according to the start time of the current bin, the preset number of bins, and the accumulation period, determine a plurality of bins within the accumulation period whose start time is later than the completion time of the most recent conversion; Determine the plurality of bins within the accumulation period whose start time is later than the completion time of the most recent conversion as the plurality of target conversion bins.

[0016] Optionally, the data processing method further includes: Determine an atomic command according to a preset accumulation period, where the atomic command is used to obtain the deduplicated value accumulated by the first data structure; Based on the atomic command, obtain a target deduplicated value from the first data structure; or Determine the current window in the second data structure based on the target event time; Based on the current window, obtain the target deduplicated value.

[0017] Another aspect of the embodiments of the present application provides a data processing device, where the device includes: An acquisition module, configured to acquire target event data, where the target event data includes target event characteristics; A determination module, configured to determine a storage type, where the storage type includes a first storage type or a second storage type; A writing module, configured to write the target event data into a first data structure or a second data structure according to the storage type; Wherein, the first data structure is used to accumulate the deduplicated value of event characteristics of a first scale, the second data structure is used to accumulate the deduplicated value of event characteristics of a second scale, the first scale is smaller than the second scale, and the first data structure is converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold.

[0018] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0019] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0020] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0021] The embodiments of the present application adopting the above technical solutions may include the following advantages: Obtain target event data including target event features, and determine that its storage type is the first storage type or the second storage type. According to the storage type, write the target event data into the first data structure or the second data structure. Among them, the first data structure is used to accumulate the deduplicated values of low-cardinality event features, the second data structure is used to accumulate the deduplicated values of high-cardinality event features, and the first data structure will be converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold. It can be seen that the embodiments of the present application realize real-time deduplication and accumulation under different scales by combining the first data structure and the second data structure, ensure the minimum data storage occupancy under any scale, improve the overall efficiency and stability of the real-time system, and support large-scale real-time user behavior analysis. Description of the Drawings

[0022] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0023] Figure 1 Schematically shows a flowchart of the data processing method according to Embodiment 1 of the present application; Figure 2 Schematically shows an overall flowchart of data writing according to Embodiment 1 of the present application; Figure 3 Schematically shows the data reading and writing process of HLL according to Embodiment 1 of the present application; Figure 4 Schematically shows the window Key writing process according to Embodiment 1 of the present application; Figure 5 Schematically shows an overall flowchart of data reading according to Embodiment 1 of the present application; Figure 6 Schematically shows a flowchart of converting Zset to HLL according to Embodiment 1 of the present application; Figure 7 Schematically shows a flowchart of data writing according to Embodiment 1 of the present application; Figure 8 Schematically shows another flowchart of converting Zset to HLL according to Embodiment 1 of the present application; Figure 9 Schematically shows a block diagram of a data processing device according to Embodiment 2 of the present application; and Figure 10 Schematically shows a schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0024] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0025] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0026] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0027] First, provide the term explanations involved in the present application: Sliding window: A technology for processing data streams through a window of a fixed size.

[0028] Bucketing: A data processing method that divides data into multiple buckets based on a specific rule, and the data within each bucket has similar attributes.

[0029] Redis: An open-source, ANSI C language-written, network-supported, log-type, Key-Value database that can be memory-based or persistent, and is widely used in scenarios such as caching, counters, and real-time analysis. Redis can be deployed in a cluster mode for horizontal capacity expansion.

[0030] HashTag: A naming pattern provided by Redis that takes effect in the Redis cluster mode. Keys with the same HashTag are assigned to the same slot. The specific format can be "{HashTag}".

[0031] Slot: A mechanism used for sharding in the Redis cluster mode. Data in the same Slot is stored in the same node. Many Redis commands require the keys to be operated on to be in the same Slot, such as rename, etc.

[0032] Floor function: A mathematical operation used to discard the fractional part of a real number and take the largest integer not greater than the real number. Specifically, for any real number x, the result of its floor function is denoted as ⌊x⌋, representing the largest integer less than or equal to x.

[0033] Hash: A cryptographic technique used to convert a message of any length into a fixed-length digest information and ensure that any modification to the original message will produce a different digest information. Its calculation result is a smaller integer or string, which can be called a "hash value" or "hash result".

[0034] IO operation: Input / Output operation.

[0035] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: An event is a specific action or activity executed in a platform, application, or system. By cumulatively analyzing event data, the performance, stability, and resource consumption of the system can be deeply understood. Different cumulative methods have differences in cumulative timeliness, data scale, and counting accuracy. A real-time system needs to perform feature judgment and duplicate removal and accumulation on each piece of event data, requiring high timeliness and high concurrency capabilities.

[0036] However, the applicant has learned that the related real-time duplicate removal and accumulation methods have a high storage cost, are prone to causing node storage skew, affecting the overall efficiency and stability of the real-time system, and are unable to support large-scale real-time user behavior analysis.

[0037] To this end, the embodiments of the present application provide a technical solution for data processing. In this technical solution: (1) Zset (ordered set) and HyperLogLog (cardinality estimation data structure, HLL) are used to implement real-time deduplication and accumulation of different scales respectively. A preset automatic conversion threshold L0 is set. When the deduplication value is less than L0, Zset is used to store the count. When the deduplication value is greater than L0, HLL is used to store the count. When the deduplication value first reaches L0, HLL is automatically generated based on the data in Zset, and Zset is safely deleted to ensure the lowest data storage occupancy at any scale. The Lua script is used to implement the data conversion process to ensure the atomicity of the conversion process. (2) By increasing the local cache, the IO operations for the real-time system to judge the storage type (Zset or HLL) are reduced, and the CPU resource occupancy rate is lowered. (3) By centralizing the storage conversion flag (Store Flag, SF) and the conversion completion time (Data Conversation Timestamp, DCT), the storage type can be quickly judged, and the extra IO generated when judging the storage type during data reading and writing is reduced. That is, when SF does not exist, the data has not been converted, and Zset can be used for storage. Otherwise, HLL is used for storage. At the same time, when the data conversion is triggered, the start time of the data conversion is initialized based on the existing DCT, so as to solve the problem of resource waste caused by multiple triggers of data conversion under distributed high-concurrency conditions. (4) HLL is implemented using a sliding window scheme, and the number of buckets is denoted as N. Each time the window rolls, a window for the current accumulation period is generated. Since the current window is a new window, the window HLL can be generated only by merging the first N - 1 buckets before the current time. During the data writing process, by writing data to the buckets and the window in a double-write manner, both the need for window generation during sliding window switching can be met, and the window Key can be directly read to achieve efficient query under high-concurrency conditions. (5) A global sliding window is used, so that the start and end times of each bucket corresponding to different event characteristics are the same, and the start and end times of the deduplication value within a fixed period are the same at the same moment, ensuring the accuracy of the calculation between features. (6) Through the window preloading mechanism, the process of creating the window is scattered within a bucket time, avoiding the problem that all sliding windows switch simultaneously when using the global sliding window, resulting in a concentrated occurrence of the window creation behavior, forming an instantaneous high CPU occupancy and reducing the stability of the Redis cluster. See the following for details.

[0038] The technical solution of the present application will be introduced below through multiple embodiments. It should be noted that these embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments described herein.

[0039] Embodiment 1 Figure 1 The flowchart of the data processing method according to Embodiment 1 of the present application is schematically shown.

[0040] As Figure 1 shown, the data processing method may include steps S100 to S104, where: Step S100, obtaining target event data, where the target event data includes target event features.

[0041] Step S102, determining a storage type, where the storage type includes a first storage type or a second storage type.

[0042] Step S104, writing the target event data into a first data structure or a second data structure according to the storage type; where the first data structure is used to accumulate the deduplicated values of event features of a first scale, the second data structure is used to accumulate the deduplicated values of event features of a second scale, the first scale is smaller than the second scale, and the first data structure is converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold.

[0043] The data processing method provided in this embodiment obtains target event data including target event features, and determines its storage type as a first storage type or a second storage type. According to the storage type, the target event data is written into a first data structure or a second data structure. Among them, the first data structure is used to accumulate the deduplicated values of event features with a low cardinality, the second data structure is used to accumulate the deduplicated values of event features with a high cardinality, and the first data structure will be converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold. It can be seen that the embodiment of the present application realizes real-time deduplication and accumulation under different scales by combining the first data structure and the second data structure, ensures the minimum data storage occupancy under any scale, improves the overall efficiency and stability of the real-time system, and can support large-scale real-time user behavior analysis.

[0044] The following will combine Figure 1 , and elaborate on each step in steps S100 to S104 and optional other steps in detail.

[0045] Step S100, obtaining target event data, where the target event data includes target event features.

[0046] The target event can be triggered by user behavior. For example, the user accesses the service platform to watch videos, purchase goods, etc. The target event data can include target event characteristics and target event time (i.e., the occurrence time of the target event, which can be regarded as the current time in a real-time system). The event characteristics can be the text information carried in the request sent when the user accesses the platform. Exemplarily, the event characteristics can include grouping characteristics and cumulative characteristics. Among them, the grouping characteristics can be composed of one or more user behavior characteristics, which are used to distinguish the subject of the event characteristics, and each group of subjects corresponds to a deduplication value. The cumulative characteristic can be composed of one user behavior characteristic, which is used to describe the characteristic that specifically needs to be deduplicated and accumulated in real time. For example, event A can be "user kk watched video a", at this time the grouping characteristic is "user kk", and the cumulative characteristic is "video a". Event B can be "user kk watched video b", at this time the grouping characteristic is "user kk", and the cumulative characteristic is "video b". To achieve effective user behavior analysis, the user behavior characteristics can be deduplicated and accumulated in real time. The real-time deduplication and accumulation can be the deduplication statistics of the occurrence frequency accumulated from the user's previous same behavior when the user's real-time behavior occurs. Among them, the deduplication statistics can be that when the same cumulative characteristic (such as watching video b) appears multiple times, the frequency statistic value is counted as 1. For example: "accumulate the number (deduplication value) of videos (cumulative characteristic) watched by user kk (grouping characteristic) within one day (cumulative period T)", "accumulate the number (deduplication value) of users (cumulative characteristic) logged in with the same IP (grouping characteristic) within one week (cumulative period T). It should be noted that the embodiments of the present application obtain the target event data on the premise of ensuring compliance and obtaining consent, and perform desensitization processing on the obtained target event data to achieve privacy protection.

[0047] Step S102, determine the storage type, where the storage type includes a first storage type or a second storage type.

[0048] Exemplarily, the storage type can be determined in multiple ways. For example: calculate the hash tag according to the grouping characteristic in the target event characteristic, search for the corresponding storage node in the Redis cluster based on the hash tag, and then determine the storage type. The storage type can also be quickly determined by reading the local cache to reduce access to IO. The storage type can be divided into a first storage type and a second storage type.

[0049] Step S104, according to the storage type, write the target event data into a first data structure or a second data structure; where the first data structure is used to accumulate the deduplication value of the event characteristics of the first scale, the second data structure is used to accumulate the deduplication value of the event characteristics of the second scale, and the first scale is smaller than the second scale. The first data structure is converted into the second data structure when the accumulated deduplication value exceeds the preset conversion threshold.

[0050] In an alternative embodiment, the first storage type may correspond to a first data structure. The second storage type may correspond to a second data structure. The first data structure can be used to accumulate the deduplicated values of event features of the first scale (low cardinality), such as a user may watch very few videos in a day (e.g., less than 128). In contrast, the second data structure can be used to accumulate the deduplicated values of event features of the second scale (high cardinality), such as the number of users logging in with the same IP in a day (tens of thousands). Exemplarily, the first data structure can be Zset, Set, Bitmap, etc., and the second data structure can be HLL, T-Digest, etc.

[0051] Taking the first data structure Zset and the second data structure HLL as examples, the data processing method provided by the embodiments of the present application will be exemplarily introduced below.

[0052] Exemplarily, Zset is a complex data type provided by Redis. Through the score (also known as the weight parameter), the elements in the set (Zset) can be sorted in an orderly manner according to the score, and the elements can be deduplicated. The elements in Zset are unique and do not allow duplicates. When a newly added element is the same as the element in the set, only the score is updated, and the same element is not added repeatedly. Zset can save data in memory completely in chronological order, so it has advantages such as second-level timeliness, high query performance, and high write performance in real-time systems. And HLL is a probabilistic high-cardinality approximate calculation algorithm for data structures used in cardinality estimation, supporting key features such as data merging and automatic expansion. HLL only describes the cardinality information of a single data set, does not contain the data itself, has no time characteristics, can be applied to offline deduplication and accumulation of large-scale data sets, and has a constant-level space complexity and O(1) time complexity.

[0053] In some embodiments, the first data structure Zset or the second data structure HLL can also be used independently for real-time deduplication and accumulation. In practical applications, limited by the application scale, when the accumulation period of user characteristics is too long or the frequency within the period is too high, the complete storage of user behavior characteristics by Zset will occupy a large amount of storage space, resulting in high storage costs for the deduplication and accumulation process. At the same time, due to the excessive size of a single feature set, it will cause storage skew in Redis nodes, affect the stability of the Redis cluster, and reduce the read and write efficiency of the accumulation calculation, thus affecting the overall efficiency of the real-time system. The underlying implementation of Zset is related to the amount of data in the set. When the number of elements is less than a preset conversion threshold (such as 128), a compressed list is used for storage, which has better storage efficiency than HLL. However, HLL is only based on the cardinality of the complete data set and has no time characteristics. It can only be applied to the offline deduplication and accumulation of large-scale data sets and does not meet the requirements of high concurrency and high timeliness of real-time systems in terms of design. In some embodiments, a sliding window scheme can be further adopted to enable HLL to achieve real-time deduplication statistics. Exemplarily, the complete accumulation period can be evenly divided into N parts according to a preset rule using the bucketing method, and each part represents the deduplication value within its respective start and end times and is stored using a single HLL Key, called the bucketing Key. When reading, the continuous N bucketing data can be merged to form a complete HLL and stored using a single independent Key, and its result value represents the deduplication value within an accumulation period, called the window Key (Win Key). When writing target event data, it can be written into the bucket including this event time according to the time of event occurrence (i.e., the current time), and the corresponding bucket can be called the current bucket. When the window rolls when the current time is later than the cut-off time of the current bucket, a new current bucket can be generated and the earliest bucket can be eliminated. However, HLL using the sliding window scheme still has the following disadvantages: (1) Decrease in data accuracy: Since the HLL result is an approximate cumulative estimate value and does not store specific data, there is a certain error. On this basis, since expired data is only eliminated when the window rolls, it results in a higher latency compared to Zset. The maximum latency is the start and end duration of the bucket (bucket period), which further reduces the data accuracy. To obtain sufficient data accuracy, it is necessary to increase the number of buckets as much as possible. (2) Reduction in overall storage efficiency: Using buckets to store data for each time period requires a storage overhead for each bucket. The higher the accuracy of the accumulation, the more buckets are required, and the higher the total storage occupied by the buckets. In the case of low cardinality, the overall storage efficiency is worse than that of Zset. (3) The problem that the relationship between features cannot be directly calculated. To ensure the accuracy of the result, multiple event features participating in the calculation need to have the same time characteristics, including: cumulative duration, start and end times, event occurrence time, etc. If the sliding windows used for each event feature are independent of each other, they cannot be directly calculated with each other to determine the relationship between features.(4) Increased overall computing resource occupancy: To ensure the long-term stable operation of HLL in a real-time system, the additional computing resource consumption is much higher than that of using Zset alone. To a certain extent, the storage resource bottleneck is converted into a computing resource bottleneck, including: additional I / O generated by the bucketing scheme, high computing occupancy caused by HLL data merging, etc.

[0054] As can be seen from the above-mentioned multiple embodiments, there are limitations in using Zset or HLL alone for real-time deduplication and accumulation. Therefore, the embodiments of the present application creatively integrate Zset and HLL, and combine a sliding window to achieve real-time deduplication and accumulation. At the cost of adding a small amount of computing resources and reducing a certain data accuracy, it realizes the approximate calculation of real-time deduplication and accumulation values under distributed and high-concurrency conditions with extremely low storage costs and extremely high read and write efficiencies, and can solve problems such as excessive real-time cumulative storage costs and reduced system stability under high-cardinality and long-cycle conditions for user characteristics, and the data accuracy can support large-scale real-time user behavior analysis.

[0055] Exemplarily, as Figure 2 shown, when the storage type is the first storage type, the target event data can be written into the corresponding first data structure Zset to perform real-time deduplication and accumulation through Zset. When the storage type is the second storage type, the target event data can be written into the corresponding second data structure HLL to perform real-time deduplication and accumulation through HLL. Among them, when the deduplication value accumulated by the first data structure Zset exceeds the preset conversion threshold (such as L0 = 128), it will be automatically converted into the second data structure HLL to ensure the lowest data storage occupancy at any scale. Multiple exemplary solutions are provided below.

[0056] In an alternative embodiment, step S104 may include: Step S200, determining the grouping feature, the accumulation feature, and the target event time according to the target event data.

[0057] Step S202, performing a hash operation on the grouping feature to obtain a hash tag.

[0058] Step S204, determining the corresponding first data structure according to the hash tag.

[0059] Step S206, writing the accumulation feature and the target event time as a new element and its corresponding score into the corresponding first data structure.

[0060] Exemplarily, by parsing the target event data, grouping features, cumulative features, and the target event time can be determined. Performing a hash operation on the grouping features can obtain a hash tag HashTag. The hash tag can be used as the key prefix of the Zset to distinguish different Zsets. Data with the same HashTag is in the same Slot. If there is only one grouping feature, it can be represented as {HashTag}. If there are multiple grouping features, a hash operation can be performed on each grouping feature, and the obtained hash tags can be concatenated with ":" in lexicographical order. When storing using the first data structure Zset, the data is not bucketed, and the corresponding first data structure can be found according to the hash tag to store the cumulative features. Of course, a Zset with a key prefix can also be created to store the cumulative features. As Figure 2 shown, the cumulative feature can be written as a new element (member) into the Zset, and the target event time (such as the second-level timestamp of the current time, which can be denoted as CT) can be written as the score corresponding to the new element into the Zset.

[0061] In this embodiment, by writing the target event data into the ordered set Zset, the data can be completely stored in memory in chronological order, with high query performance and good write performance, and has second-level timeliness.

[0062] In an alternative embodiment, step S104 may include: Step S300, determining grouping features, cumulative features, and the target event time according to the target event data.

[0063] Step S302, performing a hash operation on the grouping features to obtain a hash tag.

[0064] Step S304, determining the corresponding second data structure according to the hash tag.

[0065] Step S306, determining the start time of the corresponding bucket according to the target event time to determine the current bucket from the multiple buckets.

[0066] Step S308, writing the cumulative feature into the current bucket.

[0067] Exemplarily, by parsing the target event data, grouping features, cumulative features, and the target event time can be determined. Performing a hash operation on the grouping features can obtain a hash tag. According to the hash tag, the corresponding second data structure HLL can be found or created. As Figure 3As shown, when the bucket or window does not exist, that is, the sliding window has scrolled, corresponding HLL Keys (bucket Key, window Key) can be generated respectively. For the bucket Key, Redis will automatically create a new HLL when writing data, and the expiration time can be set to an accumulation period. Among them, the bucket Key of the HLL is generated in the format of "Key prefix: start time". The window Key is generated in the format of "Key prefix: start time: end time". When a user behavior occurs, the start time of the corresponding bucket can be calculated using the target event time. From this, the current bucket Key and the current window Key can be determined from multiple buckets of the HLL according to the start time of the corresponding bucket, and the cumulative features are written into the current bucket.

[0068] In this embodiment, writing the cumulative features into the current bucket of the HLL can ensure that the data is written immediately, which can be used for subsequent calculations or queries without waiting until the data is fully accumulated before writing, and can meet the real-time deduplication and accumulation requirements.

[0069] In an alternative embodiment, step S306 may include: Step S400, performing an integer division operation based on the target event time and the bucket period to obtain a bucket index number; wherein, the bucket period is the difference between the start times of two adjacent buckets.

[0070] Step S402, performing a multiplication operation based on the bucket index number and the bucket period to obtain the start time of the current bucket.

[0071] Exemplarily, the start time of the bucket can be obtained in the following manner: BT = BI × BC, BI = ⌊ET / BC⌋. Where ET (Event Timestamp) can be the event time (such as a second-level timestamp). BC can be the bucket period (the unit can be seconds). BI is the result of rounding down the division of ET by BC, which can be called the bucket index number. Multiplying the bucket index number by the bucket period can obtain the start time of the current bucket. For example: the event time is 1730369506, the bucket period is 300 seconds, BI is 5767898, and the start time of the corresponding bucket is 1730369400.

[0072] In this embodiment, the start time of the corresponding bucket can be quickly and accurately determined by combining integer division and multiplication operations, so as to efficiently locate the current bucket from multiple buckets of the HLL.

[0073] In an alternative embodiment, step S104 may further include: Step S500, determining the current window according to the current bucket, where the current window includes a preset number of buckets whose start times are earlier than the current bucket.

[0074] Step S502, write the cumulative feature into the current window.

[0075] Exemplarily, the current window can be determined based on the current bin. The cut-off time of the bin is the start time of the next bin (excluding). For the window Key, as Figure 3 shown, the first N - 1 bins of the current bin can be merged into a new HLL, and the expiration time can be set to a bin period, which will expire naturally after the next window is created, further reducing memory occupancy. The start time WT of the current window is the start time of the first bin that constitutes the current window, and the cut-off time of the current window is the start time of the next window (excluding). As Figure 3 shown, the cumulative feature can be double-written to the bin and the window.

[0076] In this embodiment, double-writing the cumulative feature to the bin and the window can not only meet the need for window generation when the sliding window switches, but also directly read the window Key to achieve efficient query under high concurrency conditions.

[0077] In an alternative embodiment, step S104 may further include: Step S600, perform a modulo operation on the hash tag and the bin period to obtain an offset value; wherein, the bin period is the difference between the start times of two adjacent bins; Step S602, determine a preloading time interval according to the start time of the current bin, the offset value, and the start time of the next bin of the current bin.

[0078] Step S604, determine whether the target event time is within the preloading time interval.

[0079] Step S606, create the next window of the target window when the target event time is within the preloading time interval.

[0080] Step S608, write the cumulative feature into the next window of the target window.

[0081] Exemplarily, the grouped feature can be hashed to obtain an integer hash value (hash value), and the hash value is modulo-operated with the bin period to obtain an offset value Offset. Then, according to the start time BT of the current bin, the offset value Offset, and the start time NBT of the next bin of the current bin, a preloading time interval P: {BT + Offset, NBT} is determined. When the user behavior is triggered, if the target event time belongs to P, window preloading is performed. As Figure 3 and as Figure 4As shown, the initialization and writing of the next window are triggered simultaneously. The next window generated by preloading counts simultaneously with the current window and automatically becomes the current window after window switching.

[0082] In this embodiment, the process of window creation is scattered over a bucketing time through the window preloading mechanism, avoiding the problem that when using a global sliding window, all sliding windows switch simultaneously, resulting in a concentrated occurrence of window creation behavior, forming an instantaneous high CPU occupancy and reducing the stability of the Redis cluster.

[0083] The above-mentioned multiple embodiments introduce how to write target event data into the first data structure or the second data structure. Next, an exemplary introduction will be given on how the first data structure is converted into the second data structure.

[0084] In an alternative embodiment, the first data structure is converted into the second data structure through the following steps: Step S700: Determine the start time of the current bucket according to the target event time.

[0085] Step S702: Determine a plurality of target conversion buckets according to the start time of the current bucket, the preset number of buckets, and the accumulation period, and each target conversion bucket has a corresponding start time.

[0086] Step S704: Obtain the corresponding event data from the first data structure according to the start times of the plurality of target conversion buckets and write it into the plurality of target conversion buckets to obtain the second data structure.

[0087] Exemplarily, after writing the target event data into the corresponding data structure, the corresponding data structure can be read immediately to obtain the latest deduplication value. As Figure 5 shown, when reading the first data structure Zset, when the read deduplication value is greater than the preset conversion threshold L0, the Zset data can be safely converted into HLL. To ensure the atomicity of the entire process, a Redis Lua script can be used to implement this process. As Figure 6 shown, after triggering the conversion, the bucketing period can be calculated based on the preset number of buckets N and the accumulation period T. The start time of the current bucket can be calculated based on the target event time and the bucketing period, and a plurality of target conversion buckets and their corresponding start times can be determined. For each target conversion bucket, the data in the corresponding time interval can be read from the Zset in sequence and written into the bucket HLL. By analogy, the first data structure can be converted into the second data structure. The remaining expiration time can also be set for each bucket according to the event time to optimize the storage efficiency.

[0088] In this embodiment, an HLL is automatically generated based on the conversion of data in the Zset, and the Zset is securely deleted to ensure the lowest data storage occupancy at any scale. The Lua script is used to implement the process of data conversion to ensure the atomicity of the conversion process.

[0089] In an alternative embodiment, the data processing method may further include: when converting the first data structure to the second data structure, determining a conversion flag. Correspondingly, step S102 may include: Step S800, obtaining the conversion flag.

[0090] Step S802, when the conversion flag is not obtained, the storage type is the first storage type.

[0091] Step S804, when the conversion flag is obtained, the storage type is the second storage type.

[0092] Exemplarily, after all the bucket conversions are completed, the conversion flag can be set according to the timestamp when this conversion is completed, indicating that the data has been converted to the second data structure HLL storage, and the current storage type is the second storage type. The conversion flag can be used to distinguish the storage type. When the SF does not exist, the Zset storage is used by default, otherwise the HLL storage is used. Correspondingly, before each event data is written, the conversion flag SF can be obtained first. As Figure 7 shown, the storage type can be quickly determined through the SF, reducing the additional I / O generated when judging the storage type during data reading and writing. That is: when the SF does not exist, the data has not been converted, and the Zset storage can be used, otherwise the HLL storage is used.

[0093] In this embodiment, by centrally storing the SF, the occupancy of computing resources can be further reduced.

[0094] In an alternative embodiment, the conversion flag includes the completion time of the most recent conversion. Step S702 may include: Step S900, obtaining the conversion flag.

[0095] Step S902, when the conversion flag is not obtained: according to the start time of the current bucket, the preset number of buckets, and the cumulative period, determine multiple buckets within the cumulative period; determine the multiple target conversion buckets as the multiple buckets within the cumulative period.

[0096] Step S904, when the conversion flag is obtained, determine, according to the start time of the current bucket, the preset number of buckets, and the accumulation period, multiple buckets within the accumulation period whose start times are later than the completion time of the most recent conversion; determine the multiple buckets within the accumulation period whose start times are later than the completion time of the most recent conversion as the multiple target conversion buckets.

[0097] Exemplarily, as Figure 8 shown, obtain the conversion flag and determine the completion time of the most recent conversion. For those with SF, it is considered that a conversion has occurred, and the data before the completion time of the most recent conversion has all been converted. To save computing resources, only trigger the bucket conversion after the completion time of the most recent conversion. Therefore, multiple target conversion buckets can be determined according to the start time of the current bucket and the completion time of the most recent conversion. In the case where there is no SF, it is considered that no conversion has occurred, and the initial bucket start time is calculated using the full period.

[0098] In this embodiment, when triggering data conversion, the start time of data conversion is initialized based on the existing DCT, thereby solving the problem of resource waste caused by possible multiple triggers of data conversion under distributed high-concurrency conditions.

[0099] In some embodiments, the expiration time of the Zset can be reset to the local cache expiration time to ensure that under distributed conditions, after the local cache of each instance expires, the incremental data that has not been converted in time and written into the Zset can be correctly supplemented and converted into the HLL bucket.

[0100] In an alternative embodiment, the data processing method may further include: when the conversion flag is obtained, reset the expiration time of the conversion flag to the sum of the target event time and the accumulation period; and / or after writing the target event data into the first data structure or the second data structure, write the corresponding storage type into the local cache.

[0101] As Figure 7 shown, when writing data, it can first check whether there is a conversion flag Store Flag to determine the storage type. If SF does not exist, it is default to store in the first data structure Zset. If SF exists, renew SF, that is, reset its expiration time to the sum of the event time and the preset accumulation period, to ensure that the storage type can be determined based on SF at any time within the same period. As Figure 2As shown, after writing the target event data into the first data structure or the second data structure, the corresponding storage type can be returned to create a local cache, thereby further reducing the I / O for determining the storage type. For the Zset storage type, since there is a possibility of automatic conversion to the HLL storage, the local cache time can be set to a relatively small value, such as 1 s. For the HLL storage type, there is no change in the storage type within each bucketing period, so the local cache can be set to one bucketing period.

[0102] In this embodiment, by renewing the SF and / or increasing the local cache, the I / O for determining the storage type can be effectively reduced.

[0103] In an alternative embodiment, the data processing method may further include: Step S1000: Determine an atomic command according to a preset cumulative period, where the atomic command is used to obtain the deduplicated value accumulated in the first data structure; based on the atomic command, obtain the target deduplicated value from the first data structure.

[0104] Step S1002: Determine the current window in the second data structure based on the target event time; based on the current window, obtain the target deduplicated value.

[0105] As Figure 5 shown, after each write of the event data into the corresponding data structure, the corresponding data structure can be read immediately to obtain the latest deduplicated value. First, it can be queried whether there is a storage type in the local cache and the cache value is HLL. If it exists, it means using the HLL storage, and the current window can be determined according to the target time. Take the HLL count of the current window, which is the deduplicated value within the most recent period, that is, the target deduplicated value. If the local cache is the first storage type Zset, then based on the preset cumulative period, the atomic command zcount can be used to obtain the deduplicated value within the start and end times of the cumulative period, that is, the target deduplicated value. When reading Zset, the number of members with Score in [WT - T, WT] is read according to the cumulative period T. Among them, WT is the start time of the current window, ensuring that in the calculation of the relationship between features, the Zset and HLL values have the same cumulative duration.

[0106] In this embodiment, a global sliding window is used, so that the start and end times of each bucket corresponding to different event features are the same, and the start and end times of the deduplicated value within a fixed period are the same at the same moment, ensuring the accuracy of the calculation between features.

[0107] Embodiment 2 Figure 9Schematically shown is a block diagram of a data processing apparatus according to Embodiment 2 of the present application. The apparatus can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 9 shown, the apparatus 1000 may include: an acquisition module 1100, a determination module 1200, and a writing module 1300, where: The acquisition module 1100 is configured to acquire target event data, and the target event data includes target features; The determination module 1200 is configured to determine a storage type, and the storage type includes a first storage type or a second storage type; The writing module 1300 is configured to write the target event data into a first data structure or a second data structure according to the storage type; Wherein, the first data structure is used to accumulate the deduplicated values of event features of a first scale, the second data structure is used to accumulate the deduplicated values of event features of a second scale, the first scale is smaller than the second scale, and the first data structure is converted into the second data structure when the accumulated deduplicated value exceeds a preset conversion threshold.

[0108] As an optional embodiment, the first storage type corresponds to the first data structure, and the first data structure includes an ordered set; the second storage type corresponds to the second data structure, and the second data structure includes a cardinality estimation data structure.

[0109] As an optional embodiment, the event features include a grouping feature and an accumulation feature; the target event data further includes a target event time; the first data structure includes a plurality of elements, and each element has a corresponding score; Correspondingly, writing the target event data into the first data structure according to the storage type includes: Determining the grouping feature, the accumulation feature, and the target event time according to the target event data; Performing a hash operation on the grouping feature to obtain a hash tag; Determining the corresponding first data structure according to the hash tag; Writing the accumulation feature and the target event time as a new element and its corresponding score into the corresponding first data structure.

[0110] As an optional embodiment, the event features include a grouping feature and an accumulation feature; the target event data further includes a target event time; the second data structure includes a plurality of buckets, and each bucket has a corresponding start time; Correspondingly, according to the storage type, writing the target event data into a second data structure includes: Determining a grouping feature, an accumulation feature, and a target event time according to the target event data; Performing a hashing operation on the grouping feature to obtain a hash tag; Determining a corresponding second data structure according to the hash tag; Determining a start time of a corresponding bucket according to the target event time, so as to determine a current bucket from the multiple buckets; Writing the accumulation feature into the current bucket.

[0111] As an optional embodiment, determining a start time of a corresponding bucket according to the target event time includes: Performing an integer division operation on the target event time and a bucket period to obtain a bucket index number; wherein, the bucket period is a difference between start times of two adjacent buckets; Performing a multiplication operation on the bucket index number and the bucket period to obtain the start time of the current bucket.

[0112] As an optional embodiment, the apparatus 1000 is further configured to: Determining a current window according to the current bucket, where the current window includes a preset number of buckets whose start times are earlier than the current bucket; Writing the accumulation feature into the current window.

[0113] As an optional embodiment, the apparatus 1000 is further configured to: Performing a modulo operation on the hash tag and a bucket period to obtain an offset value; wherein, the bucket period is a difference between start times of two adjacent buckets; Determining a preloading time interval according to the start time of the current bucket, the offset value, and the start time of the next bucket of the current bucket; Determining whether the target event time is within the preloading time interval; Creating a next window of the target window when the target event time is within the preloading time interval; Writing the accumulation feature into the next window of the target window.

[0114] As an optional embodiment, the target event data further includes a target event time; the second data structure includes multiple buckets, each bucket having a corresponding start time; the first data structure includes multiple elements, each element having a corresponding score; wherein, the element is event data, and the score is an event time; Correspondingly, the first data structure is converted into the second data structure through the following operations: Determine the start time of the current bucket according to the target event time; Determine a plurality of target conversion buckets according to the start time of the current bucket, the preset number of buckets, and the accumulation period, and each target conversion bucket has a corresponding start time; Obtain the corresponding event data from the first data structure according to the start times of the plurality of target conversion buckets and write it into the plurality of target conversion buckets to obtain the second data structure.

[0115] As an optional embodiment, the apparatus 1000 is further configured to: Determine a conversion flag when converting the first data structure into the second data structure; Correspondingly, determine the storage type, including: Obtain the conversion flag; When the conversion flag is not obtained, the storage type is the first storage type; When the conversion flag is obtained, the storage type is the second storage type.

[0116] As an optional embodiment, the apparatus 1000 is further configured to: When the conversion flag is obtained, reset the expiration time of the conversion flag to the sum of the target event time and the accumulation period; and / or After writing the target event data into the first data structure or the second data structure, write the corresponding storage type into the local cache.

[0117] As an optional embodiment, the conversion flag includes the completion time of the most recent conversion; Correspondingly, determining a plurality of target conversion buckets according to the start time of the current bucket, the preset number of buckets, and the accumulation period, and each target conversion bucket has a corresponding start time, includes: Obtain the conversion flag; When the conversion flag is not obtained: Determine a plurality of buckets within the accumulation period according to the start time of the current bucket, the preset number of buckets, and the accumulation period; Determine the plurality of buckets within the accumulation period as the plurality of target conversion buckets; When the conversion flag is obtained, determine a plurality of buckets within the accumulation period whose start time is later than the completion time of the most recent conversion according to the start time of the current bucket, the preset number of buckets, and the accumulation period; Determine the plurality of buckets within the accumulation period whose start time is later than the completion time of the most recent conversion as the plurality of target conversion buckets.

[0118] As an alternative embodiment, the apparatus 1000 is further configured to: Determine an atomic command according to a preset accumulation period, where the atomic command is used to obtain the deduplicated value accumulated in the first data structure; based on the atomic command, obtain a target deduplicated value from the first data structure; or Determine a current window in the second data structure based on the target event time; and obtain the target deduplicated value based on the current window.

[0119] Embodiment III Figure 10 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the data processing method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers). As Figure 10 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the data processing method. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.

[0120] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0121] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal via a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0122] It should be noted that Figure 10 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.

[0123] In this embodiment, the data processing method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0124] Embodiment 4 The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data processing method in the embodiment are implemented.

[0125] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the data processing method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0126] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.

[0127] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0128] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire target event data, wherein the target event data includes target event features; determining a storage type, the storage type comprising a first storage type or a second storage type; Writing the target event data into a first data structure or a second data structure according to the storage type; Among them, the first data structure is used to accumulate the deduplication values ​​of event features of a first scale, and the second data structure is used to accumulate the deduplication values ​​of event features of a second scale. The first scale is smaller than the second scale, and the first data structure is converted to the second data structure when the accumulated deduplication value exceeds a preset conversion threshold.

2. The method according to claim 1, characterized in that The first storage type corresponds to the first data structure, which includes an ordered set; the second storage type corresponds to the second data structure, which includes a cardinality estimation data structure.

3. The method according to claim 1, characterized in that The event features include grouping features and cumulative features; the target event data also includes the target event time; the first data structure includes a plurality of elements, each element having a corresponding score; Correspondingly, according to the storage type, writing the target event data into the first data structure includes: Determining grouping characteristics, cumulative characteristics and target event time according to the target event data; Performing a hash operation on the grouping feature to obtain a hash tag; Determine a corresponding first data structure according to the hash tag; The accumulated features and the target event time are written into the corresponding first data structure as new elements and their corresponding scores.

4. The method according to claim 1, characterized in that: The event features include grouping features and cumulative features; the target event data also includes the target event time; the second data structure includes a plurality of buckets, each bucket having a corresponding start time; Correspondingly, according to the storage type, writing the target event data into the second data structure includes: Determining grouping characteristics, cumulative characteristics and target event time according to the target event data; Performing a hash operation on the grouping feature to obtain a hash tag; Determine a corresponding second data structure according to the hash tag; Determine a start time of a corresponding bucket according to the target event time, so as to determine a current bucket from the multiple buckets; Write the accumulated features into the current bucket.

5. The method according to claim 4, characterized in that Determine the start time of the corresponding bucket according to the target event time, including: An integer division operation is performed based on the target event time and the bucket period to obtain a bucket index number; wherein the bucket period is the difference between the start times of two adjacent buckets; A multiplication operation is performed based on the bucket index number and the bucket period to obtain the start time of the current bucket.

6. The method according to claim 4, characterized in that Also includes: Determine a current window according to the current bucket, where the current window includes a preset number of buckets whose start time is earlier than the current bucket; The accumulated features are written into the current window.

7. The method according to claim 6, characterized in that Also includes: Performing a modulo operation based on the hash tag and the bucket period to obtain an offset value; wherein the bucket period is the difference between the start times of two adjacent buckets; Determine a preloading time interval according to the start time of the current bucket, the offset value, and the start time of the next bucket of the current bucket; Determining whether the target event time is within the preload time interval; When the target event time is within the preload time interval, creating a next window of the target window; The accumulated features are written into a window next to the target window.

8. The method according to claim 1, characterized in that The target event data also includes the target event time; the second data structure includes a plurality of buckets, each bucket having a corresponding start time; the first data structure includes a plurality of elements, each element having a corresponding score; wherein the element is event data, and the score is event time; Correspondingly, the first data structure is converted into the second data structure by the following operations: Determine the start time of the current bucket according to the target event time; Determine a plurality of target conversion buckets according to the start time of the current bucket, the preset number of buckets and the accumulation period, each target conversion bucket having a corresponding start time; According to the start time of the multiple target conversion buckets, corresponding event data is obtained from the first data structure and written into the multiple target conversion buckets to obtain the second data structure.

9. The method according to claim 8, characterized in that Also includes: In case of converting the first data structure to the second data structure, determining a conversion mark; Accordingly, determine the storage type, including: Obtaining the conversion mark; When the conversion mark is not obtained, the storage type is the first storage type; When the conversion mark is obtained, the storage type is the second storage type.

10. The method according to claim 9, characterized in that Also includes: When the conversion mark is obtained, the expiration time of the conversion mark is reset to the sum of the target event time and the accumulation period; and / or After the target event data is written into the first data structure or the second data structure, the corresponding storage type is written into the local cache.

11. The method according to claim 8, characterized in that The conversion mark includes the completion time of the most recent conversion; Correspondingly, according to the start time of the current bucket, the preset number of buckets and the accumulation period, a plurality of target conversion buckets are determined, each target conversion bucket having a corresponding start time, including: Get conversion token; In the case where the conversion mark is not obtained: determining a plurality of buckets within the accumulation period according to the start time of the current bucket, the preset number of buckets and the accumulation period; determining the plurality of buckets within the accumulation period as the plurality of target conversion buckets; When the conversion mark is obtained, multiple buckets whose start time is later than the completion time of the most recent conversion within the accumulation period are determined according to the start time of the current bucket, the preset number of buckets and the accumulation period; and multiple buckets whose start time is later than the completion time of the most recent conversion within the accumulation period are determined as the multiple target conversion buckets.

12. The method according to any one of claims 1 to 11, characterized in that Also includes: Determining an atomic command according to a preset accumulation period, wherein the atomic command is used to obtain the accumulated deduplication value of the first data structure; Based on the atomic command, obtaining a target deduplication value from the first data structure; or determining a current window in the second data structure based on the target event time; Based on the current window, the target deduplication value is obtained.

13. A data processing device, characterized in that: The device comprises: An acquisition module, used for acquiring target event data, wherein the target event data includes target event features; A determination module, configured to determine a storage type, wherein the storage type includes a first storage type or a second storage type; A writing module, used for writing the target event data into a first data structure or a second data structure according to the storage type; Among them, the first data structure is used to accumulate the deduplication values ​​of event features of a first scale, and the second data structure is used to accumulate the deduplication values ​​of event features of a second scale. The first scale is smaller than the second scale, and the first data structure is converted to the second data structure when the accumulated deduplication value exceeds a preset conversion threshold.

14. A computer device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 12 are implemented.