Network independent visitor quantity statistical method based on super logarithm data structure
By combining hyperlogarithmic data structure with time bucketing mechanism, the existing UV statistical methods in terms of calculation efficiency and accuracy are solved, efficient and flexible statistics on the number of independent visitors in the network are achieved, storage costs and misjudgment rates are reduced, and data reliability and accuracy are improved.
Patent Information
- Application Number
- CN202510441264.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-01
AI Technical Summary
The existing UV statistics methods have insufficient calculation efficiency and time dimensions, making it difficult to achieve real-time, efficient and high-precision statistics on the number of independent visitors in the network, especially when the calculation overhead is huge during non-full-cycle data statistics, and traditional methods have the problem of high misjudgment rate for identification of unlogged users.
The method of combining hyperlogarithmic data structure with time bucketing mechanism is adopted, by establishing multiple time buckets to store visitor identification and time stamps, the hyperlogarithmic algorithm is used to estimate the number of independent visitors, and combining multi-dimensional identification generation strategies and confidence grading, dynamically correct statistical results, supporting flexible time dimensions and statistical requirements under complex conditions.
It achieves extremely low storage overhead (12KB can count up to 64th power data volumes up to 2), significantly improves statistical efficiency and accuracy, supports flexible statistics at any time span, reduces server load and misjudgment rate, and improves data reliability and robustness.
Smart Images

Figure CN120407959A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network unique visitor count statistics, and particularly relates to a method for counting network unique visitor numbers based on a hyperloglog data structure. Background Art
[0002] In today's digital age, the booming development of network services has given rise to a vast amount of user access data. Accurately counting the number of unique visitors (UV) to a website is crucial for website operators, as it is a key basis for evaluating website traffic, user activity, market influence, and formulating precise marketing strategies. With the exponential growth of the user scale and the complexity of business scenarios, real-time, efficient, and high-precision UV statistics have become the core requirements for enterprises to optimize operations and enhance user experience. However, traditional statistical methods have significant limitations in terms of technical implementation, resource consumption, and statistical flexibility, making it difficult to meet the stringent requirements of modern Internet businesses.
[0003] Currently, the main technical means relied on by mainstream UV statistical methods include: the Cookie method tracks visitors by setting a unique identifier (such as a Cookie ID) in the user's browser; the IP address method uses the user's IP address as an identifier to record access behaviors; the database precise deduplication method directly stores user identifiers (such as IDs, IP addresses) and calculates the deduplicated quantity in real time.
[0004] Although the above methods can partially meet the requirements in specific scenarios, there are still problems of low calculation efficiency. When counting UV, it is necessary to screen and deduplicate the full amount of data, and the time complexity is linearly related to the data scale, making it difficult to achieve real-time response. For example, counting data in the order of hundreds of millions may take several minutes and cannot support dynamic business decisions.
[0005] In addition, if the above methods need to count the UV of non-integer periods (such as a partial period of a certain day or a cross-period interval), it is necessary to rescan the full amount of data and deduplicate it, resulting in huge calculation overhead and insufficient statistical capabilities in the time dimension. Summary of the Invention
[0006] The present invention provides a method for counting network unique visitor numbers based on a hyperloglog data structure to solve the problems of low calculation efficiency caused by screening and deduplicating the full amount of data when counting network unique visitor numbers, and the problem of insufficient statistical capabilities when counting non-integer period data.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A method for counting network unique visitor numbers based on a hyperloglog data structure, comprising:
[0009] Establishing a plurality of said hyperloglog data structures, and obtaining a plurality of time buckets corresponding to time;
[0010] According to the time when the visitor accesses the network terminal, store the identifier and timestamp of the visitor correspondingly in the time bucket.
[0011] According to the statistical time of the network unique visitor count and the cycle time of the time bucket, extract the number of non-repeating visitor identifiers within the statistical time to obtain the network unique visitor count.
[0012] The method for counting the network unique visitor count based on the hyperloglog data structure in the present invention further includes the following additional technical features:
[0013] The identifier of the visitor is specifically:
[0014] When the visitor logs in to the network terminal, obtain the identifier according to the ID of the visitor;
[0015] When the visitor does not log in to the network terminal, obtain the identifier through the IP address of the visitor.
[0016] When the visitor does not log in to the network terminal, obtaining the identifier through the IP address of the visitor is specifically:
[0017] Obtain the identifier according to the device identifier in combination with the IP address, and correspondingly set the first confidence level, or,
[0018] Obtain the identifier according to the user agent string in combination with the IP address, and correspondingly set the second confidence level, or,
[0019] Obtain the identifier according to the fingerprint information of the browser in combination with the IP address, and correspondingly set the third confidence level, or,
[0020] Obtain the identifier according to the IP address, and correspondingly set the fourth confidence level;
[0021] Among them, the first confidence level, the second confidence level, the third confidence level, and the fourth confidence level decrease in sequence.
[0022] Extracting the number of non-repeating visitor identifiers within the statistical time is specifically:
[0023] According to the generation method of the identifier, obtain the similarity threshold. When the similarity of the visitor identifier is greater than or equal to the similarity threshold, it is determined as a repeated visitor identifier;
[0024] Among them, when obtaining the identifier according to the ID of the visitor, the similarity threshold is 100%; when the visitor does not log in to the network terminal, the similarity threshold is the confidence level corresponding to the generation method of the identifier.
[0025] Extract the number of non-repeating visitor identifiers within the statistical time according to the statistical time of the network unique visitor count and the cycle time of the time binning, specifically as follows:
[0026] Obtain the time period corresponding to the statistical time according to the cycle time;
[0027] For the time period consistent with the cycle time, obtain the number of non-repeating visitor identifiers through the corresponding time binning;
[0028] For the time period less than the cycle time, call the visitor identifiers in the corresponding time period within the time binning through the time stamp to obtain a temporary hyperloglog data structure, and obtain the number of non-repeating visitor identifiers in the temporary hyperloglog data structure;
[0029] Obtain the number of non-repeating visitor identifiers within the statistical time according to the number of non-repeating visitor identifiers corresponding to the time period.
[0030] The method for counting the network unique visitor number based on the hyperloglog data structure further includes:
[0031] When counting the network unique visitor numbers of multiple network terminals,
[0032] According to the hyperloglog data structure corresponding to the network terminal, in combination with the statistical time, call the visitor identifier, merge to obtain a temporary hyperloglog data structure, and obtain the number of non-repeating visitor identifiers in the temporary hyperloglog data structure.
[0033] Extract the number of non-repeating visitor identifiers within the statistical time to obtain the network unique visitor number, specifically as follows:
[0034] According to the number of non-repeating visitor identifiers, perform correction through the slope coefficient and the intercept coefficient to obtain the network unique visitor number;
[0035] Wherein the slope coefficient and the intercept coefficient are obtained by fitting multiple historical data.
[0036] The time binning is specifically as follows:
[0037] Set one time bin for one cycle time, or,
[0038] Set multiple time bins for one cycle time, and the number of the time bins is determined according to the number of the identifiers of the visitors stored within the cycle time.
[0039] The present invention also provides a storage medium,
[0040] A computer program is stored on the storage medium, and when the computer program is executed, the steps of the method for counting the number of network independent visitors based on the hyperloglog data structure are implemented.
[0041] The present invention further provides a processing device, including:
[0042] A memory for storing a computer program;
[0043] A processor for implementing the steps of the method for counting the number of network independent visitors based on the hyperloglog data structure when executing the computer program.
[0044] Due to the adoption of the above technical solution, the beneficial effects obtained by the present invention are as follows:
[0045] 1. In the present invention, a plurality of the hyperloglog data structures are established, and a plurality of time buckets are obtained according to the time correspondence; according to the time when the visitor accesses the network terminal, the identifier and timestamp of the visitor are correspondingly stored in the time bucket. By combining the hyperloglog data algorithm with the time bucket mechanism, the present invention realizes the ultimate optimization of storage resources. The hyperloglog data algorithm replaces the storage requirement for all user identifiers of the traditional method with extremely low storage overhead (only 12KB is required to count up to 2 to the 64th power of data volume), directly reducing the storage cost by more than 99% compared with the traditional method (which requires storing all data).
[0046] In addition, the time bucket mechanism further avoids redundant storage of all data by storing data in segments on demand. For example, after bucketing by time granularity such as daily, weekly, monthly, etc., only an independent hyperloglog data structure needs to be maintained for each time period, rather than storing all the identifiers and timestamps of all visitors, so as to allocate storage resources on demand and further reduce the storage overhead of the server. This design realizes the efficient utilization of storage resources while ensuring the statistical accuracy.
[0047] 2. In the present invention, according to the time when the visitor accesses the network terminal, the identifier and timestamp of the visitor are correspondingly stored in the time bucket; according to the statistical time for counting the number of network independent visitors and the cycle time of the time bucket, the number of non-repeated visitor identifiers within the statistical time is extracted to obtain the number of network independent visitors. By combining the flexible time bucket design with the timestamp recording mechanism, the present invention significantly improves the statistical ability in the time dimension.
[0048] First, it supports bucketing by periodic time, meeting the requirements of conventional time series analysis (such as counting peak traffic hours by day and analyzing user active cycles by week). According to the statistical time of the number of network independent visitors and the periodic time of the time bucketing, the corresponding time bucket is selected to quickly obtain the visitor identifier for the corresponding statistical time without full-scanning the data of the entire cycle. In addition, the collaborative design of time bucketing and timestamps also supports flexible combinations of complex time conditions (such as counting the number of network independent visitors within a fragmented time). By quickly filtering the visitor identifiers corresponding to the fragmented time within the time bucket through timestamps, the flexibility and practicality of statistics are further expanded.
[0049] 3. As a preferred implementation manner of the present invention, the number of non-repeating visitor identifiers within the statistical time is extracted as follows: According to the generation method of the identifier, a similarity threshold is obtained. When the similarity of the visitor identifier is greater than or equal to the similarity threshold, it is determined as a repeating visitor identifier; wherein, when the identifier is obtained according to the ID of the visitor, the similarity threshold is 100%; when the visitor has not logged in to the network terminal, the similarity threshold is the confidence level corresponding to the generation method of the identifier. First, for logged-in users, the system directly generates an identifier using the unique user ID and sets the similarity threshold to 100% to ensure the uniqueness of the identifier for the same user on different devices or sessions, completely avoiding double counting.
[0050] For non-logged-in users, the algorithm adopts a multi-dimensional data fusion strategy: combining the IP address, device fingerprint (such as hardware identifier), browser information (such as User-Agent), and browser fingerprint (such as screen resolution, plugin information) to generate a composite identifier, and filtering based on the level of confidence. For example, if a user only provides an IP address, the confidence level of its identifier is the lowest, and it is regarded as a repeating visitor only when the lower similarity threshold condition is met, thereby significantly reducing the probability that multiple users under the same IP address are misjudged as the same visitor. This multi-dimensional precise control solves the misjudgment problem of traditional methods (such as simply relying on the IP address).
[0051] 4. As a preferred implementation manner of the present invention, according to the number of non-repeating visitor identifiers, it is corrected through the slope coefficient and the intercept coefficient to obtain the number of network independent visitors; wherein the slope coefficient and the intercept coefficient are obtained by fitting multiple historical data. The present invention introduces an error correction mechanism driven by historical data: By fitting the relationship between the historical true UV and the estimated value of the hyperlogarithmic algorithm through a linear regression model, the slope coefficient (a) and the intercept coefficient (b) are obtained, and based on this, the current statistical result is dynamically corrected to make the statistical result closer to the true value. This dynamic calibration compensates for the inherent error of the probability algorithm (hyperlogarithmic algorithm) and finally realizes the dual improvement of data reliability and robustness. Description of the Drawings
[0052] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0053] Figure 1 It is a schematic flowchart of the method for counting the number of independent visitors to a network based on the HyperLogLog data structure under an embodiment of the present invention. Detailed implementation manners
[0054] In order to more clearly illustrate the overall concept of the present invention, the following will be described in detail by way of examples in conjunction with the drawings of the specification.
[0055] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0056] As Figure 1 shown, a method for counting the number of independent visitors to a network based on the HyperLogLog data structure includes:
[0057] S100: Establish a plurality of the HyperLogLog data structures and obtain a plurality of time buckets according to the time correspondence.
[0058] The core objective of this step is to combine the time bucket mechanism with the HyperLogLog data structure, utilize the low storage characteristic of the HyperLogLog algorithm, store data in time slices, avoid storing user identifiers in full, efficiently store a large amount of time series data; extract independent visitor data for a specific time period as needed through time buckets, without full-scale scanning, improve the statistical efficiency, and flexibly support time dimension statistics.
[0059] Among them, the HyperLogLog data structure refers to a probabilistic cardinality estimation algorithm used to efficiently count the number of unique elements in a data set (such as the independent visitor UV). It can be understood that the HyperLogLog data structure supports the statistics of data volume up to the 64th power of 2 with a very small storage space (such as 12KB), and the error rate is only ±0.81%. The method of storing user identifiers by the present invention using the HyperLogLog data structure significantly reduces the storage cost.
[0060] Time Buckets refers to dividing data into multiple independent time periods according to a preset time granularity (such as daily, weekly, monthly), and allocating an independent HyperLogLog data structure for each time period. It can realize the time dimension management of data and support the rapid extraction of statistical results for a specific time period.
[0061] In this step, when establishing the hyperloglog data structure, it is necessary to determine the time granularity and the bucketing strategy. Select the granularity of time bucketing according to business requirements (such as monthly, weekly, daily, hourly), and dynamically adjust the number of buckets according to the number of visitors within the time period corresponding to the time bucket (such as increasing the bucket density when the data volume surges).
[0062] Initialize the time bucket and the hyperloglog data structure, assign a unique identifier to each time bucket (such as bucket_2023-10-01), and initialize the corresponding hyperloglog data structure. It should be noted that the identifier of the time bucket is generated according to the corresponding time period, and bucket_2023-10-01 is the time bucket with a time period of October 1, 2023. If there are multiple time buckets corresponding to a time period, the identifier of each time bucket includes a time identifier and a sequence number identifier (for example, bucket_2023-10-01-01), where the time identifier is 2023-10-01 and the sequence number identifier is 01. The time period corresponding to the time bucket is marked by the time identifier, and the time buckets within the same time period are distinguished by the sequence number identifier.
[0063] In a specific embodiment, a certain platform needs to count the number of independent visitors per hour during the "Double Eleven" period. At this time, the preset time granularity is hourly, and the time buckets are divided hourly (such as from bucket_2023-11-11-00 to bucket_2023-11-11-23). The visitor identifiers per hour are added to the hyperloglog data structure of the corresponding bucket in real time. If it is necessary to count the number of independent visitors from 10:00 to 12:00, only the data of bucket_2023-11-11-10 and bucket_2023-11-11-11 need to be merged, without scanning the entire volume of data.
[0064] In this step, through the hyperloglog data structure and the time bucket, 12KB can support a data volume of up to 2 to the 64th power, reducing the storage cost. When extracting the number of independent visitors within the statistical time, there is no need to scan the entire volume of data. The statistics can be completed directly by merging the data of the target time bucket, improving the statistical efficiency.
[0065] This step combines the hyperloglog algorithm with the time bucket, solving the core defects of traditional methods in terms of storage, statistical efficiency, and time flexibility, providing an efficient and lightweight data basis for subsequent steps such as identifier storage and error correction.
[0066] S200: Store the identifier and timestamp of the visitor corresponding to the time when the visitor accesses the network terminal in the time bucket.
[0067] The core goal of this step is to accurately map the visitor's identification and time information to the corresponding time bucket, ensuring that data can be quickly located by time dimension during subsequent statistics. By associating timestamps with buckets, efficient extraction of any time period can be achieved.
[0068] It should be noted that the Visitor Identifier is used to uniquely identify the visitor's feature combination, ensuring the uniqueness of the identification of different visitors and avoiding duplicate counting.
[0069] Timestamps are used to record the precise time a visitor accesses a network terminal (e.g., ISO 8601 format: 2023-10-01T14:30:00). Timestamp parsing can be used to determine the time bucket to which a visitor belongs. Timestamps can also be used to precisely locate a visitor within a time bucket, enabling the counting of visitors within a specific timeframe.
[0070] In this step, when a visitor accesses a terminal, the access timestamp is first parsed and the bucket to which it belongs is determined based on the preset time granularity (such as day or hour). For example, if the bucket granularity is "day", the timestamp 2023-10-01T14:30:00 will be matched to the bucket_2023-10-01.
[0071] The association between the identifier and the timestamp is stored in the corresponding bucket, and the identifier is converted into a hash value to adapt to the cardinality estimation algorithm of the hyperlogarithmic data structure. The timestamp is indexed to facilitate subsequent fast retrieval of buckets by time range.
[0072] It is understandable that when the bucketing period ends (such as 24:00 for daily bucketing), the system automatically archives the current bucketed data and creates a new bucket. At this time, only the new bucket needs to be maintained, which reduces the data maintenance workload.
[0073] In one specific embodiment, a platform needs to monitor the number of unique visitors per hour in real time to optimize server resource allocation. Buckets are created by hour (e.g., bucket_2023-10-01-14). Visitor A visits the website at 2:30 PM. The system parses the timestamp and matches it to bucket_2023-10-01-14. Visitor A's identifier is stored in the corresponding bucket. This method can directly count the data in bucket_2023-10-01-14 and quickly obtain the number of unique visitors from 2:00 PM to 3:00 PM.
[0074] In this step, through the precise association between identifiers and time, the exact matching of each identifier with time is ensured. Through the timestamp index, the target bucket can be directly located without full-scale scanning, improving the ability to locate data in the time dimension and solving the core defects of traditional methods in terms of time association accuracy and storage efficiency.
[0075] S300: According to the statistical time of the number of network unique visitors and the cycle time of the time bucket, extract the number of non-repeated visitor identifiers within the statistical time to obtain the number of network unique visitors.
[0076] The core objective of this step is to quickly and accurately extract the number of unique visitors within a specific time period based on the time bucket mechanism and the hyperloglog algorithm. Through the parallel processing of time buckets and the cardinality estimation ability of the hyperloglog algorithm, full-scale data scanning is avoided, and the statistical time is shortened. Through the comparison and positioning of the statistical time and the cycle time, the statistics of the number of network unique visitors for any time span (such as partial cycles, cross-cycle intervals) are supported.
[0077] In this step, through the user-specified statistical time (such as "2023-10-01 10:00 to 2023-10-01 12:00"), all relevant buckets covering the target time are determined according to the cycle of the time bucket (such as hourly buckets). For example: If the statistical time is "10:00 - 12:00", then bucket_2023-10-01-10 and bucket_2023-10-01-11 are matched.
[0078] Based on the comparison of visitor identifiers within the corresponding time bucket, the number of non-repeated visitor identifiers is determined, which is the preliminarily estimated number of network unique visitors. After directly adopting it or making corrections, the number of network unique visitors can be obtained.
[0079] In a specific embodiment, a certain platform needs to count the number of unique visitors from 10:00 to 12:00 on "Double Eleven" to optimize server resources. Buckets bucket_2023-11-11-10 and bucket_2023-11-11-11 are matched. The data structures of the two buckets are merged, and the estimated number of non-repeated visitor identifiers is 9,800,000.
[0080] In this step, by merging the data structures of the buckets, the statistical time is shortened. While traditional methods require full-scale scanning and deduplication of hundreds of millions of data, which is more time-consuming in minutes; while this step (such as only 1 second for counting hundreds of millions of data). In addition, it can provide precise support for complex time ranges. For example, when counting a local time period of a certain day (such as 10:00 - 12:00), only the data of 2 buckets need to be merged, without scanning all the data, reducing the server load.
[0081] As a preferred embodiment of the present invention, the identification of the visitor is specifically as follows:
[0082] When the visitor logs in to the network terminal, the identification is obtained according to the ID of the visitor;
[0083] When the visitor does not log in to the network terminal, the identification is obtained through the IP address of the visitor.
[0084] The core objective of this embodiment is to achieve a coarse-grained statistics of unlogged users with minimal resource consumption while ensuring the accuracy of logged user statistics through a differential identification generation strategy. Through identification, duplicate counting is avoided, and the accuracy of unique visitor number statistics is improved.
[0085] For the generation of logged user identification, according to the user ID (such as account ID), the user ID is directly used as the identification, or a unique hash value is generated through a hash function (such as SHA-256) as the identification. It can be understood that the unique ID ensures the uniqueness of the identification of the same user on different devices or sessions, completely avoiding duplicate counting.
[0086] For the generation of unlogged user identification, an identification is generated according to the IP address of the visitor. At this time, under certain conditions, the device used under this IP address or other information should be combined for comprehensive generation of the identification to reduce the misjudgment that multiple users under the same IP are the same visitor.
[0087] In a specific embodiment, an e-commerce platform needs to count the number of unique visitors on a certain day. The user logs in with the account ID, and the identification is user_12345, which is stored in the corresponding bucket. The identification of logged users is unique, ensuring no duplicate counting.
[0088] As a preferred embodiment of this embodiment, when the visitor does not log in to the network terminal, the identification is obtained through the IP address of the visitor, specifically as follows:
[0089] The identification is obtained according to the device identification in combination with the IP address, and the first confidence level is correspondingly set, or,
[0090] The identification is obtained according to the user agent string in combination with the IP address, and the second confidence level is correspondingly set, or,
[0091] The identification is obtained according to the fingerprint information of the browser in combination with the IP address, and the third confidence level is correspondingly set, or,
[0092] The identification is obtained according to the IP address, and the fourth confidence level is correspondingly set;
[0093] Among them, the first confidence level, the second confidence level, the third confidence level, and the fourth confidence level decrease in sequence.
[0094] The core objective of this embodiment is to enhance the uniqueness and statistical accuracy of unregistered user identification through multi-dimensional data fusion and confidence grading mechanism. Specifically, the probability that multiple users (such as in a home network or public place) under the same IP address are misjudged as the same visitor is significantly reduced. Dynamically adjust the credibility of the identification according to the reliability (confidence) of the data source to balance precision and resource consumption.
[0095] It can be understood that device identification refers to hardware information (such as device model, IMEI, MAC address). User agent string refers to browser type, operating system version, etc. Browser fingerprint refers to screen resolution, plugin list, time zone, font information, etc.
[0096] The uniqueness of device identification (such as IMEI) combined with the IP address can uniquely identify the access of the same device under the same network. For example, device ID: ABC123 + IP: 192.168.1.100.
[0097] The combination of user agent strings (such as browser + operating system) can distinguish different device types, but may be forged. For example, User-Agent: Chrome / 100 + IP: 192.168.1.100.
[0098] The multi-dimensional features of browser fingerprint (such as resolution + plugins) can approximately uniquely identify the device, but there is a risk of dynamic change. For example, resolution: 1920x1080 + plugin list: AdBlock + IP: 192.168.1.100.
[0099] Using only the IP address has the highest misjudgment rate (such as multiple people sharing the IP in a home network).
[0100] It can be understood that according to the credibility ranking of the above identification generation methods, the first confidence level > the second confidence level > the third confidence level > the fourth confidence level.
[0101] In a specific embodiment, in a home network, 3 users use the same IP address (192.168.1.100) to access the website, through mobile phone, tablet and computer respectively.
[0102] For user A (mobile phone), the identification is device ID: PHONE123 + IP: 192.168.1.100 (the first confidence level). For user B (tablet), the identification is User-Agent: SamsungTablet + IP: 192.168.1.100 (the second confidence level). For user C (computer), only the IP address is used (the fourth confidence level).
[0103] In this embodiment, through multi-dimensional identification, the misjudgment rate of traditional IP address statistics is reduced. In addition, only the necessary features of high-confidence identification (such as device ID, fingerprint) are stored, rather than the full amount of data. For example, the storage space of device identification + IP is only twice that of pure IP, but the misjudgment rate is reduced by 80%. Through hierarchical confidence, high-confidence identification (such as device ID) is preferentially processed, reducing redundant calculations for low-confidence identification. This embodiment solves the problems of high misjudgment rate and resource consumption in the statistics of unlogged users through a multi-dimensional feature combination and confidence grading mechanism.
[0104] Specifically, the number of non-repeating visitor identifications within the statistical time is extracted, specifically:
[0105] According to the generation method of the identification, a similarity threshold is obtained. When the similarity of the visitor identification is greater than or equal to the similarity threshold, it is determined as a repeating visitor identification;
[0106] Among them, when the identification is obtained according to the ID of the visitor, the similarity threshold is 100%; when the visitor does not log in to the network terminal, the similarity threshold is the confidence corresponding to the generation method of the identification.
[0107] The core objective of this embodiment is to accurately distinguish repeating and independent visitor identifications through differential threshold settings. Specifically, the unique user ID ensures the uniqueness of the identification. Setting the threshold to 100% can completely avoid duplicate counting and there is no misjudgment for logged-in users. Through the confidence grading threshold, the balance between accuracy and statistical efficiency is achieved, reducing the probability of misjudgment for multiple users under the same IP, and the misjudgment of unlogged users is controllable. The threshold is dynamically adjusted according to the reliability of the identification generation method to support statistical requirements in different scenarios (such as high-precision or resource-constrained scenarios).
[0108] Convert the identification (such as device ID: ABC123 + IP: 192.168.1.100) into a hash value (such as SHA-256). Calculate the similarity through the matching degree of the hash value or feature similarity (such as the matching degree of device identification, browser fingerprint).
[0109] For logged-in users, the identification is generated by the unique user ID, and the similarity threshold is 100%. It is determined as repeating only when the hash values are exactly the same (such as a user accessing across devices).
[0110] For unlogged users, for device identification + IP: the first 95% of the hash value matches (such as 243 bits are the same in a 256-bit hash). For user agent + IP: the first 90% of the hash value matches. For browser fingerprint + IP: the first 85% of the hash value matches. For pure IP: the first 80% of the hash value matches.
[0111] It is understandable that for logged-in users, it is determined as a duplicate only when the hash values are exactly the same. For the pure IP generation method, if the first 80% of the hash value matches, it is determined as the same visitor. Other identifiers are determined according to the corresponding thresholds.
[0112] Merge the identifiers determined as duplicates and only retain the count of unique identifiers.
[0113] In a specific embodiment, in a home network, 3 users use the same IP address (192.168.1.100) to access a website, through a mobile phone (device identifier), a tablet (user agent), and a computer (pure IP) respectively.
[0114] Among them, user A (mobile phone), the identifier is device ID: PHONE123 + IP: 192.168.1.100 (the first confidence level, threshold 95%). User B (tablet) is identified as User-Agent: SamsungTablet + IP: 192.168.1.100 (the second confidence level, threshold 90%). User C (computer) only uses the IP address (the fourth confidence level, threshold 80%).
[0115] User A and B are identified as independent visitors due to high-confidence identifiers. The hash value of user C's identifier (pure IP) may match the hash values of user A / B in the first 80% (such as the same IP address, but other features are missing), so it is determined as a duplicate.
[0116] In this embodiment, through hierarchical thresholds, high-confidence identifiers (such as device IDs) are processed preferentially, reducing redundant calculations of low-confidence identifiers. Only the hash values of the identifiers and the threshold parameters need to be stored, and there is no need to retain the full amount of original feature data.
[0117] It should be noted that in this embodiment, during the identifier generation process, the high-confidence generation method is preferentially adopted, that is, according to the generation method of the visitor ID, when the visitor ID cannot be generated, the generation method based on the device identifier combined with the IP address is preferentially adopted. After the high-confidence identifier is generated, no subsequent identifier generation method is performed.
[0118] As a preferred embodiment of the present invention, according to the statistical time of the network independent visitor count and the cycle time of the time bucket, extract the number of non-duplicate visitor identifiers within the statistical time, specifically:
[0119] According to the cycle time, obtain the time period corresponding to the statistical time;
[0120] For the time period consistent with the cycle time, through the corresponding time bucket, obtain the number of non-duplicate visitor identifiers;
[0121] For a time period shorter than the cycle time, call the visitor identifiers corresponding to the corresponding time period within the time bucket through the time stamp to obtain a temporary super-logarithmic data structure, and obtain the number of non-repeated visitor identifiers in the temporary super-logarithmic data structure;
[0122] According to the number of non-repeated visitor identifiers corresponding to the time period, obtain the number of non-repeated visitor identifiers within the statistical time.
[0123] The core objective of this embodiment is to support efficient UV statistics for any time span, while taking into account the flexibility of complete cycles and partial cycles. Specifically, for the efficiency of complete cycle statistics, directly merge the bucket data to avoid redundant calculations. For the feasibility of partial cycle statistics, through time stamp filtering and temporary data structures, accurately count non-integer cycle data. This embodiment does not require full-scale scanning, only processes necessary data, and reduces computing and storage overhead.
[0124] Time period parsing and bucket matching, statistical time (such as "2023-10-01 10:00-12:30"). Assume the bucket is hourly (bucket_2023-10-01-10, bucket_2023-10-01-11, bucket_2023-10-01-12).
[0125] Classify the time period into complete cycles: 10:00-11:00 (corresponding to bucket_2023-10-01-10); partial cycles: 12:00-12:30 (the first 30 minutes of data need to be extracted from bucket_2023-10-01-12).
[0126] For the extraction of complete cycle data, merge the super-logarithmic structures of the complete cycle buckets (such as bucket_2023-10-01-10 and bucket_2023-10-01-11), and directly obtain the cardinality.
[0127] For the processing of partial cycle data. Screen the identifiers with time stamps within 12:00-12:30 from the bucket (such as bucket_2023-10-01-12). Store the screened identifiers into a temporary super-logarithmic data structure and count the cardinality.
[0128] Add the cardinalities of the complete cycle and the partial cycle to obtain the final UV quantity.
[0129] In this embodiment, the traditional method requires a full-scale scan and filtering of timestamps, which takes minutes; this step reduces the statistics time to milliseconds (for example, it only takes 1 second to count data in the order of hundreds of millions) through bucket merging and timestamp filtering. Through timestamp filtering and a temporary hyperloglog data structure, non-integer cycle data can be accurately counted. In addition, there is no need to store full-scale identifiers, and the statistics can be completed by only merging buckets and the temporary structure. Moreover, by avoiding a full-scale scan and only processing data for necessary time periods, the computational overhead is reduced.
[0130] Generally speaking, through the collaborative design of bucket matching, timestamp filtering, and temporary structure construction in this embodiment, the efficiency and accuracy problems of the traditional method in non-integer cycle statistics are solved, and efficient UV statistics for any time span are achieved. By separating the processing logics of complete cycles and partial cycles, resource consumption is significantly reduced while ensuring accuracy.
[0131] As a preferred embodiment of the present invention, the method for counting the number of independent network visitors based on the hyperloglog data structure further includes:
[0132] When counting the number of independent network visitors of multiple network terminals,
[0133] According to the hyperloglog data structure corresponding to the network terminal, combined with the statistical time, the visitor identifier is called, and a temporary hyperloglog data structure is merged to obtain the number of non-repeated visitor identifiers in the temporary hyperloglog data structure.
[0134] The core objective of this embodiment is to support multi-terminal collaborative statistics of the number of independent visitors while ensuring high efficiency and accuracy. Specifically, through cross-terminal data merging, visitor data scattered on different terminals is uniformly counted to avoid duplicate counting (such as the same user accessing different servers or APPs). Through the merging characteristics of the hyperloglog data structure, a full-scale data scan is avoided, and millisecond-level statistics are achieved.
[0135] First, according to multiple network terminals (such as servers A, B, and C) and the statistical time (such as "2023-10-01"). The hyperloglog data structures (such as bucket_2023-10-01_A, bucket_2023-10-01_B) are extracted from the corresponding time buckets of each terminal.
[0136] Use the pfmerge command of Redis to merge the hyperloglog data structures of multiple terminals into a temporary structure (temp_bucket). The merging operation of the hyperloglog data structure retains the cardinality estimation of unique identifiers, and the error rate is still ≤ 0.81%.
[0137] Obtain the cardinality of the temporary structure through the pfcount command to get the UV quantity after multi-terminal merging.
[0138] In a specific embodiment, the order system (Server A) and the product browsing system (Server B) of a certain e-commerce platform need to count the total UV for a certain day. Extract bucket_2023-10-01_A from Server A and bucket_2023-10-01_B from Server B. Generate temp_bucket through pfmerge. After merging, the UV is 100,000 (A) + 150,000 (B) - overlapping users. Assuming the overlapping users are 30,000, the final UV is 220,000. It can avoid double counting (such as when a user accesses both systems simultaneously), and the statistical time only takes 1 second.
[0139] In this embodiment, the traditional method requires full-scale data merging (such as merging and deduplicating the identification lists of Server A and B). This embodiment merges through a hyperloglog data structure, reducing the processing time. In addition, when a new terminal is added, only its hyperloglog data structure needs to be added to the merging list, without modifying the existing architecture. It can merge multi-terminal data in real time to meet dynamic business requirements (such as adding servers during a promotion period).
[0140] As a preferred embodiment of the present invention, the number of non-repeating visitor identifiers within the statistical time is extracted to obtain the number of network independent visitors, specifically:
[0141] Based on the number of non-repeating visitor identifiers, it is corrected through the slope coefficient and the intercept coefficient to obtain the number of network independent visitors;
[0142] Wherein the slope coefficient and the intercept coefficient are obtained by fitting multiple historical data.
[0143] The core objective of this embodiment is to correct the estimation error of the hyperloglog algorithm and improve the accuracy of the independent visitor number statistics. There is an error of ±0.81% in the cardinality estimation of the hyperloglog data structure (for example, when estimating 1 million UVs, the actual value may be 992,000 or 1,008,000). The error can be further reduced through the historical data fitting coefficient.
[0144] Stat the true UV values (obtained through full-scale data scanning) and the corresponding estimated values for the past N statistics. The data example is as follows:
[0145] Statistical period Estimated UV True UV 2023-01 1,000,000 995,000 2023-02 1,200,000 1,190,000
[0146] Taking the estimated value as the independent variable (X) and the true UV as the dependent variable (Y), the fitting equation:
[0147] Y = aX + b
[0148] Wherein, the slope coefficient (a) is calculated through the least squares method. For example, in the example data, a ≈ 0.995.
[0149] Intercept coefficient (b), such as b≈ -5,000 in the example data.
[0150] When the estimated value for the current statistic is 1,500,000
[0151] Corrected UV = (0.995×1,500,000)+(-5,000)=1,487,500
[0152] It should be noted that new data points are supplemented in each cycle, and the coefficients are refitted to adapt to the changes. Or, the coefficients are fitted separately for different business scenarios (such as promotions, non - promotions).
[0153] This embodiment uses a history - data - driven linear correction model, which reduces the estimation error and significantly improves the statistical accuracy while maintaining high efficiency.
[0154] As a preferred embodiment of the present invention, the time bucketing is specifically as follows:
[0155] One time bucket is set for one cycle time, or,
[0156] Multiple time buckets are set for one cycle time, and the number of the time buckets is determined according to the number of identifiers of the visitors stored within the cycle time.
[0157] The core objective of this embodiment is to optimize the storage and query efficiency through a dynamic bucketing strategy while adapting to data volume fluctuations.
[0158] Determine the cycle time (such as hourly level) and the preset identifier quantity threshold (such as 1 million). If the identifier quantity ≤ threshold: Set 1 bucket (fixed mode). If the identifier quantity > threshold: Split into multiple buckets as needed (dynamic mode).
[0159] When dynamically bucketing, execute the equal - proportion splitting rule to evenly distribute the identifiers to multiple buckets (for example, when the data volume in 1 hour is 2 million, split into 2 buckets, each with 1 million). Or, execute hash bucketing, and distribute to different buckets according to the high - order digits of the identifier hash value (such as the first 2 hash value digits determine the bucket number).
[0160] Dynamically adjust the number of buckets according to the real - time data volume (such as automatic splitting triggered by sudden traffic). In addition, buckets can be merged during off - peak periods to release resources.
[0161] In a specific embodiment, during the "Double Eleven" period on an e - commerce platform, the number of visitor identifiers in a certain hour surges to 3 million (exceeding the threshold of 1 million). This hour is split into 3 buckets (each bucket has approximately 1 million identifiers).
[0162] When counting the hourly UV, only the super-logarithmic data structures of 3 buckets need to be merged, instead of processing a single extremely large data bucket, which reduces the query time and resource consumption.
[0163] In the traditional method, the data volume of a single bucket is too large at high traffic, resulting in waste of storage and computing resources. In this embodiment, through dynamic splitting, the resource utilization rate is improved. Moreover, merging multiple small buckets is faster than processing a single large data bucket. In addition, it avoids the accumulation of identifiers in a single bucket and reduces the probability of hash collisions.
[0164] Generally speaking, dynamically binding the number of buckets to the data volume not only ensures the query performance but also realizes the elastic allocation of resources, solves the problems of resource waste and query latency of traditional fixed buckets in high-traffic scenarios, and at the same time maintains the efficient utilization of resources in low-traffic scenarios.
[0165] The present invention also provides a storage medium.
[0166] A computer program is stored on the storage medium, and when the computer program is executed, the steps of the method for counting the number of unique visitors to a network based on the super-logarithmic data structure are implemented.
[0167] Therefore, any effects of the method for counting the number of unique visitors to a network based on the super-logarithmic data structure can be achieved, which will not be elaborated here.
[0168] The present invention also provides a processing device, including:
[0169] A memory for storing a computer program;
[0170] A processor for implementing the steps of the method for counting the number of unique visitors to a network based on the super-logarithmic data structure when executing the computer program.
[0171] Therefore, any effects of the method for counting the number of unique visitors to a network based on the super-logarithmic data structure can be achieved, which will not be elaborated here.
[0172] What is not described in the present invention can be implemented by adopting or referring to the existing technologies.
[0173] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.
[0174] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for counting the number of independent visitors to a network based on a hyperloglog data structure, characterized in that, Including: Create multiple of the above hyperloglog data structures, and obtain multiple time buckets according to the corresponding time. According to the time when the visitor accesses the network terminal, store the identifier and timestamp of the visitor in the corresponding time bucket. According to the statistical time of the network unique visitor count and the cycle time of the time bucket, extract the number of non-repeating visitor identifiers within the statistical time to obtain the network unique visitor count.
2. The method for counting the number of independent network visitors based on the hyperloglog data structure according to claim 1, wherein The identifier of the visitor is specifically: When the visitor logs in to the network terminal, obtain the identifier according to the ID of the visitor. When the visitor does not log in to the network terminal, obtain the identifier through the IP address of the visitor.
3. The method for counting the number of independent network visitors based on the hyperloglog data structure according to claim 2, characterized in that, When the visitor does not log in to the network terminal, obtaining the identifier through the IP address of the visitor is specifically: Obtain the identifier by combining the device identifier with the IP address, and correspondingly set the first confidence level, or Obtain the identifier by combining the user agent string with the IP address, and correspondingly set the second confidence level, or Obtain the identifier by combining the fingerprint information of the browser with the IP address, and correspondingly set the third confidence level, or Obtain the identifier according to the IP address, and correspondingly set the fourth confidence level; Among them, the first confidence level, the second confidence level, the third confidence level, and the fourth confidence level decrease in sequence.
4. The method for counting the number of independent visitors to a network based on a hyperloglog data structure according to claim 3, wherein Extracting the number of non-repeating visitor identifiers within the statistical time is specifically: According to the generation method of the identifier, obtain the similarity threshold. When the similarity of the visitor identifier is greater than or equal to the similarity threshold, it is determined as a repeating visitor identifier. Among them, when obtaining the identifier according to the ID of the visitor, the similarity threshold is 100%; when the visitor does not log in to the network terminal, the similarity threshold is the confidence level corresponding to the generation method of the identifier.
5. The method for counting the number of independent network visitors based on the hyperloglog data structure according to claim 1, wherein According to the statistical time of the network unique visitor count and the cycle time of the time bucket, extracting the number of non-repeating visitor identifiers within the statistical time is specifically: According to the cycle time, obtain the time period corresponding to the statistical time; For the time period consistent with the cycle time, obtain the number of non-repeating visitor identifiers through the corresponding time bucket; For the time period less than the cycle time, call the visitor identifiers in the corresponding time period within the time bucket through the timestamp to obtain a temporary hyperloglog data structure, and obtain the number of non-repeating visitor identifiers in the temporary hyperloglog data structure; According to the number of non-repeating visitor identifiers corresponding to the time period, obtain the number of non-repeating visitor identifiers within the statistical time.
6. The method for counting the number of independent visitors to a network based on a hyperloglog data structure according to claim 1, wherein It also includes: When counting the network unique visitor counts of multiple network terminals, According to the hyperloglog data structure corresponding to the network terminal, combine with the statistical time, call the visitor identifier, and merge to obtain a temporary hyperloglog data structure, and obtain the number of non-repeating visitor identifiers in the temporary hyperloglog data structure.
7. The method for counting the number of independent network visitors based on the hyperloglog data structure according to claim 1, wherein Extracting the number of non-repeating visitor identifiers within the statistical time to obtain the network unique visitor count is specifically: According to the number of non-repeating visitor identifiers, perform correction through the slope coefficient and intercept coefficient to obtain the network unique visitor count; The slope coefficient and the intercept coefficient are obtained by fitting multiple historical data.
8. The method for counting the number of independent network visitors based on the hyperloglog data structure according to claim 1, wherein The time binning is specifically as follows: One time bin is set for one cycle time, or Multiple time bins are set for one cycle time, and the number of the time bins is determined according to the number of identifiers of the visitors stored within the cycle time.
9. A storage medium, characterized in that A computer program is stored on the storage medium, and when the computer program is executed, the steps of the method for counting the number of network independent visitors based on the hyperloglog data structure according to any one of claims 1 to 8 are implemented.
10. A processing device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the method for counting the number of network independent visitors based on the hyperloglog data structure according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Cross-platform live broadcast visitor data fusion and routing load balancing method and system
CN119420696A