A user data protection method and system of an online promotion platform

By analyzing the promotion factors and conversion rates of user browsing records, and utilizing hash ring mapping and diluted similarity preference significance, the problem of user data being easily cracked in online promotion platforms was solved, achieving higher data storage security.

CN120850336BActive Publication Date: 2026-03-24BEIJING RIYUEZHEN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

User data with similar preferences on online promotion platforms is easily compromised by attackers due to these similar preferences, and existing technologies are insufficient to effectively protect user data privacy.

Method used

By analyzing user browsing history, quantifying the promotion factors and conversion rates of browsing tags, and mapping users to a hash ring using a hash function, the distribution of user data on the hash ring is adjusted based on the dilution of similar preferences significance by data breach risk factors, thereby reducing the risk of data breach.

Benefits of technology

It improves the storage security of user data in online promotion platforms, reduces the probability of data from users with similar preferences being compromised at storage nodes, and enhances data protection effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850336B_ABST
    Figure CN120850336B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data protection, and discloses a user data protection method and system for an online promotion platform, which comprises the following steps: obtaining browsing records of a plurality of users in the online promotion platform as user data of the users; obtaining promotion factors of the browsing labels on the users; determining conversion proportions of the browsing labels on the users; obtaining preference significances of the users on the browsing labels; mapping the users to hash rings through a hash function to obtain initial distribution positions of the users on the hash rings of the browsing labels; determining data cracking risk factors of the users on the hash rings of the browsing labels; diluting similar preference significances between the users under the same browsing label to obtain final distribution positions of the users on the hash rings of the browsing labels; and protecting the user data in combination with nodes on the hash rings. The application aims to solve the problem that user data with similar preferences in a promotion platform is easily cracked by attackers due to the similar preferences.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data protection, in particular to a user data protection method and system of an online promotion platform. BACKGROUND

[0002] The total server of the online promotion platform stores the browsing records of the users based on the storage nodes in a distributed manner, and the browsing records are analyzed and promoted by the online promotion platform as user data. During the promotion process, the browsing records can reflect the preference degree of the users for different browsing tags in the browsing process, so as to promote online. A large amount of user data is stored in a distributed manner by the online promotion platform. Since the analysis and promotion process involves the browsing records of the users, the browsing records have certain privacy, and therefore it is necessary to protect the data of the browsing records in the user data in the online promotion platform.

[0003] During the distributed storage of the browsing records of the users by the online promotion platform, the prior art analyzes the user preferences by using a consistent hash algorithm, and stores the user data of the users with similar preferences in the same node. Since the users with similar preferences are similar, when the user data is attacked, the users with similar preferences will provide certain reference for the attacker to further piece together, so that the user data with similar preferences is more likely to be cracked by the attacker. However, the user preferences are usually not single, and the dilution of multiple types of preferences can reduce the similarity between the user preferences, so as to improve the protection effect of the user data of the online promotion platform. SUMMARY

[0004] The present application provides a user data protection method and system of an online promotion platform to solve the problem that the user data with similar preferences in the existing promotion platform is easily cracked by the attacker due to similar preferences. The technical solution adopted is as follows:

[0005] The present application provides a user data protection method and system of an online promotion platform to solve the problem that the user data with similar preferences in the existing promotion platform is easily cracked by the attacker due to similar preferences. The technical solution adopted is as follows:

[0006] Obtain the browsing records of each browsing of a plurality of users in the online promotion platform as user data of each user. The browsing records include browsing tags, browsing time, browsing duration and consumption records.

[0007] Analyze the distribution of the browsing time under the same browsing tag in each browsing of the same user and the correlation between the browsing durations to obtain the promotion factors of each browsing tag for each user. Determine the conversion rate of each browsing tag for each user based on the distribution of the browsing records of the same browsing tag of the same customer and the change of the consumption records. Obtain the preference significance of the user for each browsing tag based on the promotion factors and the conversion rate.

[0008] According to the preference significance of each user to each browsing label, each user is mapped to a hash ring through a hash function to obtain an initial distribution position of the user on the hash ring of each browsing label; the relationship between the initial distribution positions of the users with similar preference significance to the same browsing label on the hash ring is analyzed to determine a data cracking risk factor of the user on the hash ring of each browsing label;

[0009] Based on the data cracking risk factor of the user on the hash ring of each browsing label, the difference between the initial distribution positions of the same user on the hash rings of different browsing labels is combined to dilute the similar preference significance between the users under the same browsing label to obtain a final distribution position of the user on the hash ring of each browsing label, and the nodes on each hash ring are combined for user data protection.

[0010] Optionally, the method for obtaining the promotion factor of each browsing label to each user comprises the following specific method:

[0011] For a plurality of browsing records of any user, a plurality of browsing records of any browsing label are obtained, the browsing time of each browsing record is obtained as the corresponding time in each day as the distribution time of each browsing record, the mean and standard deviation of the distribution time of all browsing records of the user and the browsing label are obtained;

[0012] For all browsing records of the user and the browsing label, the difference between the browsing times of any two adjacent browsing records is obtained as the time interval between the latter browsing record and the adjacent former browsing record, and the calculation method of the promotion factor Q of the browsing label to the user is as follows:

[0013]

[0014] Wherein, σ(t) represents the standard deviation of the distribution time of all browsing records of the user and the browsing label, I represents the number of browsing records of the user and the browsing label, t i represents the distribution time of the i-th browsing record of the user and the browsing label, represents the mean of the distribution time of all browsing records of the user and the browsing label, Δt i represents the time interval between the i-th browsing record and the adjacent former browsing record of the user and the browsing label, T i represents the browsing duration of the i-th browsing record of the user and the browsing label, T i-1 represents the browsing duration of the i-1-th browsing record of the user and the browsing label; || represents the absolute value function; softmax() represents the weight normalization function; exp() represents the exponential function with the natural constant as the base.

[0015] Optionally, the conversion ratio of each browsing label to each user is obtained by the following specific method:

[0016]

[0017] wherein P represents a conversion rate of any browsing label to any user, I represents a number of browsing records of the user for the browsing label, a i represents a consumption amount of the i-th browsing record of the user for the browsing label, represents a mean value of time intervals corresponding to all browsing records of the user for the browsing label except the first browsing record, Δt i represents a time interval between the i-th browsing record and the adjacent previous browsing record of the user for the browsing label.

[0018] Optionally, the method for obtaining the preference significance of the user for each browsing label comprises the following specific method:

[0019] a three-dimensional sample space of the promotion factor, the conversion rate and the browsing label is constructed, the promotion factor and the conversion rate of each browsing label to each user are mapped, the user and the browsing label are taken as a combination, one combination corresponds to one sample point, and a plurality of sample points in the three-dimensional sample space are obtained; density clustering is performed on the sample points of the same browsing label, a Euclidean distance between the sample points is used as a distance measurement, and a plurality of class clusters of each browsing label are obtained.

[0020] based on the distribution of the sample points in the class cluster of the same browsing label, the class cluster preference of the sample point corresponding to the user for the browsing label is obtained.

[0021] the product of the promotion factor and the conversion rate of any browsing label to each user is linearly normalized, and the obtained result is taken as a conversion factor of the browsing label to each user; the product of the class cluster preference of any user for the browsing label and the conversion factor of the browsing label to the user is taken as the preference significance of the user for the browsing label.

[0022] Optionally, the method for obtaining the class cluster preference of the sample point corresponding to the user for the browsing label comprises the following specific method:

[0023] a centroid of any class cluster of any browsing label is obtained, the class cluster preference of the sample point corresponding to the user for the browsing label is obtained according to the distance between any sample point in the class cluster and the centroid and the mean value of the distances between the sample point and each sample point in the class cluster; the class cluster preference is in a negative correlation with the distance from the centroid, and the class cluster preference is in a negative correlation with the mean value of the distances from each sample point in the class cluster.

[0024] Optionally, the method for obtaining the initial distribution position of the user on the hash ring of each browsing label comprises the following specific method:

[0025] For any browsing tag, a hash ring is constructed, and users existing in the browsing record of the browsing tag are taken as mapping users of the browsing tag. The preference significance of each mapping user to the browsing tag is processed by a hash function in a consistent hash algorithm to obtain a hash value of each mapping user to the browsing tag. The hash value is mapped to the hash ring to obtain an initial distribution position of each mapping user on the hash ring of the browsing tag.

[0026] Optionally, the data cracking risk factor of the user on the hash ring of each browsing tag is obtained by the following method:

[0027] The preference significance between users corresponding to the same cluster of sample points under the same browsing tag is analyzed to obtain similar users of the user under the browsing tag and preference similarity weights of each similar user.

[0028] The distance of any user and any similar user of the user along the hash ring on the hash ring of any browsing tag is taken as a position difference between the similar user and the user. The ratio of the position difference and the length of 1 / 2 of the hash ring is taken as a position difference degree between the similar user and the user. The difference obtained by subtracting the position difference degree from 1 is taken as a position similarity degree between the similar user and the user.

[0029] The position similarity degrees between all similar users and the user are weighted and summed by using the preference similarity weights to obtain the data cracking risk factor of the user on the hash ring of the browsing tag.

[0030] Optionally, the method for obtaining the similar users of the user under the browsing tag and the preference similarity weights of each similar user includes the following method:

[0031] Users corresponding to other sample points in the cluster of the sample point corresponding to any user under any browsing tag are taken as similar users of the user under the browsing tag. The absolute value of the difference in preference significance of the user and any similar user of the user to the browsing tag is obtained. The difference obtained by subtracting the absolute value of the difference in preference significance from 1 is taken as a preference similarity factor between the similar user and the user. The preference similarity factors of all similar users and the user are weight normalized to obtain the preference similarity weights of each similar user.

[0032] Optionally, the method for obtaining the final distribution position of the user on the hash ring of each browsing tag includes the following method:

[0033] The data cracking risk factor and the initial distribution position of any user on the hash ring of each browsing tag are obtained. For any browsing tag, the data cracking risk factor of the user on the hash ring of the browsing tag is smaller than that of other browsing tags of the user on the hash ring of the browsing tag, which is taken as a dilution browsing tag of the user on the browsing tag.

[0034] The product of the difference between the data breach risk factor of the user at the browsing tag and any dilution browsing tag, the position of the dilution browsing tag, and the preference significance of the user to the dilution browsing tag is taken as the adjusted position of the user at the browsing tag affected by the dilution browsing tag;

[0035] The mean of the corresponding adjusted positions of all dilution browsing tags is obtained, and the mean of the initial distribution position of the browsing tag and the initial adjusted position of the user at the browsing tag are calculated as the initial adjusted position of the user at the browsing tag;

[0036] The data risk breach factor is recalculated based on the initial adjusted position of the user at the browsing tag, and if the recalculated data risk breach factor is still greater than or equal to the risk breach threshold, the adjusted position is iteratively obtained based on the initial adjusted position until the data risk breach factor is less than the threshold, and the adjusted position of the user at the browsing tag when the iteration is stopped is taken as the final distribution position of the user at the browsing tag on the hash ring.

[0037] The present application also provides a user data protection system of an online promotion platform, which comprises a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the above method when executing the computer program.

[0038] The beneficial effects of the present application are: the present application obtains the browsing records of users in the online promotion platform as user data; by analyzing the browsing records of users, the promotion factor of the browsing label to the user is quantified through the similar relationship of the time period distribution of different days of browsing time and the change of the browsing time length, to reflect the time sequence change characteristics of the user to the browsing label; the overall intensive distribution of the browsing time can reflect the frequent browsing characteristics of the user to the browsing label, the conversion ratio is obtained by combining the change of the consumption record, to quantify the preference conversion of the user to the browsing label, and based on the promotion factor and the conversion ratio, the preference significance of the user to the browsing label is quantified by combining the browsing label, to provide a basis for subsequent user data mapping of overall preference of the user; the initial distribution position is obtained by mapping the preference significance of the user by the consistent hash algorithm through the hash ring, and based on the initial distribution position of the other users with similar preference significance on the hash ring under the corresponding browsing label, the feature that the closer the initial distribution position is, the easier the user data with similar preference significance is cracked is considered, the data cracking risk factor is quantified, to provide a basis for subsequent adjustment of the distribution position to reduce the data cracking risk; the initial distribution position of the user on the hash ring of different browsing labels is referred to and adjusted, the preference significance of the browsing label with larger data risk cracking factor is diluted by the other browsing labels with larger preference significance and smaller data risk cracking factor, the probability that the similar preference significance is stored in the same node and is easily deduced and cracked is reduced, and then the protection of the user data and the storage security of the user data are improved under the condition that the preference of the user in the online promotion platform is similar. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only show some embodiments of the present application, and all other drawings obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0040] Figure 1 A user data protection method flow chart of an online promotion platform provided by an embodiment of the present application. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0042] Please refer to Figure 1Fig. 1 shows a flow chart of a method for protecting user data of an online promotion platform according to an embodiment of the present application, which comprises the following steps:

[0043] In step S001, the browsing records of each browsing of a plurality of users in the online promotion platform are obtained as the user data of each user.

[0044] The purpose of the embodiment is to dilute the similar preferences of the user data of similar preference users in the online promotion platform in the storage process by relying on other unique preferences of the users, so as to reduce the possibility that the user data of the users is easily cracked by attackers due to the similarity of the storage nodes and the storage mode under similar preferences, and thus the data protection effect is poor. Therefore, the browsing records of a plurality of users in the online promotion platform are first obtained, and the browsing records constitute the user data.

[0045] It should be further explained that the online promotion platform promotes related data based on the browsing process of the users, so the tags of the contents such as goods and videos browsed by the users are obtained in the browsing records, which are used to classify the contents browsed by the users. The browsing time of each browsing can reflect the browsing state and the browsing frequency of the users, and the browsing duration can reflect the preferences of the users for the contents. The consumption records combined with the browsing time can reflect the habits of the users for consuming the preferred contents, and further reflect the preferences of the users. Therefore, the related data such as the browsing tags, the browsing time, the browsing duration and the consumption records are obtained in the browsing records.

[0046] Specifically, the browsing records of each browsing of a plurality of users are recorded in the total server of the online promotion platform for analysis and promotion. The browsing records include the tags of the contents such as goods and videos browsed by the users in the browsing webpages, which are used as the browsing tags. The time when each browsing starts is used as the browsing time, and the difference between the time when each browsing ends and the time when it starts is obtained as the browsing duration of each browsing. At the same time, the consumption occurs in the browsing process, and the consumption amount of the corresponding browsing tags with the consumption is recorded as the consumption records. Therefore, for each browsing of the users, the corresponding browsing tags, the browsing time, the browsing duration and the consumption records are obtained, which together constitute the browsing records of each browsing. All the browsing records of the same user constitute the user data of the user.

[0047] In step S002, the distribution of the browsing time under the same browsing tags in each browsing of the same user and the correlation between the browsing durations are analyzed to obtain the promotion factors of each browsing tag for each user. The distribution of the browsing records of the same browsing tags of the same user is determined in combination with the change of the consumption records to determine the conversion proportion of each browsing tag for each user. The preference significance of the users for each browsing tag is obtained based on the promotion factors and the conversion proportion.

[0048] It should be noted that the user's browsing record under the same browsing tag can reflect the user's preference for the browsing tag, and the closer the distribution of the browsing time in different days for different browsing of the same browsing tag, that is, the browsing time is mostly distributed in the same time period of each day, the stronger the correlation between the user and the browsing time for the browsing tag. On this basis, analyzing the browsing time length change of the user for the same browsing tag in different browsing, the stable or increasing change trend of the browsing time length can reflect the gradual increase of the user's preference for the corresponding browsing tag. Therefore, the time performance characteristics of the user's preference for the browsing tag are quantified, and used as a promotion factor of the browsing tag for the user.

[0049] Preferably, in an embodiment of the present application, the distribution of the browsing time of the same browsing tag in each browsing of the same user and the correlation between the browsing time length are analyzed to obtain the promotion factor of each browsing tag for each user, including the specific method:

[0050] For a plurality of browsing records of any user, a plurality of browsing records of any browsing tag are obtained, the browsing time of each browsing record is obtained, and the corresponding time in each day is obtained, that is, the corresponding 24-hour time in each day of each browsing record is obtained as the distribution time of each browsing record. The mean and standard deviation of the distribution time of all browsing records of the user and the browsing tag are obtained.

[0051] Further, for all browsing records of the user and the browsing tag, the difference between the browsing time of any two adjacent browsing records is obtained as the time interval between the adjacent previous browsing record and the subsequent browsing record (the difference obtained by subtracting the previous browsing time from the subsequent browsing time, not based on the distribution time to calculate the difference). The calculation method of the promotion factor Q of the browsing tag for the user is:

[0052]

[0053] Wherein, σ(t) represents the standard deviation of the distribution time of all browsing records of the user and the browsing tag, I represents the number of browsing records of the user and the browsing tag, t i represents the distribution time of the i-th browsing record of the user and the browsing tag, represents the mean of the distribution time of all browsing records of the user and the browsing tag, Δt i represents the time interval between the i-th browsing record of the user and the browsing tag and the adjacent previous browsing record, T i represents the browsing time length of the i-th browsing record of the user and the browsing tag, T i-1represents the time length of the i-1th browsing record of the user for the browsing label; || represents an absolute value function; softmax() represents a weight normalization function, and the normalization object is each browsing record of the user for the browsing label except the first browsing record exp() represents an exponential function with a natural constant as a base, an exp(-x) model is adopted in the embodiment to present an inverse proportional relationship and normalization processing, x is an input of the model, and an implementer can set an inverse proportional function and a normalization function according to actual conditions.

[0054] It is required to be explained that the smaller the standard deviation is, the more concentrated the distribution time is, and the weight is constructed by the deviation of the distribution time and the time interval, the smaller the deviation of the distribution time is and the smaller the time interval is, the greater the reference of the change of the time length of the browsing record is, that is, on the basis of the concentrated distribution time of the browsing record, the increase of the time length in a short time can better reflect the promotion factor, and the greater the growth of the time length under the adjacent browsing record is, the greater the preference of the user in time sequence is, and then the promotion factor is obtained.

[0055] It is further required to be explained that the promotion factor reflects the change characteristics of the user in time sequence for the browsing label, and the distribution of the browsing record of the user for the same browsing label, that is, the distribution of the browsing time can reflect the frequent browsing degree of the user for the corresponding browsing label, and the change of the consumption record of each browsing is combined, the preference conversion of the user for the browsing label is reflected through the frequent browsing and the consumption condition, that is, the more frequent the browsing is and the more the change of the consumption record exists, the higher the conversion rate of the corresponding browsing label to the user is, that is, the higher the conversion proportion is.

[0056] Preferably, in an embodiment of the present application, the conversion proportion of each browsing label to each user is determined according to the distribution of the browsing record of the same browsing label of the same customer, and the change of the consumption record, and the specific method comprises:

[0057] For all browsing records of any user for any browsing label, the calculation method of the conversion proportion P of the browsing label to the user is:

[0058]

[0059] Wherein, I represents the number of the browsing record of the user for the browsing label, a i represents the consumption amount of the i-th browsing record of the user for the browsing label, represents the mean value of the time interval corresponding to all browsing records of the user for the browsing label except the first browsing record, Δt irepresents the time interval between the i-th browsing record of the user for the browsing label and the adjacent previous browsing record; || represents the absolute value function; exp() represents the exponential function with the natural constant as the base, and the embodiment adopts an exp(-x) model to present an inverse proportional relationship and normalization processing, x being the input of the model, and the implementer can set the inverse proportional function and the normalization function according to the actual situation.

[0060] It is required to be explained that, by taking the consumption record as the weight, the greater the consumption amount, the higher the reflection degree of the browsing time change in the corresponding browsing record on the preference conversion, and the smaller the mean value and the smaller the overall fluctuation of the time interval, the more frequent the browsing, and the larger the overall browsing record quantity, indicating that the overall high-frequency browsing, and thus the conversion ratio is quantified, that is, the high-frequency browsing is accompanied by a higher conversion rate.

[0061] It is further required to be explained that, based on the promotion factor and the conversion ratio, the preference significance of the user for the browsing label is comprehensively reflected, the promotion factor reflects the performance characteristics of the user in time sequence for the browsing label, the conversion ratio reflects the overall preference of the user for the browsing label, and the three-dimensional space is established in combination with the browsing label itself for clustering analysis, and then the preference significance of the user in the corresponding browsing label is quantified based on the user distribution in the cluster and the promotion factor and the conversion ratio, so as to obtain the preference significance of the user for each browsing label, so as to reflect the preference performance of the user and provide a basis for subsequent construction of the hash ring based on the preference and mapping of the user data.

[0062] Preferably, in an embodiment of the present application, based on the promotion factor and the conversion ratio, the preference significance of the user for each browsing label is obtained, including the specific method:

[0063] A three-dimensional sample space of the promotion factor, the conversion ratio and the browsing label is constructed, and the promotion factor and the conversion ratio of each browsing label to each user are mapped, so that the user and the browsing label are combined as a group, one group corresponds to one sample point, and a plurality of sample points in the three-dimensional sample space are obtained; the sample points of the same browsing label are clustered by DBSCAN, the Euclidean distance between the sample points is used for distance measurement, and a plurality of clusters of each browsing label are obtained.

[0064] Further, the centroid of any cluster of any browsing label is obtained, the distance between any sample point in the cluster and the centroid is obtained, and the mean value of the distance between the sample point and each sample point in the cluster is obtained, and the product of the distance from the centroid and the mean value of the distance from each sample point in the cluster is inversely proportional to the normalization result, which is used as the cluster preference of the corresponding user for the browsing label. The embodiment adopts an exp(-x) model to present an inverse proportional relationship, exp() represents the exponential function with the natural constant as the base, x is the input of the model, and the implementer can set the inverse proportional function according to the actual situation.

[0065] Further, the product of the promotion factor and the conversion ratio of the browsing label for each user is linearly normalized to obtain a conversion factor of the browsing label for each user; and the product of the cluster preference of any user for the browsing label and the conversion factor of the browsing label for the user is taken as the preference significance of the user for the browsing label.

[0066] It should be noted that the more the sample point deviates from the centroid in the cluster, and the greater the distance between the sample point and other sample points, the more the sample point deviates to the periphery in the cluster, and the smaller the preference feature of the browsing label in the cluster; in combination with the promotion factor and the conversion ratio of the corresponding user of the sample point, both can reflect the preference feature of the user for the browsing label, that is, the time preference enhancement feature, that is, the frequent browsing feature, so as to quantify the preference significance.

[0067] At this point, by analyzing the browsing records of the user, the promotion factor of the browsing label for the user is quantified through the similar relationship of the time period distribution of the browsing time on different days and the change of the browsing time, to reflect the time sequence change feature of the user for the browsing label; the frequent browsing feature of the user for the browsing label can be reflected by the overall dense distribution of the browsing time, the conversion ratio is obtained in combination with the change of the consumption record, to quantify the preference conversion of the user for the browsing label, and based on the promotion factor and the conversion ratio, the preference significance of the user for the browsing label is quantified in combination with the browsing label, to provide a basis for subsequent user data mapping of the overall preference of the user.

[0068] Step S003, according to the preference significance of the user for each browsing label, mapping each user to the hash ring through a hash function to obtain the initial distribution position of the user on the hash ring of each browsing label; analyzing the relationship between the initial distribution positions of the users with similar preference significance for the same browsing label on the hash ring to determine the data cracking risk factor of the user on the hash ring of each browsing label.

[0069] It should be noted that by using the consistent hash algorithm, the hash ring is constructed and the user is mapped according to its characteristics, and the related user data of the mapped user is stored in a distributed manner based on the nodes on the hash ring, which can reduce the amount of cached and deleted data of the stored data when the number of nodes increases or decreases. In the online promotion platform, the number of browsing labels is fixed, and a hash ring is constructed for each type of browsing label, and the user is mapped on the hash ring through the preference significance of the user for the browsing label. First, the initial distribution position of the user on the hash ring of each browsing label is obtained.

[0070] Preferably, in an embodiment of the present application, according to the preference significance of the user for each browsing label, each user is mapped to the hash ring through a hash function to obtain the initial distribution position of the user on the hash ring of each browsing label, which includes the following specific method:

[0071] For any browsing tag, the users existing in the browsing record of the browsing tag are mapped users of the browsing tag, the preference significance of each mapped user to the browsing tag is processed by a hash function in a consistent hashing algorithm, the hash value of each mapped user to the browsing tag is obtained by mapping the hash value to a hash ring, and then the initial distribution position of each mapped user on the hash ring of the browsing tag is obtained; the consistent hashing algorithm is a known technology, and the existing steps in the hash function mapping to the hash ring are not described herein.

[0072] It should be further explained that the users with similar preference significance of the same browsing tag may belong to the same node on the hash ring, and due to the similar preference significance, the users are more likely to be attacked or cracked in the data protection process due to the cracking of other users, and therefore, based on the relationship between the initial distribution positions of the users on the hash ring of the same browsing tag, the similar relationship between the preference significances is combined to analyze the possibility that the initial distribution positions of the users on the hash ring of each browsing tag are attacked or cracked, so as to quantify the data cracking risk factor and provide a basis for subsequent multi-browsing tag dilution of the overall hash ring.

[0073] Preferably, in an embodiment of the present application, the relationship between the initial distribution positions of the users with similar preference significance of the same browsing tag on the hash ring is analyzed, and the data cracking risk factor of the users on the hash ring of each browsing tag is determined, including the following specific method:

[0074] For the initial distribution position of any user on the hash ring of any browsing tag, the users corresponding to other sample points in the cluster where the sample point corresponding to the user is located are obtained as the similar users of the user under the browsing tag; the absolute value of the difference between the preference significances of the user and any similar user to the browsing tag is obtained, the difference obtained by subtracting the absolute value of the difference from 1 is taken as the preference similarity factor of the similar user and the user; the preference similarity factors of all similar users and the user are weight normalized to obtain the preference similarity weight of each similar user.

[0075] Further, the distance of the user and any of the similar users along the hash ring on the hash ring of the browsing label is obtained as the location difference of the similar user and the user, the ratio of the location difference and the 1 / 2 circumference of the hash ring is taken as the location difference degree of the similar user and the user, and the difference obtained by subtracting the location difference degree from 1 is taken as the location proximity degree of the similar user and the user; the hash ring is distributed with a plurality of nodes, and each initial distribution position belongs to each node; the location proximity degrees of all similar users and the user are weighted and summed with the preference proximity weight, and a judgment condition is added in the summation process; if the initial distribution positions of the user and the similar user belong to the same node on the hash ring of the browsing label, the weighted summation is normal, if not, the product of the preference proximity weight and the location proximity degree is multiplied by 0, that is, there is no attack cracking relationship between different nodes, and the sum value obtained is taken as the data cracking risk factor of the user on the hash ring of the browsing label.

[0076] It should be noted that by analyzing the difference between the initial distribution positions of the users in the same cluster under any browsing label on the hash ring, and taking the corresponding preference significance difference as the weight, the smaller the preference significance difference, the more similar the preferences, the closer the initial distribution positions on the hash ring, and the closer the initial distribution positions belong to the same node, the user data of the similar users in the same node is cracked, and the user data of the user is cracked under the premise that the user data of the similar users is cracked, the user data of the user is cracked by the attacker with slight deduction due to the similar preference and the close initial distribution position, and the data cracking risk factor is quantified.

[0077] So far, the preference significance of the user is mapped to the initial distribution position on the hash ring by using the consistent hash algorithm, and based on the initial distribution positions of the users with similar preference significance on the hash ring under the corresponding browsing label, the closer the initial distribution positions, the easier the data of the users with similar preference significance is cracked, the data cracking risk factor is quantified, and the basis is provided for adjusting the distribution position to reduce the data cracking risk.

[0078] Step S004, based on the data cracking risk factor of the user on the hash ring of each browsing label, combining the difference between the initial distribution positions of the same user on the hash rings of different browsing labels, diluting the similar preference significance between the users under the same browsing label, obtaining the final distribution position of the user on the hash ring of each browsing label, and protecting the user data in combination with the nodes on each hash ring.

[0079] It should be noted that the user has preference significance for each browsing label, and the browsing label with larger preference significance and smaller data cracking risk factor on the corresponding hash ring has better effect on diluting the preference significance of other browsing labels as a reference, and the initial distribution position is adjusted.

[0080] Specifically, the data cracking risk factor and the initial distribution position of any user on the hash ring of each browsing tag are obtained. For any browsing tag, the data cracking risk factor of the user on the hash ring of the browsing tag is less than that of other browsing tags of the user on the hash ring of the browsing tag, which is taken as the dilution browsing tag of the user on the browsing tag. The difference between the data cracking risk factor of the user on the browsing tag and any dilution browsing tag is multiplied by the position of the dilution browsing tag, and then multiplied by the preference significance of the user to the dilution browsing tag, and the result is taken as the influence adjustment position of the user on the browsing tag affected by the dilution browsing tag. The mean value of the corresponding influence adjustment positions of all dilution browsing tags is obtained, and the mean value and the initial distribution position of the browsing tag are calculated as the initial adjustment position of the user on the browsing tag.

[0081] It should be noted that the initial distribution position on the hash ring of other browsing tags with larger preference significance and smaller data cracking risk factor is adjusted to dilute the position cracking effect caused by similar preference significance.

[0082] Further, the data cracking risk factor is recalculated based on the initial adjustment position of the user on the browsing tag, a preset risk cracking threshold is set, and the risk cracking threshold of the embodiment is 0.4. If the recalculated data cracking risk factor is still greater than or equal to the risk cracking threshold, the initial adjustment position is iteratively obtained (the adjustment position is obtained by calculating the mean value of the cumulative influence adjustment position and the initial adjustment position), until the data cracking risk factor is less than the threshold, and the adjustment position of the user on the browsing tag when the iteration stops is taken as the final distribution position of the user on the browsing tag.

[0083] Further, the final distribution position of each user on the hash ring of each browsing tag is obtained according to the above method. The hash ring has a plurality of nodes, and the nodes at the same position on the hash rings of different browsing tags correspond to the same distributed storage node. Therefore, the browsing records of the users corresponding to the plurality of final distribution positions of the corresponding nodes of each browsing tag on the hash ring of any distributed storage node are stored, that is, the relevant browsing records of the users on the corresponding browsing tag are stored according to the final distribution position, so as to realize the data protection of the user data of the online promotion platform.

[0084] So far, by referring to and adjusting the initial distribution position of the user on the hash ring of different browsing tags, the preference significance of the browsing tags with larger data risk cracking factors is diluted, the probability that the similar preference significance is easily deduced and cracked under the storage of the same node is reduced, and then the protection of the user data and the storage security of the user data are improved under the similar user preference in the online promotion platform.

[0085] Another embodiment of the present application provides a user data protection system of an online promotion platform, which comprises a memory, a processor and a computer program stored in the memory and running on the processor, and the processor implements the above method steps S001 to S004 when executing the computer program.

[0086] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for protecting user data on an online promotion platform, characterized in that, The method includes the following steps: The browsing records of several users on an online promotion platform are obtained each time, and are used as user data for each user; the browsing records include browsing tags, browsing time, browsing duration and consumption records; Analyze the distribution of browsing time under the same browsing tags in each browsing session of the same user, and the correlation between browsing durations to obtain the promotion factor of each browsing tag for each user; based on the distribution of browsing records of the same browsing tags for the same customer, combined with the changes in consumption records, determine the conversion ratio of each browsing tag for each user; based on the promotion factor and conversion ratio, obtain the significance of user preference for each browsing tag. Based on the significance of users' preferences for each browsing tag, each user is mapped to a hash ring using a hash function to obtain the initial distribution position of users on the hash ring of each browsing tag; the relationship between the initial distribution positions of users with similar significance of preferences for the same browsing tag on their hash rings is analyzed to determine the data breach risk factors of users on the hash ring of each browsing tag. Based on the data cracking risk factor of users on the hash ring of each browsing tab, and combined with the difference between the initial distribution positions of the same user on the hash ring of different browsing tabs, the significance of similar preferences among users under the same browsing tab is diluted to obtain the final distribution position of users on the hash ring of each browsing tab, and user data protection is carried out in combination with the nodes on each hash ring. The specific method for obtaining the data breach risk factor of the user on the hash ring of each browsing tag is as follows: Analyze the preference significance among users in the same cluster of corresponding sample points under the same browsing tag to obtain the similar users of the user under the browsing tag and the preference similarity weight of each similar user; obtain the distance along the hash ring of any user and any similar user of any browsing tag as the position difference between the similar user and the user; take the ratio of the position difference to 1 / 2 of the perimeter of the hash ring as the degree of position difference between the similar user and the user; subtract the degree of position difference from 1 as the degree of position similarity between the similar user and the user; and sum the degree of position similarity between all similar users and the user using the preference similarity weight to obtain the data breach risk factor of the user on the hash ring of the browsing tag. The specific method for obtaining the final distribution position of a user on the hash ring of each browsing tag includes: obtaining the data breach risk factor and initial distribution position of any user on the hash ring of each browsing tag; for any browsing tag, taking other browsing tags whose data breach risk factor of the user on its hash ring is less than that of the user on the hash ring of that browsing tag as the diluted browsing tag of that user; multiplying the difference in data breach risk factor between the user on the browsing tag and any diluted browsing tag by the position of the diluted browsing tag, and then multiplying the user's preference significance for the diluted browsing tag by the result, the result is taken as the user's data breach risk factor on the hash ring of each browsing tag. The user adjusts their position on the browsing tag due to the influence of the diluted browsing tag; the average of the adjustment positions corresponding to all diluted browsing tags is obtained, and the average of the average of the average and the initial distribution position of the browsing tag is calculated as the user's initial adjustment position on the browsing tag; the data risk cracking factor is recalculated based on the user's initial adjustment position on the browsing tag; if the recalculated data risk cracking factor is still greater than or equal to the risk cracking threshold, the adjustment position is iteratively obtained from the initial adjustment position until the data risk cracking factor is less than the threshold, and the user's adjustment position on the browsing tag at the time of stopping the iteration is taken as the user's final distribution position on the hash ring of the browsing tag; The specific method for obtaining the initial distribution position of users on the hash ring of each browsing tag is as follows: construct a hash ring for any browsing tag, take the users who have the browsing tag in the browsing record as the mapped users of the browsing tag, process the significance of each mapped user's preference for the browsing tag through the hash function in the consistent hash algorithm, obtain the hash value of each mapped user for the browsing tag, and map the hash value to the hash ring to obtain the initial distribution position of each mapped user on the hash ring of the browsing tag.

2. The user data protection method for an online promotion platform according to claim 1, characterized in that, The specific method for obtaining the promotion factor of each browsing tag for each user is as follows: For any user's browsing history, obtain several browsing records for any browsing tag, and obtain the corresponding time of each browsing record within each day as the distribution time of each browsing record; obtain the mean and standard deviation of the distribution time of all browsing records for that user's browsing tag. For all browsing records of this user with this browsing tag, the difference between the browsing time of any two adjacent browsing records is obtained, which is used as the time interval between the later browsing record and the adjacent previous browsing record. This browsing tag is the promotion factor for this user. The calculation method is as follows: ; in, This represents the standard deviation of the time distribution of all browsing records for this user under this browsing tab. This indicates the number of browsing records for this user's browsing tab. This indicates the user's browsing history for that tab. The distribution time of each browsing record This represents the mean of the time distribution of all browsing records for this user under this browsing tab. This indicates the user's browsing tab number. The time interval between each browsing record and the adjacent previous browsing record. This indicates the user's browsing tab number. The duration of each browsing record. This indicates the user's browsing tab number. The duration of each browsing record; Represents the absolute value function; This represents the weight normalization function; This represents an exponential function with the natural constant as its base.

3. The user data protection method for an online promotion platform according to claim 2, characterized in that, The conversion rate of each browsing tag for each user is obtained using the following method: ; in, This indicates the conversion rate of any browsing tab to any user. This indicates the number of browsing records for this user's browsing tab. This indicates the user's browsing tab number. The amount spent on each browsing record. This represents the average of the time intervals corresponding to all browsing records for this user's browsing tab, excluding the first browsing record. This indicates the user's browsing tab number. The time interval between each browsing record and the adjacent previous browsing record.

4. The user data protection method for an online promotion platform according to claim 1, characterized in that, The specific methods for obtaining the significance of user preferences for each browsing tag are as follows: A three-dimensional sample space is constructed based on promotion factors, conversion ratios, and browsing tags. The promotion factors and conversion ratios of each user are mapped based on each browsing tag. Users and browsing tags are combined, and each combination corresponds to a sample point, resulting in several sample points in the three-dimensional sample space. Density clustering is performed on the sample points of the same browsing tag, and the distance metric is the Euclidean distance between the sample points, resulting in several clusters for each browsing tag. Based on the distribution of sample points in the same browsing tag cluster, the cluster preference of users for browsing tags corresponding to the sample points is obtained; For any browsing tag, the product of the promotion factor and conversion rate for each user is linearly normalized, and the result is used as the conversion factor for each user for that browsing tag; the product of any user's cluster preference for that browsing tag and the conversion factor for that browsing tag for that user is used as the significance of that user's preference for that browsing tag.

5. The user data protection method for an online promotion platform according to claim 4, characterized in that, The specific methods for obtaining the cluster preference of users for browsing tags corresponding to sample points are as follows: Obtain the centroid of any cluster of any browsing tag. Based on the distance between any sample point in the cluster and the centroid, and the average distance between the sample point and all sample points in the cluster, obtain the cluster preference of the user corresponding to the sample point for that browsing tag. The cluster preference is negatively correlated with the distance to the centroid and negatively correlated with the average distance to all sample points in the cluster.

6. The user data protection method for an online promotion platform according to claim 1, characterized in that, The specific method for obtaining the similar users under the browsing tags and the similarity weights of each similar user's preferences includes: Obtain users corresponding to other sample points in the cluster where the sample point corresponding to any user under any browsing tag belongs, and take them as users close to that user under that browsing tag; obtain the absolute value of the difference in the significance of the user's preference for that browsing tag with any of its close users, and subtract the absolute value of the difference in the significance of the preference from 1, and take the difference as the similarity factor between the user and the user's preferences; The similarity factors of all similar users and the user's preferences are normalized to obtain the similarity weights of each similar user's preferences.

7. A user data protection system for an online promotion platform, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the user data protection method for an online promotion platform as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data processing method, data processing device, data processing equipment and computer-readable storage medium

    CN108345659A

  • User preference analysis and identification method based on big data

    CN115187344A