A data protection method, device and storage medium for a fingerprint browser

Through dynamic hierarchical modeling and generation of adversarial networks, the contradiction between privacy protection and functional compatibility in browser fingerprint technology is solved, and the browser function integrity is maintained while protecting user privacy, and the ability to dynamically adapt to tracking system changes is achieved.

CN119622819BActive Publication Date: 2025-07-25GUANGZHOU JEEKUP INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510147329.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-25
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art cannot effectively balance privacy protection and functional compatibility, and it is difficult to achieve dynamic adaptation of tracking systems.

Method used

Dynamic hierarchical modeling method is used to divide fingerprint features into high-frequency and low-frequency layers, and a confusion strategy is adjusted in real time to enhance privacy protection and functional integrity by calculating feature entropy values and designing.

Benefits of technology

It realizes the ability to maintain browser function integrity while protecting user privacy, and has the ability to dynamically adapt to tracking system changes, solving compatibility and adaptability problems in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622819B_ABST
    Figure CN119622819B_ABST
Patent Text Reader

Abstract

The present invention provides a data protection method, device, and medium for a fingerprint browser. The method includes: inputting device fingerprint data to construct a device fingerprint data set; calculating the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, and constructing an entropy value set for each feature; designing a low-frequency feature perturbation formula for the low-frequency feature set to obtain the perturbed low-frequency feature values and the perturbed low-frequency feature set; constructing a generative adversarial network according to the high-frequency feature set and the perturbed low-frequency feature set; collecting the dynamic feedback information of the tracking system in real time, and finally outputting the updated generative adversarial network and the verified obfuscated feature set. The present invention solves the contradiction between privacy protection and function compatibility in the prior art and significantly improves the dynamic adaptation ability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data protection, and particularly relates to a data protection method, device and storage medium for a fingerprint browser. Background Art

[0002] With the rapid development of Internet and big data technologies, user privacy protection is facing unprecedented challenges. As an important means of user identification and tracking, browser fingerprint technology has been widely used in fields such as advertising push, user behavior analysis, and security verification. Browser fingerprints generate a unique identifier by collecting specific parameters of the user's device (such as screen resolution, time zone settings, language preferences, font library, and Canvas drawing characteristics), which is used to mark users in multiple scenarios. The imperceptibility, persistence, and cross-site adaptability of this technology make it one of the difficulties in privacy protection.

[0003] In the prior art, the solutions to browser fingerprint privacy problems are mainly divided into two categories: one is to block the acquisition of fingerprint features, that is, to reduce the generated information of fingerprints by prohibiting access to specific APIs or blocking certain parameters, such as restricting Canvas drawing data or blocking plugin information. Although this method weakens the tracking ability to a certain extent, it will cause partial loss of browser functions, significantly reduce the user experience, and even affect the normal use of applications (such as online payment and identity authentication) in some scenarios. The other is the immobilization processing of fingerprint features, that is, to avoid unique identifiers by setting unified fingerprint parameters. For example, the time zone, resolution, font, etc. of all users are fixed to unified values. However, this method is not only easily recognized as abnormal behavior by tracking systems, but also reduces the functional compatibility of users due to the consistency of features, especially in application scenarios that require personalized settings (such as multi-language support). More importantly, these two types of methods lack dynamics and cannot cope with the continuously evolving technical means of tracking systems.

[0004] Facing the above problems, the prior art cannot achieve an effective balance between privacy protection and functional compatibility, and it is also difficult to achieve dynamic adaptation to tracking systems. This contradiction urgently requires an innovative method that can protect privacy without affecting the user experience and has the ability to dynamically adjust to adapt to changes in tracking technology at any time. Summary of the Invention

[0005] The object of the present invention is to provide a data protection method, device and storage medium for a fingerprint browser, which solves the contradiction between privacy protection and functional compatibility in the prior art and significantly improves the dynamic adaptation ability of the system.

[0006] To achieve the above object, the present invention provides a data protection method for a fingerprint browser, the method comprising:

[0007] S1. Input device fingerprint data to construct a device fingerprint data set. Each feature in the device fingerprint data set represents a characteristic of the user device. The contribution of each feature to user discrimination is calculated by the method of variance normalization to obtain the feature score of each feature. A hierarchical threshold is designed according to the feature score, and the device fingerprint features are divided into two layers, including a high-frequency feature set and a low-frequency feature set;

[0008] S2. Calculate the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, construct an entropy value set of each feature, and calculate the average entropy value of the low-frequency feature set and the high-frequency feature set respectively according to the entropy value of each feature. An optimization objective function is constructed based on the average entropy values of the low-frequency feature set and the high-frequency feature set to balance the privacy protection, functional integrity and computational overhead of feature perturbation, and the expected target entropy value of each feature is obtained;

[0009] S3. Design a low-frequency feature perturbation formula for the low-frequency feature set to obtain the perturbed low-frequency feature values and the perturbed low-frequency feature set. Recalculate the entropy value set of the perturbed low-frequency feature set to evaluate the perturbation effect, and obtain the complexity index;

[0010] S4. Construct a generative adversarial network based on the high-frequency feature set and the perturbed low-frequency feature set to enhance the interference ability against the tracking system while maintaining the integrity of the fingerprint browser function. The generative adversarial network finally outputs a confused feature set;

[0011] Among them, the objective of the generative adversarial network is designed as follows:

[0012] Generator Receives the high-frequency feature set and the perturbed low-frequency feature set and outputs a confused feature set;

[0013] Discriminator Receives the confused feature set and outputs a probability value indicating the possibility that the input feature is the real feature set;

[0014] Output the optimized generator and discriminator ;

[0015] S5. Real-time collect the dynamic feedback information of the tracking system, including the actual recognition rate and the characteristics of the actual recognition sample distribution, calculate the feedback deviation between the target recognition rate set by the system and the actual recognition rate, and dynamically adjust the regularization term weight of the objective function of the optimized generator according to the feedback deviation to optimize the loss function of the generator. At the same time, add the dynamic feedback information to the optimized discriminator of the objective function of the optimized generator to optimize the loss function of the generator. At the same time, add the dynamic feedback information to the optimized discriminator Optimize the discriminator by maintaining the ability to distinguish between the real feature set and the obfuscated feature set 's loss function, and use it to optimize the generator 's loss function and the optimized discriminator 's loss function for the generator and the discriminator for iterative training, and finally output the updated generative adversarial network and the verified obfuscated feature set.

[0016] Furthermore, calculate the contribution of each feature to user distinguishability through variance normalization to obtain the feature score of each feature, specifically:

[0017] ;

[0018] where represents the variance of the values of feature in the device fingerprint dataset, reflecting the distinguishability of this feature. The largest score value indicates the highest importance of this feature in the user identifier. is the importance score of feature i and constructs a score set represents the variance of the values of feature in the device fingerprint dataset; n is the total number of features;

[0019] The high-frequency feature set includes all features with scores ; is the hierarchical threshold;

[0020] The low-frequency feature set includes all features with scores ;

[0021] At the same time, design dynamic hierarchical update:

[0022] For scenarios with high real-time requirements, design a sliding window dynamic update mechanism. Based on the user's latest interaction behavior, re-evaluate the feature importance score. Among them, the updated score calculation formula is:

[0023] ;

[0024] where is the weight coefficient, used to balance the influence of new data and historical data, is the previous round of score result. The updated ultra-high-frequency feature set and low-frequency feature set are appropriately applied to the user's latest device state to improve the dynamic nature of feature stratification.

[0025] Furthermore, calculate the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, expressed as:

[0026] ;

[0027] Among them, represents the information entropy of the feature . The highest value indicates the maximum randomness of the feature and the minimum dependence on user recognition; is the set of possible values of the feature ; is the probability of the value ; is the weight coefficient of the compensation term, which is used to balance the essential information of the entropy value and the fingerprint protection requirement; is the randomness compensation term, which is used to increase the entropy value additionally for some specific features. Its definition is:

[0028] ;

[0029] Among them, is the number of possible values of the feature , indicating the inherent diversity of the feature; is the set of low-frequency features;

[0030] The average entropy values are calculated for the set of low-frequency features and the set of high-frequency features respectively according to the entropy value of each feature , which is expressed as:

[0031] ;

[0032] Among them, is the average entropy value of the low-frequency features, which mainly reflects the randomness of the feature set with a lower protection level; is the average entropy value of the high-frequency features, which mainly reflects the randomness of the feature set with functional priority; is the number of low-frequency features, is the number of high-frequency features; is the set of low-frequency features; is the set of high-frequency features;

[0033] The optimization objective function is expressed as:

[0034] ;

[0035] Among them, is the trade-off coefficient between privacy and function, which determines the emphasis on protection intensity and function priority; is the regularization weight coefficient, which is used to limit the complexity of feature perturbation and avoid unnecessary computational overhead; is the perturbation complexity regularization term, which is used to limit the overhead of low-frequency features in the randomization process. Its definition is:

[0036] ;

[0037] Among them, represents the possible number of values of the feature, indicating the relative complexity required for feature perturbation; is the feature 's variance, representing the scalability of its original distribution. The smaller the variance, the greater the difficulty of perturbation.

[0038] Furthermore, the S3 also includes calculating the entropy difference of each feature based on the expected target entropy value of each feature and the actual entropy value of each feature. Among them, the perturbation intensity is proportional to the entropy difference of each feature, and the feature with the largest randomness gap is preferentially assigned the largest perturbation weight.

[0039] Furthermore, the S3 specifically includes:

[0040] For the low-frequency feature set the low-frequency feature perturbation formula is expressed as:

[0041] ;

[0042] Among them, is the perturbed feature value, is the original feature value; is the random perturbation term, sampled from the normal distribution and dynamically adjusted to control the perturbation amplitude; represents the variance of the feature reflecting its distribution characteristics; is the regularization factor of the patent design, controlling the dynamic adjustment of the perturbation amplitude with the feature distribution to ensure more significant perturbations are applied to features with a more uniform distribution; is the feature with the largest variance in the low-frequency feature set, used for normalization;

[0043] Design the perturbation complexity index for the perturbed feature set and recalculate its entropy value to evaluate the perturbation effect, expressed as:

[0044] ;

[0045] Among them, represents the size of the set of possible values of the feature used to evaluate the feature complexity; is the square of the perturbation amplitude, representing the perturbation energy.

[0046] Furthermore, the target design of the generative adversarial network is specifically:

[0047] The generator Optimization objective Design:

[0048] ;

[0049] The first item Represents the ability of the generator to confuse the discriminator. The generator hopes to confuse the feature set to be misjudged as real features by the discriminator; Represents the probability value of the discriminator judging whether the confused features are real features;

[0050] The second item Is a regularization term used to enhance the randomness of the confused features, Represents the weight coefficient of the regularization term, where, Represents the average entropy value of the confused feature set:

[0051] ;

[0052] Among them, Is the entropy value of each feature;

[0053] Optimization objective of discriminator D Design:

[0054] ;

[0055] Among them, Represents the real feature set, Represents the generated confused feature set.

[0056] Furthermore, the regularization term weight of the objective function of the optimized generator dynamically adjusted according to the feedback deviation optimizes the loss function of the generator Specifically includes:

[0057] ;

[0058] Among them, Is the loss function of the optimized generator ;

[0059] The first item Is the main generative adversarial network loss of the generator, ensuring that the confused features are difficult to be distinguished by the discriminator;

[0060] The second item Is a feedback-driven randomness regularization term, which enhances the entropy value and randomness of the confused features by dynamically adjusting the regularization weight:

[0061] ;

[0062] Among them, is the entropy value of the obfuscation feature , measuring the randomness of the feature; is the distribution variance of the obfuscation feature , ensuring the diversity of the feature; is the diversity adjustment coefficient of the feature distribution;

[0063] At the same time, add dynamic feedback information to the optimized discriminator to maintain the ability to distinguish between the real feature set and the obfuscation feature set and optimize the discriminator 's loss function, expressed as:

[0064] ;

[0065] Among them, is the loss function of the optimized discriminator ;

[0066] The first and second terms are the standard generative adversarial network losses, measuring the probability that real features are judged to be real and the probability that obfuscation features are judged to be generated, respectively;

[0067] The third term is the feedback control regularization term, used to limit the overfitting of the discriminator to the obfuscation features:

[0068] ;

[0069] This term avoids overfitting and simplification by controlling the degree to which the output of the discriminator deviates from the central value .

[0070] Furthermore, the S5 further includes verifying the updated generator and discriminator to see if they meet the following conditions:

[0071] The recognition rate of the obfuscation feature set ;

[0072] The average entropy value of the obfuscation feature set reaches the set range;

[0073] If the set range is not reached or the recognition rate is not reached then recalculate the feedback deviation and train the generative adversarial network.

[0074] In addition, to achieve the above object, the present invention further provides a data protection device for a fingerprint browser, the device comprising: a memory, a processor, and a data protection program for a fingerprint browser stored on the memory and executable on the processor, the data protection for a fingerprint browser configured to implement the steps of a data protection method for a fingerprint browser as described above.

[0075] In addition, to achieve the above object, the present invention further provides a storage medium, on which a data protection program for a fingerprint browser is stored, and when the data protection program for a fingerprint browser is executed by a processor, it implements the steps of a data protection method for a fingerprint browser as described above.

[0076] The beneficial technical effects of the present invention are at least as follows:

[0077] (1) The present invention first introduces a dynamic hierarchical modeling method for fingerprint features. According to the importance and usage frequency of features, fingerprint features are divided into a high-frequency layer and a low-frequency layer. By maintaining dynamic legal output for high-frequency features, the integrity of browser functions is ensured; for low-frequency features, forged features or obfuscation parameters are injected to enhance the privacy protection effect. This dynamic hierarchical mechanism can effectively balance privacy protection and functional compatibility, and solves the compatibility problem caused by fingerprint immobilization in the prior art.

[0078] (2) The present invention dynamically adjusts the obfuscation strategy of different features by calculating the entropy value of fingerprint features. For features with lower information entropy (such as plugin order or font information), the randomness and diversity are preferentially increased to interfere with the analysis of the tracking system; for features with higher information entropy (such as screen resolution or Canvas drawing), a moderate perturbation is maintained to ensure normal functions. The entropy optimization method makes the obfuscated features closer to real data in statistical distribution, thereby effectively improving the protection effect and solving the problem that single obfuscated features in the prior art are easily recognized.

[0079] (3) The present invention designs a dynamically updated generative adversarial network (GAN) for real-time generation of obfuscated fingerprint features. The generator dynamically generates high-entropy forged fingerprints according to user behavior and tracking scenarios, while the discriminator simulates the recognition logic of the tracking system to ensure that the generated features achieve the best balance between privacy protection and authenticity. This network has an adaptive ability and can automatically adjust the obfuscation strategy according to the technological evolution of the tracking system, solving the problem that static protection methods in the prior art are difficult to cope with dynamic threats. Description of the Drawings

[0080] The present invention will be further described with reference to the accompanying drawings. However, the embodiments shown in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on the following drawings without creative efforts.

[0081] Figure 1 It is a flowchart of a data protection method for a fingerprint browser according to the present invention. Specific embodiments

[0082] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0083] As Figure 1 shown, a data protection method for a fingerprint browser provided by an embodiment of the present invention includes the following steps S1 - S5:

[0084] S1. Input device fingerprint data to construct a device fingerprint data set. Each feature in the device fingerprint data set represents a characteristic of the user device. The contribution of each feature to user discrimination is calculated by the method of variance normalization to obtain the feature score of each feature. A hierarchical threshold is designed according to the feature score, and the device fingerprint features are divided into two layers, including a high - frequency feature set and a low - frequency feature set.

[0085] Specifically, this step aims to collect device fingerprint features and provide support for subsequent steps through dynamic hierarchical modeling. The differences in data ranges are eliminated through feature standardization, the sensitivity of features is quantified by combining with a feature importance scoring model, and then the features are stratified according to the scoring results. The design of dynamic stratification not only ensures the identification of key protected features but also provides clear inputs for subsequent entropy optimization and obfuscation strategies.

[0086] The input is a device fingerprint data set , where each feature represents a characteristic of the user device, such as screen resolution, time zone, language setting, font library, etc. First, each feature is standardized to eliminate the influence of different feature value ranges. The standardization formula is:

[0087] ;

[0088] where is the standardized feature, and are the mean and standard deviation of feature respectively, both calculated from the training data set. The standardized result It is the basic data for subsequent feature scoring.

[0089] Furthermore, a feature scoring model is used to calculate the relative importance of each normalized feature, and the contribution of each feature to user discrimination is quantified through the method of variance normalization:

[0090] ;

[0091] Among them, represents the variance of the values of feature in the device fingerprint dataset, reflecting the discriminability of this feature. The larger the scoring value, the higher the importance of this feature in user identification. The calculated scoring set is the core basis for dividing high-frequency and low-frequency features.

[0092] Furthermore, according to the importance score a hierarchical threshold is set to divide the device fingerprint features into two layers:

[0093] The high-frequency feature set includes all features with scores , and these features are usually the core features related to user functions, such as resolution, language options, etc.

[0094] The low-frequency feature set includes all features with scores , and these features contribute relatively little to user identification and can be perturbed through subsequent obfuscation operations.

[0095] After stratification, the high-frequency feature set and the low-frequency feature set are output for use in the next entropy value calculation.

[0096] ;

[0097] Among them, is the weight coefficient used to balance the influence of new data and historical data, is the scoring result of the previous round. The updated feature stratification results and adapt to the latest device status of users and improve the dynamics of feature stratification.

[0098] S2. Calculate the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, construct an entropy value set for each feature, calculate the average entropy value of the low-frequency feature set and the high-frequency feature set respectively according to the entropy value of each feature, and construct an optimization objective function based on the average entropy values of the low-frequency feature set and the high-frequency feature set to balance privacy protection, functional integrity, and the computational overhead of feature perturbation, so as to obtain the expected target entropy value of each feature.

[0099] Further, for the input and , calculate the entropy value of each feature respectively. The definition of the entropy value is:

[0100] ;

[0101] where, represents the information entropy of feature . The higher it is, the greater the randomness of the feature and the smaller the dependence on user identification. is the set of possible values of feature , is the probability of value . is the weight coefficient of the compensation term, which is used to balance the essential information of the entropy value and the fingerprint protection requirements. is an additional designed randomness compensation term, which is used to increase the entropy value for some specific features (such as fonts, plugin lists) additionally. These features are more likely to be exploited in the tracking system and need to further enhance their randomness. Its definition is:

[0102] ;

[0103] where, is the number of possible values of feature , indicating the inherent diversity of the feature.

[0104] By introducing the compensation term, the entropy value calculation is more adapted to the scenario requirements of fingerprint protection. Especially for low-frequency features in the patent scenario, higher randomness is required to interfere with tracking.

[0105] Further, after calculating the entropy value of each feature, calculate the average entropy value of the low-frequency feature set and the high-frequency feature set respectively. The formula is:

[0106] ;

[0107] where, is the average entropy value of the low-frequency features, which mainly reflects the randomness of the feature set with a lower protection level; is the average entropy value of the high-frequency features, which mainly reflects the randomness of the feature set with functional priority; is the number of low-frequency features, is the number of high-frequency features; a higher corresponds to a higher privacy protection effect, while should be maintained at a certain level to ensure functional integrity.

[0108] Furthermore, to balance privacy protection and functional requirements, an optimization objective function is constructed, taking into account global randomness, protection requirements, and tracking interference capabilities. The optimization objective function is:

[0109] ;

[0110] where, is the trade-off coefficient between privacy and function, which determines the emphasis on protection intensity and function priority. is the regularization weight coefficient, which is used to limit the complexity of feature perturbation and avoid unnecessary computational overhead. is the perturbation complexity regularization term for innovative design, which is used to limit the overhead of low-frequency features in the randomization process, and its definition is:

[0111] ;

[0112] where, represents the number of possible values of the feature, represents the relative complexity required for feature perturbation; is the feature 's variance, which represents the scalability of its original distribution. The smaller the variance, the greater the difficulty of perturbation. By introducing , the optimization objective function can more effectively balance privacy protection, functional integrity, and the computational overhead of feature perturbation.

[0113] It can be understood that the following content is finally output:

[0114] The entropy value set of each feature , which clarifies the randomness level of each feature.

[0115] The average entropy values of low-frequency and high-frequency features and , which provide a global feature randomness index for subsequent protection strategies.

[0116] The optimization objective function , which serves as an important reference basis for subsequent feature perturbation and obfuscation strategies.

[0117] S3. Design a low-frequency feature perturbation formula for the low-frequency feature set, obtain the perturbed low-frequency feature values and the perturbed low-frequency feature set, recalculate the entropy value set of the perturbed low-frequency feature set to evaluate the perturbation effect, and obtain the complexity index.

[0118] Specifically, the optimized objective function calculated according to step S2 and the entropy difference of each feature , for the low-frequency feature set assign perturbation intensities. Define the entropy difference of each feature :

[0119] ;

[0120] Among them, is the expected target entropy value, usually determined by the optimized objective function , for example, selecting the target average entropy value of low-frequency features as the benchmark. is the current entropy value calculated in step S2.

[0121] Among them, the perturbation intensity is proportional to , and features with a larger randomness gap are preferentially assigned a larger perturbation weight.

[0122] Furthermore, for the low-frequency feature set , adopt the following innovative perturbation formula:

[0123] ;

[0124] Among them, is the perturbed feature value, is the original feature value. is the random perturbation term, sampled from the normal distribution , dynamically adjusted to control the perturbation amplitude. represents the variance of feature , reflecting its distribution characteristics. is the regularization factor of the patent design, controlling the dynamic adjustment of the perturbation amplitude with the feature distribution, ensuring more significant perturbations are applied to features with a more uniform distribution. is the feature with the largest variance in the low-frequency feature set, used for normalization. This formula effectively improves the pertinence and effect of the perturbation by introducing this special term and adaptively adjusting the perturbation amplitude according to the distribution characteristics of the features.

[0125] Furthermore, for the perturbed feature set , recalculate its entropy value to evaluate the perturbation effect. To reduce formula repetition, the entropy value calculation formula here follows the definition in step S2, but the distribution of the perturbed features is re-counted.

[0126] In addition, for the control requirement of complexity in the patent scenario, design a perturbation complexity index:

[0127] ;

[0128] Among them, represents the size of the set of possible values of the feature , and is used to evaluate the feature complexity. is the square of the perturbation amplitude, representing the perturbation energy. The complexity index can be used in subsequent steps to limit the negative impact of excessive perturbation.

[0129] It can be understood that the set of low-frequency features after output perturbation is provided as input for subsequent generation of obfuscated features. The set of updated entropy values is output for evaluating the optimized result. The complexity index is output to provide a basis for controlling the consumption of perturbation resources.

[0130] S4. Construct a generative adversarial network based on the set of high-frequency features and the set of low-frequency features after perturbation, enhance the interference ability against the tracking system, and at the same time maintain the integrity of the fingerprint browser function. The generative adversarial network finally outputs a set of obfuscated features.

[0131] Specifically, construct a generative adversarial network (GAN), where the generator and the discriminator are designed as follows in terms of input and target:

[0132] Generator : Receive and , and output a set of obfuscated features . The goal of the generator is to generate features with high randomness and difficulty in distinction, while retaining key functional features.

[0133] Discriminator : Receive the set of features and output a probability value, representing the possibility that the input features are the set of real features (such as ). The goal of the discriminator is to maximize the ability to distinguish between real features and generated features.

[0134] Furthermore, to balance the authenticity and randomness of the generated features, an innovative optimization objective function is introduced:

[0135] ;

[0136] The first term represents the ability of the generator to confuse the discriminator. The generator hopes that the set of obfuscated features is misjudged by the discriminator as real features.

[0137] The second term It is a regular term designed specifically for patents to enhance the randomness of obfuscation features. Among them, represents the average entropy value of the obfuscation feature set:

[0138] ;

[0139] This item guides the generator to output high-entropy features, making the generated features more difficult to be utilized by the tracking system.

[0140] Furthermore, the goal of the discriminator is to distinguish real features from the generated obfuscation features, and its optimization goal is:

[0141] ;

[0142] Among them, represents the real feature set (such as and part of ). represents the generated obfuscation feature set.

[0143] Furthermore, to adapt to the dynamic changes of the tracking system, a dynamic feedback mechanism is designed:

[0144] Collect the behavior feedback of the tracking system, such as the change in the recognition rate of .

[0145] Dynamically adjust the weights in the optimization goal of the generator using the feedback information to adapt to the evolution of the tracking algorithm and improve the robustness of the obfuscation features.

[0146] It can be understood that the output obfuscation feature set provides input for the subsequent dynamic feedback link and privacy protection deployment. The output optimized generator and discriminator , as the core modules in the patent protection system, support long-term operation.

[0147] S5. Real-time collect the dynamic feedback information of the tracking system, including the actual recognition rate and the characteristics of the actual recognition sample distribution, calculate the feedback deviation between the target recognition rate set by the system and the actual recognition rate, and dynamically adjust the regular term weight of the objective function of the optimized generator to optimize the loss function of the generator , and at the same time add the dynamic feedback information to the optimized discriminator to maintain the ability to distinguish between the real feature set and the obfuscation feature set and optimize the loss function of the discriminator , and use the loss function of the optimized generator and the loss function of the optimized discriminator to optimize the loss function of the generator and discriminator Iterative training is performed, and finally an updated generative adversarial network and a verified set of obfuscation features are output.

[0148] Specifically, dynamic feedback information of the tracking system is collected, including the recognition rate (the probability that the tracking system correctly identifies the obfuscation features ) and the characteristics of the distribution of the recognition samples .

[0149] Among them, the feedback deviation is defined :

[0150] ;

[0151] Among them, is the target recognition rate set by the system (for example ), indicating the upper limit of the recognition of the expected set of obfuscation features. is the optimization demand of the current state of the model. The greater the deviation, the more insufficient the current obfuscation effect, and it is necessary to increase randomness and complexity to interfere with the tracking system.

[0152] Furthermore, according to the feedback deviation the regularization term weight of the generator objective function is dynamically adjusted, and the optimized generator objective function is:

[0153] ;

[0154] The first term is the main GAN loss of the generator, ensuring that the obfuscation features are difficult to be distinguished by the discriminator.

[0155] The second term is a feedback-driven randomness regularization term, which improves the entropy value and randomness of the obfuscation features through the dynamically adjusted regularization weight:

[0156] ;

[0157] Among them, is the entropy value of the obfuscation feature , measuring the randomness of the feature. is the distribution variance of the obfuscation feature , ensuring the diversity of the features. is the feature distribution diversity adjustment coefficient. This design combines the entropy value and the variance, and through dynamic adjustment, enables the generator to adapt to the evolution of the tracking system in real time.

[0158] Furthermore, the discriminator target maintains the ability to distinguish between the real feature set and the obfuscation feature set, and at the same time feedback information is added to optimize the loss function:

[0159] ;

[0160] The first and second terms are standard GAN losses, which measure the probability that real features are judged to be real and the probability that confused features are judged to be generated, respectively.

[0161] The third term is a feedback control regularization term used to limit the overfitting of the discriminator to confused features:

[0162] ;

[0163] This term avoids overfitting and simplification by controlling the degree to which the output of the discriminator deviates from the central value .

[0164] It can be understood that using the updated and , the generator and the discriminator are iteratively trained:

[0165] In each round of training, the latest feedback data and are sampled, and the weights and are dynamically updated.

[0166] The outputs of the optimized generator and discriminator are reused to verify the feedback deviation , ensuring that the update effect gradually approaches the target state.

[0167] Finally, verify whether the updated generator and discriminator meet the following conditions:

[0168] The recognition rate of the set of confused features .

[0169] The average entropy value of the set of confused features reaches the set range.

[0170] Finally, output the optimized generator and discriminator , as well as the verified set of confused features for subsequent privacy protection deployment.

[0171] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.

[0172] In addition, for the technical details not described in detail in this embodiment, reference may be made to the parameter running method provided in any embodiment of the present invention, which will not be elaborated here.

[0173] Other embodiments or specific implementation manners of the data protection device for a fingerprint browser according to the present invention may refer to the above method embodiments, which will not be elaborated here.

[0174] In addition, an embodiment of the present invention further provides a storage medium, on which a data protection program for a fingerprint browser is stored. When the data protection program for a fingerprint browser is executed by a processor, the steps of the data protection method for a fingerprint browser as described above are implemented.

[0175] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0176] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0178] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A data protection method for a fingerprint browser, characterized in that The method includes: S1. Input device fingerprint data to construct a device fingerprint data set. Each feature in the device fingerprint data set represents a characteristic of the user device. The contribution of each feature to user discrimination is calculated by the method of variance normalization to obtain the feature score of each feature. A hierarchical threshold is designed according to the feature scores, and the device fingerprint features are divided into two layers, including a high-frequency feature set and a low-frequency feature set; S2. Calculate the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, and construct an entropy value set of each feature. The average entropy values of the low-frequency feature set and the high-frequency feature set are calculated according to the entropy value of each feature respectively. An optimization objective function is constructed according to the average entropy values of the low-frequency feature set and the high-frequency feature set to balance privacy protection, functional integrity and the computational overhead of feature perturbation, and the expected target entropy value of each feature is obtained; S3. Design a low-frequency feature perturbation formula for the low-frequency feature set to obtain the perturbed low-frequency feature values and the perturbed low-frequency feature set. Re-calculate the entropy value set of the perturbed low-frequency feature set to evaluate the perturbation effect, and obtain a complexity index; S4. Construct a generative adversarial network according to the high-frequency feature set and the perturbed low-frequency feature set to enhance the interference ability against the tracking system while maintaining the integrity of the fingerprint browser function. The generative adversarial network finally outputs a confused feature set; Among them, the objective of the generative adversarial network is designed as follows: Generator Receives a high-frequency feature set and a perturbed low-frequency feature set and outputs a confused feature set; Discriminator Receives the set of obfuscated features and outputs a probability value indicating the likelihood that the input features are the set of true features; The optimized generator for output and the discriminator ; S5. Dynamically collect and track the feedback information of the system, including the actual recognition rate and the characteristics of the actual recognition sample distribution, calculate the feedback deviation between the target recognition rate set by the system and the actual recognition rate, and dynamically adjust the optimized generator according to the feedback deviation to optimize the regular term weight of the objective function of the generator and the loss function, and at the same time add the dynamic feedback information to the optimized discriminator to optimize the discriminator while maintaining the ability to distinguish between the real feature set and the confused feature set and its loss function, and use the optimized generator and its loss function and the optimized discriminator and its loss function to perform iterative training on the generator and the discriminator Finally, output the updated generative adversarial network and the verified confused feature set.

2. The data protection method for a fingerprint browser according to claim 1, characterized in that, The contribution of each feature to user discrimination is calculated by the method of variance normalization to obtain the feature score of each feature. Specifically: ; Among them, represents the variance of the value of the feature in the device fingerprint dataset, reflecting the distinguishability of the feature. The largest scoring value indicates the highest importance of the feature in the user identification. is the importance score of the feature and constructs a scoring set , represents the variance of the value of the feature in the device fingerprint dataset; n is the total number of features; The high-frequency feature set includes all scores of the features, which are hierarchical thresholds; The low-frequency feature set includes all scores of the features; At the same time, a dynamic hierarchical update is designed: For scenarios with high real-time requirements, a sliding window dynamic update mechanism is designed. Based on the user's latest interaction behavior, the feature importance score is re-evaluated. The updated score calculation formula is: ; Among them, is the weight coefficient, which is used to balance the influence of new data and historical data. is the previous round of scoring results; the updated high-frequency feature set and low-frequency feature set are adapted to the user's latest device status to improve the dynamic nature of feature stratification.

3. A data protection method for a fingerprint browser according to claim 2, characterized in that, Calculate the entropy value of each feature according to the high-frequency feature set and the low-frequency feature set respectively, which is expressed as: ; Among them, represents the information entropy of the feature . The highest value indicates the greatest randomness of the feature and the least dependence on user recognition. is the set of possible values of the feature , is the probability of the value . is the weight coefficient of the compensation term, which is used to balance the essential information of the entropy value and the fingerprint protection requirements. is the randomness compensation term, which is used to additionally increase the entropy value for some specific features, and its definition is: ; Among them, is the number of possible values of the feature , representing the inherent diversity of the feature; is the set of low-frequency features; The entropy value according to each feature Calculate the average entropy values for the low-frequency feature set and the high-frequency feature set respectively, which are expressed as: ; Among them, is the average entropy value of low-frequency features, reflecting the randomness of the feature set with a lower protection level; is the average entropy value of high-frequency features, reflecting the randomness of the functional priority feature set; is the number of low-frequency features, is the number of high-frequency features; is the low-frequency feature set; is the high-frequency feature set; The optimized objective function , is expressed as: ; Among them, is the trade-off coefficient between privacy and function, which determines the emphasis on protection intensity and function priority; is the regularization weight coefficient, which is used to limit the complexity of feature perturbation and avoid unnecessary computational overhead; is the perturbation complexity regular term, which is used to limit the overhead of low-frequency features in the randomization process, and its definition is: ; Among them, represents the possible number of values of the feature, represents the relative complexity required for feature perturbation; is the variance of the feature which represents the extensibility of its original distribution. The smaller the variance, the greater the difficulty of perturbation.

4. A data protection method for a fingerprint browser according to claim 1, characterized in that In step S3, it also includes calculating the entropy difference of each feature according to the expected target entropy value of each feature and the actual entropy value of each feature. The perturbation intensity is proportional to the entropy difference of each feature, and the feature with the largest randomness gap is preferentially assigned the largest perturbation weight.

5. A data protection method for a fingerprint browser according to claim 4, characterized in that, Step S3 specifically includes: For the low-frequency feature set The low-frequency feature perturbation formula is expressed as: ; Among them, is the perturbed eigenvalue, is the original eigenvalue; is the random perturbation term, sampled from the normal distribution and dynamically adjusted to control the perturbation amplitude; represents the variance of the feature reflecting its distribution characteristics; is the regularization factor, controlling the dynamic adjustment of the perturbation amplitude with respect to the feature distribution, ensuring more significant perturbations are imposed on features with a more uniform distribution; is the feature with the largest variance in the low-frequency feature set, used for normalization;​ For the perturbed feature set Design the perturbation complexity index , and recalculate its entropy value to evaluate the perturbation effect, expressed as: ; Among them, represents the size of the set of possible values of the feature and is used to evaluate the feature complexity; is the square of the perturbation amplitude and represents the perturbation energy.

6. A data protection method for a fingerprint browser according to claim 1, characterized in that The objective design of the generative adversarial network is specifically: Generator Optimization objective Design: ; The first item Indicates the ability of the generator to confuse the discriminator. The generator hopes to confuse the feature set and be misjudged as real features by the discriminator; Indicates the probability value of the discriminator judging whether the confused features are real features; The second item is a regularization term used to enhance the randomness of the obfuscation features, represents the weight coefficient of the regularization term, where represents the average entropy value of the obfuscation feature set: ; Among them, is the entropy value of each feature; Optimization objective of discriminator D Design: ; Among them, represents the set of real features, represents the set of generated obfuscated features.

7. A data protection method for a fingerprint browser according to claim 6, characterized in that The generator optimized by dynamically adjusting according to the feedback deviation The generator for optimizing the regularization term weight of the objective function The loss function, specifically including: ; Among them, is the loss function of the optimized generator ; is the optimization demand for the current state of the model The first item is the generator adversarial network loss of the generator, ensuring that the confused features are difficult to be distinguished by the discriminator; The second item is a feedback-driven randomness regularization term that enhances the entropy value and randomness of obfuscated features by dynamically adjusting the regularization weight: ; Among them, is the entropy value of the obfuscation feature to measure the randomness of the feature; is the distribution variance of the obfuscation feature to ensure the diversity of the feature; is the diversity adjustment coefficient of the feature distribution; Meanwhile, add the dynamic feedback information to the optimized discriminator Optimize the discriminator while maintaining the ability to distinguish between the real feature set and the confused feature set The loss function of, expressed as: ; Among them, is the loss function of the optimized discriminator ; The first item and the second item are the losses of the standard generative adversarial network, which respectively measure the probability that the real feature is judged to be real and the probability that the confused feature is judged to be generated; The third item is a feedback control regularization term used to limit the overfitting of the discriminator to the confusing features: ; By controlling the degree to which the output of the discriminator deviates from the central value overfitting and oversimplification are avoided.

8. A data protection method for a fingerprint browser according to claim 7, characterized in that, The S5 further includes verifying the updated generator and discriminator to determine whether the following conditions are met: Average entropy value of the confusion feature set Reach the set range; If the set range or the recognition rate is not reached, the feedback deviation is recalculated to train the generative adversarial network.

9. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for realizing security privacy calculation based on browser client

    CN112765578A

  • Adversarial simulation attack method and device based on attention frequency domain GAN

    CN118429689A