User behavior data protection method and device, equipment and storage medium

By using feature encoding of user behavior data and a strategy decision model based on deep reinforcement learning algorithms, a data protection strategy is generated, which solves the problem that static rule-based protection mechanisms cannot balance privacy protection and user experience, and achieves a balance between privacy protection and user experience.

CN121935952APending Publication Date: 2026-04-28HANGZHOU XINZHONGDA ENTERPRISE MANAGEMENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU XINZHONGDA ENTERPRISE MANAGEMENT TECHNOLOGY CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing static rule-based protection mechanisms struggle to balance the strength of privacy protection with user experience, resulting in user behavior data being unable to be linked to specific individuals during the privacy protection process, thus affecting the effectiveness of recommendation systems and personalized settings.

Method used

By encoding user behavior data into features to generate behavior feature vectors, and combining historical data of group users with model operating environment data, deep reinforcement learning algorithms are used to determine data leakage risks, and data protection strategies are generated based on policy decision models to obfuscate user behavior data.

Benefits of technology

It achieves a balance between protecting user privacy and maintaining user experience, ensuring the effectiveness of the recommendation system and personalized settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935952A_ABST
    Figure CN121935952A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data protection, in particular to a user behavior data protection method and device, equipment and a storage medium, and the method comprises the steps: carrying out the feature coding of user behavior data, and obtaining a behavior feature vector; determining a data leakage risk at least based on the behavior feature vector and historical data of group users; determining a data protection strategy based on the behavior feature vector, model operation environment data, the data leakage risk and a strategy decision model preset based on a deep reinforcement learning algorithm; and performing confusion processing on the user behavior data based on the data protection strategy to obtain behavior confusion data. The privacy protection intensity and the user experience can be balanced conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data protection technology, and in particular to a method, apparatus, device and storage medium for protecting user behavior data. Background Technology

[0002] In today's internet environment, user behavior data is frequently collected and analyzed in large quantities, leading to increasingly serious risks of privacy breaches. Common types of user behavior data include shopping data, search data, and navigation data.

[0003] Currently, static rule-based protection mechanisms are generally used to protect user behavior data. The specific steps are as follows: 1. User behavior collection, such as user operation behavior, location information, search keywords, historical records, etc., to generate a behavior log; 2. Privacy risk identification: such as the system identifying potential privacy risk events based on manually set static rules (keyword matching, sensitive API calls) or thresholds (access frequency restrictions); 3. Policy execution: if the rules in the previous step are triggered, the corresponding fixed operations are executed, such as: 1) access request, 2) de-identification or masking of fields, 3) pop-up prompts or user confirmation, 4) adding fixed pseudo data to the log, thereby achieving the protection of user behavior data.

[0004] However, protecting user behavior data through the aforementioned static rule-based protection mechanism makes it difficult to balance the strength of privacy protection with user experience. For example, if personal identifiers (such as name, ID, IP address) are removed from user behavior data, the data cannot be associated with specific individuals, thus causing recommendation systems, content filtering, and personalized settings to all fail, and users will see a "generic" interface that is irrelevant to them. Summary of the Invention

[0005] To facilitate a balance between the strength of privacy protection and user experience, this application provides a method, apparatus, device, and storage medium for protecting user behavior data.

[0006] Firstly, this application provides a method for protecting user behavior data, including:

[0007] User behavior data is feature-encoded to obtain a behavior feature vector;

[0008] The risk of data leakage should be determined based at least on the aforementioned behavioral feature vectors and historical data of the user group.

[0009] Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model pre-set by a deep reinforcement learning algorithm, a data protection strategy is determined.

[0010] The user behavior data is obfuscated based on the data protection strategy to obtain obfuscated behavior data.

[0011] Secondly, this application provides a user behavior data protection device, comprising:

[0012] The vector calculation module is used to encode user behavior data to obtain behavior feature vectors.

[0013] The risk calculation module is used to determine the risk of data leakage based at least on the behavioral feature vector and the historical data of the group of users;

[0014] The strategy generation module is used to determine a data protection strategy based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm.

[0015] The data obfuscation module is used to obfuscate the user behavior data based on the data protection policy to obtain obfuscated behavior data.

[0016] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the method described above.

[0017] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0018] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0019] The aforementioned user behavior data protection method, apparatus, device, and storage medium obtain a behavior feature vector by feature encoding user behavior data; determine the data leakage risk based at least on the behavior feature vector and historical data of the user group; determine a data protection strategy based on the behavior feature vector, model operating environment data, the data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm; and obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data. Through the above implementation, the data leakage risk of user behavior data is first calculated, and then the behavior feature vector corresponding to the user behavior data, model operating environment data, and the data leakage risk are processed by the strategy decision model to generate a data protection strategy for obfuscating user behavior data (data leakage prevention processing). Since this data protection strategy fully considers the data leakage risk of user behavior data, it can effectively achieve a balance between privacy protection strength and user experience.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a user behavior data protection method provided in the embodiments of this application;

[0023] Figure 2 This is a schematic diagram of the structure of a user behavior data protection device provided in the embodiments of this application;

[0024] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0025] Figure 4 This is an internal structural diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure.

[0027] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0028] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a user behavior data protection method provided in Embodiment 1 of this application, with reference to... Figure 1 The method can be executed by a device that performs the method, which can be implemented in software and / or hardware, and the method includes:

[0031] S110. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0032] It should be noted that the abundance and widespread availability of smart devices and functional or entertainment software have greatly enriched people's lives and effectively improved their convenience. Smart devices include, but are not limited to, smartphones, tablets, laptops, smartwatches, and in-vehicle infotainment systems. Functional software includes, but is not limited to, search software, navigation software, and shopping software. Entertainment software includes, but is not limited to, short video software, music software, and game software. Users of smart devices frequently interact with the internet through these software programs. When using software on smart devices, users inevitably need to perform operations on the software, generating a wealth of behavioral data, which is recorded as user behavior data. For example, user behavior data includes, but is not limited to: click operations, swipe operations, search terms, preferences, travel routes, page dwell time, keyword text, session ID, sequence number, operation frequency, and prompts input by the user into the large model.

[0033] During software use, users generate the aforementioned user behavior data. It should be noted that this data is easily and silently acquired by some software or programs through improper means, not only compromising user data security but also potentially leading to subsequent informational disturbances. For example, some software or programs may silently acquire a user's click frequency on a particular product category and / or dwell time on a specific page within a shopping app. Subsequently, the software or program may analyze user preferences based on these click frequencies and / or dwell times, and then use methods such as phone calls, text messages, or advertising pages to promote products to the user, thus disturbing them. As another example, a user's travel routes while using navigation software may also be silently acquired by some software or programs through improper means, resulting in long-term tracking of their commute locations and exposure of their residence, workplace, and points of interest, making them vulnerable to exploitation by malicious trackers or advertisers.

[0034] To protect user data security and reduce the degree of information interference to users, this example aims to protect user behavior data. The risk of data leakage for user behavior data varies; some user behavior data with a high risk of leakage requires effective protection after its generation, while some user behavior data with a low risk of leakage does not require very effective protection after its generation. Therefore, it is first necessary to determine the data leakage risk of user behavior data. To facilitate the quantification of the data leakage risk of user behavior data, this example first analyzes the acquired user behavior data... Perform feature encoding to convert user behavior data Encode it as a vector, and denote this vector as the behavior feature vector. .

[0035] S120. Based at least on the behavioral feature vector and the historical data of the group of users, determine the risk of data leakage.

[0036] It should be noted that, in order to calculate the data leakage risk of user behavior data, it is necessary to obtain at least a preset group of users' historical data. Historical data of this user group This includes user behavior data stored in the history of multiple users.

[0037] In an optional embodiment, the behavioral feature vector can be... Historical data of group users Information entropy calculation is performed to determine the corresponding mutual information, which is then used as a risk indicator for the leakage of user behavior data.

[0038] In an optional embodiment, after obtaining the behavioral feature vector Historical data of group users Subsequently, data on the operating environment of the model corresponding to the smart device was further obtained. For example, the model operates on environmental data. Including but not limited to: device type, operating system, network type (WiFi / 4G), location information (latitude and longitude or grid ID), IP subnet, connection quality, current time, geographic semantics (home / office / public place), etc.; furthermore, using appropriate leakage risk calculation models to analyze behavioral feature vectors. Historical data of group users and model runtime environment data The data is processed to calculate the data leakage risk of user behavior data; for example, the leakage risk calculation model in this embodiment can be either an attacker model or a differential privacy model, while in other embodiments, the specific model is not limited.

[0039] S130. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm, determine a data protection strategy.

[0040] It should be noted that data breach risk is used to quantify and measure the degree of risk of user behavior data being silently acquired and used by other software or programs. After determining the data breach risk of user behavior data, it is necessary to further protect the user behavior data according to the data breach risk to prevent it from being silently acquired and used by other software or programs.

[0041] It should also be noted that different data protection strategies are required for different types of user behavior data and user behavior data with different data leakage risks. Therefore, after obtaining the user behavior data and the corresponding data leakage risks, it is necessary to further generate the corresponding data protection strategies.

[0042] To generate the data protection strategy, this embodiment pre-sets a trained strategy decision model. For example, in this embodiment, the strategy decision model is a strategy agent trained using a deep reinforcement learning algorithm (such as Proximal Policy Optimization, PPO), and this strategy agent is modeled as a Markov Decision Process (MDP). This strategy decision model is used to process behavioral feature vectors. Model runtime environment data And address the risk of data breaches to generate appropriate data protection strategies.

[0043] It should be noted that the model runtime environment data Data cannot be directly input into the strategy decision-making model for processing. It requires feature encoding, similar to user behavior data, to obtain corresponding vectors, and the model's runtime environment data must also be included. After feature encoding, a vector is obtained, denoted as the context feature vector. ;

[0044] Specifically, behavioral feature vectors Context feature vector The risk of data breaches is input into the aforementioned strategy decision model for processing, and the strategy decision model outputs corresponding data protection strategies. For example, the corresponding data protection strategies include, but are not limited to: inserting fake behavior, delaying operations, replacing behavior, and non-intervention, etc., without specific limitations.

[0045] In this process, the insertion of false behavioral representations involves inserting other pre-defined false behavioral data into the user behavior data. The generated false behavioral data must conform to the user's historical behavior distribution to obfuscate the real user behavior data. For example, to prevent large model service providers from profiling users based on the user's input prompt, noise unrelated to the true intent can be added to the prompt, such as splicing adversarial interference text or false context before and after the user's input prompt. Specifically, when the user behavior data is the prompt input by the user to the large model, and the user's input prompt is specifically "cold medicine", the system automatically inserts a false query fragment about "car repair" to disrupt the large model's profile of the user's interests, but does not affect the answer to the real question (because the context can be isolated through Prompt Engineering).

[0046] In this context, delayed operation means that user data is not processed, analyzed, or transmitted immediately upon collection. Instead, it is temporarily stored locally (on the user's device) and the corresponding operation is performed after a preset period of time or when specific conditions are met. For example, assuming that the acquired user behavior data is the user's exercise record, which includes exercise trajectory, steps, heart rate, etc., one way to perform delayed operation on the exercise record is to upload the exercise record stored on the local device (such as a smartwatch or smartphone) to the cloud server in response to the exercise record being stored on the local device for 24 hours or the number of exercise records reaching a corresponding threshold.

[0047] Among them, the replacement behavior representation replaces user behavior data with preset behavior data to mask the real user behavior data; the non-intervention representation does not process the user behavior data when the risk of leakage is very small; for example, when the user behavior data is a prompt input by the user to the large model, before sending the prompt to the large model, in response to the recognition that the prompt contains sensitive entities (such as names, card numbers, etc.), the sensitive entities in the prompt are replaced with preset data using placeholder replacement technology, such as replacing Zhang San with [USER_NAME]; it should be noted that after the large model returns the result, the sensitive entities are restored locally.

[0048] In this context, "no intervention" means that when the risk of leakage of user behavior data is too low, no other protective measures are taken for the user behavior data. For example, if the user behavior data is that a user stays on a technology article page for 30 seconds, since this user behavior data cannot reveal sensitive privacy information such as the user's health, finances, or identity, no other protective measures are needed for this user behavior data. In other words, the data protection strategy for this user behavior data is "no intervention."

[0049] S140. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0050] It should be noted that the above data protection strategies serve two purposes: first, to fully protect user behavior data and prevent it from being stolen and misused; and second, to balance user experience with software usage while protecting user behavior data.

[0051] For example, if the data breach risk calculated from user behavior data is high, the data protection strategy will be stronger, such as replacing the user behavior data (the data protection strategy of replacing behavior). If the data breach risk calculated from user behavior data is moderate, the data protection strategy will also be moderate, such as inserting fake behavior data into the user behavior data (the data protection strategy of inserting fake behavior). In this way, inserting fake behavior data can protect the user behavior data, while the retained user behavior data can be used by the software currently used by the user to provide accurate information push to the user. For example, the accurate information pushed may include more time-saving navigation routes, products that the user may like, etc., without being limited to specifics. This can achieve a balance between the strength of user behavior data protection and the user's software experience.

[0052] After obtaining the data protection strategy through the above steps, the user behavior data can be obfuscated using the data protection strategy. The obfuscation process corresponds to at least one specific processing method of the data protection strategy, such as inserting false behavior, delaying operation, replacing behavior, or not intervening. The new data formed after the user behavior data has been obfuscated is called obfuscated behavior data.

[0053] It should be noted that this embodiment obtains a behavioral feature vector by feature encoding user behavior data; determines the data leakage risk based at least on the behavioral feature vector and historical data of the group of users; determines a data protection strategy based on the behavioral feature vector, model runtime environment data, the data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm; and obfuscates the user behavior data based on the data protection strategy to obtain obfuscated behavioral data. Through the above implementation, the data leakage risk of user behavior data is first calculated, and then the behavioral feature vector corresponding to the user behavior data, model runtime environment data, and the data leakage risk are processed by the strategy decision model to generate a data protection strategy for obfuscating user behavior data (data leakage prevention processing). Since this data protection strategy fully considers the data leakage risk of user behavior data, it can effectively achieve a balance between privacy protection strength and user experience.

[0054] Example 2

[0055] This application provides a user behavior data protection method in Embodiment 2, which optimizes the "character encoding of user behavior data to obtain behavior feature vectors" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0056] S211. Encode the discrete behavior data, text behavior data, numerical behavior data and sequence behavior data in the user behavior data respectively to obtain discrete feature vector, text feature vector, numerical feature vector and sequence feature vector.

[0057] It should be noted that user behavior data may contain various types of data. For example, in this embodiment, the types of data included in user behavior data include: discrete behavior data, text behavior data, numerical behavior data, and sequence behavior data. Discrete behavior data may include event type, page ID, app category, etc.; text behavior data may include search terms, titles, comments, etc.; numerical behavior data may include dwell time, click coordinates, operation frequency, etc.; and sequence behavior data may include location sequence, etc. In other embodiments, the specific types of user behavior data are not limited.

[0058] In order to facilitate the subsequent calculation of the data leakage risk corresponding to user behavior data, it is necessary to first quantify and encode the user behavior data.

[0059] Specifically, for discrete behavior data in user behavior data, this embodiment pre-sets a One-Hot encoding method or a Learned Embedding method. The One-Hot encoding method or Learned Embedding method is used to encode the discrete behavior data into a corresponding vector, and the vector is denoted as a discrete feature vector.

[0060] For text behavior data in user behavior data, this embodiment pre-sets a language model Chinese-BERT. The Chinese-BERT language model is used to encode the text behavior data into a corresponding vector, and the vector is recorded as the text feature vector. It should be noted that if the text behavior data is short text, then CLS vector or average pooling is used, with a dimension of 128-768.

[0061] For numerical behavior data in user behavior data, the numerical behavior data is directly normalized and standardized to obtain the corresponding vector, and this vector is denoted as the numerical feature vector.

[0062] For sequential behavior data in user behavior data, this embodiment pre-sets an LSTM model / Transformer model, which is used to encode sequential behavior data into a temporal representation, and the temporal representation is denoted as a sequence feature vector.

[0063] S212. Perform feature fusion on the discrete feature vector, the text feature vector, the numerical feature vector, and the sequence feature vector to obtain the behavioral feature vector.

[0064] Among them, the discrete feature vector corresponding to user behavior data Text feature vectors Numerical eigenvectors and sequence feature vectors These are independent feature vectors. To calculate the data leakage risk corresponding to user behavior data, the discrete feature vectors need to be... Text feature vectors Numerical eigenvectors and sequence feature vectors To establish the relationship between the four elements, this embodiment uses a preset vector fusion method, concat, to combine the aforementioned discrete feature vectors. Text feature vectors Numerical eigenvectors and sequence feature vectors The feature vector is fused to obtain a fused feature vector, and this fused feature vector is denoted as the behavior feature vector. .

[0065] S220. Based at least on the behavioral feature vector and the historical data of the group of users, determine the risk of data leakage.

[0066] S230. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm, determine a data protection strategy.

[0067] S240. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0068] Example 3

[0069] This application provides a user behavior data protection method in Embodiment 3, which optimizes the "determining data leakage risk based at least on the behavior feature vector and historical data of the group of users" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0070] S310. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0071] S321. Discretize the behavioral feature vector to obtain a discrete vector.

[0072] The behavior feature vector is obtained by encoding at least one type of user behavior data. This behavior feature vector contains elements that correspond one-to-one with each type of user behavior data. For example, a behavior feature vector... For [view duration, number of clicks, average order value, and frequently purchased product categories].

[0073] It should be noted that in this embodiment, it is intended to calculate the behavior feature vector in the way of information entropy, so as to calculate the data leakage risk corresponding to the user behavior data. Therefore, it is necessary to first discretize the above-mentioned behavior feature vector firstly.

[0074] Taking the above-mentioned behavior feature vector as [browsing duration, click times, average order value, frequently purchased product categories] for example, where:

[0075] The discretization standard for browsing duration is: {short (0 - 30s), medium (31 - 180s), long (> 180s)};

[0076] The discretization standard for click times is: {few (1 - 5 times), many (> 5 times)};

[0077] The discretization standard for average order value is: {low (0 - 100 yuan), medium (101 - 500 yuan), high (> 500 yuan)};

[0078] The discretization standard for frequently purchased product categories is: {fast-moving consumer goods, digital products, luxury goods}.

[0079] If a specific behavior feature vector is [40s, 7 times, 125 yuan, digital products], then according to the above discretization standard, this behavior feature vector is discretized, and the discretized vector [medium, many, medium, digital products] can be obtained, and this discretized vector is denoted as the discrete vector.

[0080] S322. Determine the prior entropy based on the historical data of the group of users.

[0081] It should be noted that in this embodiment, the difference between the prior entropy and the conditional entropy corresponding to the user behavior data, that is, the mutual information, is used as the data leakage risk. Therefore, it is necessary to first determine the prior entropy corresponding to the user behavior data.

[0082] Suppose the privacy information of the user to be protected in this embodiment is the income level of the user, and the income level of the user is divided into three levels: high, medium, and low; it is also assumed that the historical data of the group of users contains the historical behavior data of 1000 users, among which there are 150 high-income personnel, 350 medium-income personnel, and 500 low-income personnel among these 1000 users. Then the proportion of high-income personnel is 0.15, the proportion of medium-income personnel is 0.35, and the proportion of low-income personnel is 0.5; the prior entropy H(S) corresponding to the historical data of the group of users is:

[0083] H(S)=-[0.15·log2(0.15)+0.35·log2(0.35)+0.50·log2(0.50)]≈1.44 bits.

[0084] S323. Determine the conditional entropy based on the discrete vector and the historical data of the group of users.

[0085] In step S321, the discrete vector calculated is [Medium, Many, Medium, Numeric]. Assuming there are 20 people in the group's historical user data who match this discrete vector, among these 20 people, 10 are high-income earners, 8 are middle-income earners, and 2 are low-income earners, then: the percentage of high-income earners matching the above discrete vector is 10 / 20 = 0.5, the percentage of middle-income earners matching the above discrete vector is 8 / 20 = 0.4, and the percentage of low-income earners matching the above discrete vector is 2 / 20 = 0.1. When the user behavior data obtained in step S310 also matches the discrete vector [Medium, Many, Medium, Numeric], then the conditional entropy H(S|R) corresponding to this user behavior data is:

[0086] H(S|R)=-[0.5·log2(0.5)+0.4·log2(0.4) + 0.1·log2(0.1)]≈1.36 bits.

[0087] S324. Based on the prior entropy and the conditional entropy, determine the risk of data leakage.

[0088] The difference between the prior entropy H(S) and the conditional entropy H(S|R) is the mutual information R, where R = H(S) - H(S|R). The mutual information R can be used to measure the risk of data leakage. To visually display this risk, the mutual information R needs to be normalized, that is, the normalized mutual information is used as the risk of data leakage.

[0089] S330. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm, determine a data protection strategy.

[0090] S340. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0091] Example 4

[0092] This application provides a user behavior data protection method in Embodiment 4, which optimizes the "determining data leakage risk based at least on the behavior feature vector and the historical data of the group of users" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0093] S410. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0094] S421. Encode the acquired model running environment data to obtain the environment feature vector.

[0095] It should be noted that, in order to calculate the data leakage risk corresponding to user behavior data, it is also necessary to obtain the model operating environment data corresponding to the smart device used by the user and encode the model operating environment data. In this embodiment, the model operating environment data includes, but is not limited to: device type, operating system, network type (WiFi / 4G), location information (latitude and longitude or grid ID), IP subnet, connection quality, current time, geographic semantics (home / company / public place), etc. The model operating environment data, like user behavior data, may also be divided into different data types, such as discrete behavior data and numerical behavior data, without specific limitations. In order to facilitate the quantification of the model operating environment data, so that the model operating environment data can participate in the calculation of data leakage risk.

[0096] Specifically, just as user behavior data is encoded, the acquired model operating environment data is also encoded to obtain a corresponding vector, which is denoted as the environment feature vector.

[0097] S422. Determine the state fusion vector based on the behavioral feature vector, the environmental feature vector, and the historical data of the group users.

[0098] Among them, user behavior data Data related to the model's runtime environment After encoding separately, behavioral feature vectors can be obtained. With environmental feature vectors Further calculations to assess the risk of data leakage require the inclusion of behavioral feature vectors. Environmental feature vectors and historical data of group users The unified input is fed into the model for processing; therefore, the behavioral feature vector must first be... Environmental feature vectors and historical data of group users The three elements are combined into a single vector to achieve fusion, and the resulting vector is denoted as the state fusion vector.

[0099] S423. The state fusion vector is processed based on a preset probability vector prediction model to obtain a prediction probability vector.

[0100] In order to determine the classification of the aforementioned state fusion vector, this embodiment also presets a probability vector prediction model. The probability vector prediction model can process the input state fusion vector and output the probability M that the state fusion vector belongs to a certain classification.

[0101] S424. Process the predicted probability vector based on the preset attack model to obtain the data leakage risk.

[0102] It should be noted that this embodiment intends to use an attack model to determine the data leakage risk corresponding to user behavior data. The attack model can be used to calculate the probability of whether the above-mentioned state fusion vector is in the training set of the probability vector prediction model, as a data leakage risk.

[0103] Specifically, the predicted probability vector M is input into a preset attack model for processing. The attack model outputs a probability value, which is used to characterize the probability of whether the state fusion vector is in the training set of the probability vector prediction model, and this probability value is recorded as the data leakage risk.

[0104] The higher the probability value, the more the probability vector prediction model remembers the state fusion vector, and the higher the risk of data leakage of user behavior data.

[0105] S430. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm, determine a data protection strategy.

[0106] S440. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0107] Example 5

[0108] This application provides a user behavior data protection method in Embodiment 5, which optimizes the "encoding of acquired model runtime environment data to obtain environment feature vectors" in Embodiment 4. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0109] S510. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0110] S521A. The discrete model operating environment data and the numerical model operating environment data in the acquired model operating environment data are encoded respectively to obtain discrete state features and numerical state features.

[0111] In this embodiment, the model runtime environment data Specifically, this includes discrete data and numerical data, denoted as discrete model runtime environment data and numerical model runtime environment data, respectively. To facilitate the quantification of the acquired model runtime environment data for subsequent calculation of data leakage risks, the discrete model runtime environment data and numerical model runtime environment data in the model runtime environment data must first be encoded separately.

[0112] Specifically, to encode the discrete model runtime environment data, this embodiment pre-sets a One-Hot encoding method or a Learned Embedding encoding method. By using the One-Hot encoding method or the Learned Embedding encoding method to encode the discrete model runtime environment data, the corresponding features of the discrete model runtime environment data can be obtained, which are denoted as discrete state features. To encode the numerical model runtime environment data, the numerical model runtime environment data can be directly normalized and standardized, and the obtained numerical components can be used as numerical state features.

[0113] S521B. Perform feature fusion on the discrete state features and the numerical state features to obtain an environmental feature vector.

[0114] The environmental feature vector is formed by fusing discrete state features and numerical state features.

[0115] Specifically, using the pre-defined combination method concat, discrete state features and numerical state features are combined into a vector, denoted as the environment feature vector.

[0116] S522. Based on the behavioral feature vector, the environmental feature vector, and the historical data of the group of users, determine the state fusion vector.

[0117] S523. The state fusion vector is processed based on a preset probability vector prediction model to obtain a prediction probability vector.

[0118] S524. Process the predicted probability vector based on the preset attack model to obtain the data leakage risk.

[0119] S530. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm, determine a data protection strategy.

[0120] S540. Based on the data protection strategy, the user behavior data is obfuscated to obtain obfuscated behavior data.

[0121] Example 6

[0122] This application provides a user behavior data protection method in Embodiment Six, which optimizes the "determining data leakage risk based at least on the behavior feature vector and historical data of the group of users" in Embodiment One. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0123] S610. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0124] S621. Encode the acquired model running environment data to obtain the environment feature vector.

[0125] S622. Determine the state fusion vector based on the behavioral feature vector, the environmental feature vector, and the historical data of the group users.

[0126] The specific implementation of steps S621-S622 can be found in steps S421-S422 of Example 4, and will not be repeated here.

[0127] S623. The state fusion vector is processed based on a preset differential privacy algorithm to obtain the data leakage risk.

[0128] In order to determine the data leakage risk corresponding to user behavior data, this embodiment intends to use a differential privacy algorithm to process the state fusion vector generated in S622.

[0129] Specifically, the state fusion vector is input into a preset differential privacy algorithm for processing, which outputs a privacy budget. It should be noted that this privacy budget can be used to characterize the risk of data leakage; the smaller the privacy budget, the lower the risk of data leakage of user behavior data. The privacy budget output by the differential privacy algorithm is denoted as the data leakage risk.

[0130] S630. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm, determine a data protection strategy.

[0131] S640. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0132] Example 7

[0133] This application provides a user behavior data protection method in Embodiment Seven, which supplements the method described in Embodiment One. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0134] S710. Perform feature encoding on user behavior data to obtain behavior feature vectors.

[0135] S720. Based at least on the behavioral feature vector and the historical data of the group of users, determine the risk of data leakage.

[0136] S730. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm, determine a data protection strategy.

[0137] S740. Obfuscate the user behavior data based on the data protection strategy to obtain obfuscated behavior data.

[0138] S750: Based on a preset reward function, the data leakage risk is processed, and the user experience loss and device execution overhead are obtained to obtain a reward value.

[0139] The strategy decision model is a pre-trained strategy agent (a deep learning model). When applied to this method, the strategy decision model can output a data protection strategy. This strategy is used to protect the acquired user behavior data and balance the user experience of the software. Sometimes, the data protection strategy output by the strategy decision model may be too strict. For example, the original data strategy should be "inserting fake behavior," but the strategy decision model outputs "replacing behavior," which results in the complete replacement of user behavior data. The software currently used by the user cannot provide more accurate information push based on the user behavior data. For example, in the case of a shopping app, it cannot push more accurate products to the user, which leads to a decline in user experience.

[0140] Therefore, even though the aforementioned strategy decision-making model has been applied to this method, it will be continuously and dynamically updated to improve the accuracy of the strategy decision-making model in balancing the strength of user behavior data protection and the user software experience.

[0141] Therefore, after executing steps S710-S740 of this method, user experience data is further collected and recorded as user experience penalty UX_penalty. For example, user experience penalty UX_penalty can be user experience score data or task delay data, etc., without any specific limitation. In addition, the device overhead caused by the smart device executing this method is also obtained and recorded as device execution overhead CPU_cost.

[0142] To optimize the strategy decision-making model, a corresponding reward value needs to be calculated using a pre-defined reward function, and the strategy decision-making model is then optimized based on this reward value. The formula for calculating the reward function is as follows:

[0143] ;

[0144] Where Rt is the reward value, α, β, and γ are preset weighting coefficients; P_leakage is the data leakage risk, UX_penalty is the user experience penalty, and CPU_cost is the device execution cost; where P_leakage is a normalized value; the preset weighting coefficients α, β, and γ all range from 0 to 1, and their sum is 1; UX_penalty is the user experience score after the user behavior data is protected by the data protection strategy generated by the policy decision model, and for example, the score ranges from 0 to 10; CPU_cost is the CPU task execution time, which can be in milliseconds (ms), microseconds (µs), or nanoseconds (ns), for example, CPU_cost is the time taken for the policy decision model to generate the data protection strategy. It should be noted that when calculating the reward value Rt, the units of each parameter (such as the unit of CPU_cost) are not used, only the numerical values ​​of each parameter are used.

[0145] S760. Optimize the engine parameters of the strategy decision model based on the reward value to obtain a new strategy decision model.

[0146] The reward value Rt is used to optimize the parameters of each engine in the strategic decision engine to generate a new strategic decision model.

[0147] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0148] Example 8

[0149] Based on the same inventive concept, this embodiment also provides a user behavior data protection device for implementing the user behavior data protection method described above. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more user behavior data protection device embodiments provided below can be found in the limitations of the user behavior data protection method described above, and will not be repeated here.

[0150] In this embodiment, as Figure 2 As shown, a user behavior data protection device is provided, comprising:

[0151] The vector calculation module is used to encode user behavior data to obtain behavior feature vectors.

[0152] The risk calculation module is used to determine the risk of data leakage based at least on the behavioral feature vector and the historical data of the group of users;

[0153] The strategy generation module is used to determine a data protection strategy based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm.

[0154] The data obfuscation module is used to obfuscate the user behavior data based on the data protection policy to obtain obfuscated behavior data.

[0155] Each module in the aforementioned user behavior data protection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0156] It should be noted that this embodiment obtains a behavioral feature vector by feature encoding user behavior data; determines the data leakage risk based at least on the behavioral feature vector and historical data of the group of users; determines a data protection strategy based on the behavioral feature vector, model runtime environment data, the data leakage risk, and a strategy decision model preset by a deep reinforcement learning algorithm; and obfuscates the user behavior data based on the data protection strategy to obtain obfuscated behavioral data. Through the above implementation, the data leakage risk of user behavior data is first calculated, and then the behavioral feature vector corresponding to the user behavior data, model runtime environment data, and the data leakage risk are processed by the strategy decision model to generate a data protection strategy for obfuscating user behavior data (data leakage prevention processing). Since this data protection strategy fully considers the data leakage risk of user behavior data, it can effectively achieve a balance between privacy protection strength and user experience.

[0157] In an optional embodiment, the user behavior data includes at least: discrete behavior data, text behavior data, numerical behavior data, and sequence behavior data;

[0158] Accordingly, the step of encoding user behavior data to obtain a behavior feature vector includes:

[0159] The discrete behavior data, text behavior data, numerical behavior data, and sequence behavior data in the user behavior data are encoded respectively to obtain discrete feature vectors, text feature vectors, numerical feature vectors, and sequence feature vectors.

[0160] The discrete feature vector, the text feature vector, the numerical feature vector, and the sequence feature vector are fused to obtain the behavioral feature vector.

[0161] In an optional embodiment, determining the data leakage risk based at least on the behavioral feature vector and historical data of the group of users includes:

[0162] Discretize the behavioral feature vector to obtain a discrete vector;

[0163] Determine the prior entropy based on the historical data of the aforementioned user group;

[0164] Based on the discrete vector and the historical data of the group of users, determine the conditional entropy;

[0165] Based on the prior entropy and the conditional entropy, the risk of data leakage is determined.

[0166] In an optional embodiment, determining the data leakage risk based at least on the behavioral feature vector and historical data of the group of users includes:

[0167] The acquired model runtime environment data is encoded to obtain an environmental feature vector;

[0168] Based on the behavioral feature vector, the environmental feature vector, and the historical data of the group of users, a state fusion vector is determined.

[0169] The state fusion vector is processed based on a preset probability vector prediction model to obtain a prediction probability vector;

[0170] The predicted probability vector is processed based on a preset attack model to obtain the data leakage risk.

[0171] In an optional embodiment, the model runtime environment data includes at least: discrete model runtime environment data and numerical model runtime environment data;

[0172] Accordingly, encoding the acquired model runtime environment data to obtain an environment feature vector includes:

[0173] The discrete model operating environment data and the numerical model operating environment data in the acquired model operating environment data are encoded respectively to obtain discrete state features and numerical state features;

[0174] The discrete state features and the numerical state features are fused to obtain an environmental feature vector.

[0175] In an optional embodiment, determining the data leakage risk based at least on the behavioral feature vector and historical data of the group of users includes:

[0176] The acquired model runtime environment data is encoded to obtain an environmental feature vector;

[0177] Based on the behavioral feature vector, the environmental feature vector, and the historical data of the group of users, a state fusion vector is determined.

[0178] The state fusion vector is processed based on a preset differential privacy algorithm to obtain the risk of data leakage.

[0179] In an optional embodiment, the user behavior data protection device further includes:

[0180] The reward value calculation module is used to process the data leakage risk based on a preset reward function, as well as the acquired user experience loss and device execution overhead, to obtain the reward value;

[0181] The engine optimization module is used to optimize the engine parameters of the strategy decision model based on the reward value to obtain a new strategy decision model.

[0182] Example 9

[0183] In this embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a user behavior data protection method.

[0184] Those skilled in the art will understand thatFigure 3 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0185] Example 10

[0186] In this embodiment, a computer-readable storage medium is provided, such as... Figure 4 As shown, a computer program is stored thereon, and when the computer program is executed by the processor, it implements the steps in the above-described method embodiments.

[0187] Example 11

[0188] In this embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0189] It should be noted that the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0190] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0191] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0192] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.

Claims

1. A method for protecting user behavior data, characterized in that, include: User behavior data is feature-encoded to obtain a behavior feature vector; The risk of data leakage should be determined based at least on the aforementioned behavioral feature vectors and historical data of the user group. Based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model pre-set by a deep reinforcement learning algorithm, a data protection strategy is determined. The user behavior data is obfuscated based on the data protection strategy to obtain obfuscated behavior data.

2. The method according to claim 1, characterized in that, The user behavior data includes at least: discrete behavior data, text behavior data, numerical behavior data, and sequence behavior data; Accordingly, the step of encoding user behavior data to obtain a behavior feature vector includes: The discrete behavior data, text behavior data, numerical behavior data, and sequence behavior data in the user behavior data are encoded respectively to obtain discrete feature vectors, text feature vectors, numerical feature vectors, and sequence feature vectors. The discrete feature vector, the text feature vector, the numerical feature vector, and the sequence feature vector are fused to obtain the behavioral feature vector.

3. The method according to claim 1, characterized in that, The determination of data leakage risk, based at least on the behavioral feature vector and historical data of the user group, includes: Discretize the behavioral feature vector to obtain a discrete vector; Determine the prior entropy based on the historical data of the aforementioned user group; Based on the discrete vector and the historical data of the group of users, determine the conditional entropy; Based on the prior entropy and the conditional entropy, the risk of data leakage is determined.

4. The method according to claim 1, characterized in that, The determination of data leakage risk, based at least on the behavioral feature vector and historical data of the user group, includes: The acquired model runtime environment data is encoded to obtain an environmental feature vector; Based on the behavioral feature vector, the environmental feature vector, and the historical data of the group of users, a state fusion vector is determined; The state fusion vector is processed based on a preset probability vector prediction model to obtain a prediction probability vector; The predicted probability vector is processed based on a preset attack model to obtain the data leakage risk.

5. The method according to claim 4, characterized in that, The model runtime environment data includes at least: discrete model runtime environment data and numerical model runtime environment data; Accordingly, encoding the acquired model runtime environment data to obtain an environment feature vector includes: The discrete model operating environment data and the numerical model operating environment data in the acquired model operating environment data are encoded respectively to obtain discrete state features and numerical state features; The discrete state features and the numerical state features are fused to obtain an environmental feature vector.

6. The method according to claim 1, characterized in that, The determination of data leakage risk, based at least on the behavioral feature vector and historical data of the user group, includes: The acquired model runtime environment data is encoded to obtain an environmental feature vector; Based on the behavioral feature vector, the environmental feature vector, and the historical data of the group of users, a state fusion vector is determined; The state fusion vector is processed based on a preset differential privacy algorithm to obtain the risk of data leakage.

7. The method according to claim 1, characterized in that, Also includes: The data leakage risk is processed based on a preset reward function, and the user experience loss and device execution overhead are obtained to obtain a reward value; The engine parameters of the strategy decision-making model are optimized based on the reward value to obtain a new strategy decision-making model.

8. A user behavior data protection device, characterized in that, The device includes: The vector calculation module is used to encode user behavior data to obtain behavior feature vectors. The risk calculation module is used to determine the risk of data leakage based at least on the behavioral feature vector and the historical data of the group of users; The strategy generation module is used to determine a data protection strategy based on the behavioral feature vector, model operating environment data, data leakage risk, and a strategy decision model preset based on a deep reinforcement learning algorithm. The data obfuscation module is used to obfuscate the user behavior data based on the data protection policy to obtain obfuscated behavior data.

9. A terminal device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data transmission method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.