A data flow control method and device, electronic equipment and storage medium

By assessing the sensitivity of information and trust relationships in data circulation scenarios, and dynamically determining privacy and security conditions and target privacy decisions, this approach solves the problem of balancing privacy protection and data sharing in existing data circulation control technologies, and achieves privacy risk assessment and control of data circulation.

CN119939663BActive Publication Date: 2026-05-05CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNIV OF PETROLEUM (BEIJING)
Filing Date
2025-01-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing data flow control technologies struggle to balance the privacy protection and data sharing needs of data subjects, fail to meet the requirements of different scenarios, and are cumbersome to set up, making it difficult to achieve an effective balance between privacy protection and data sharing.

Method used

By acquiring data to be circulated and the data circulation scenario, the sensitivity of information, the trust value and privacy attitude of data subjects and data audiences are assessed. Based on a set of preset privacy protection strategies and prospect theory, privacy security conditions and target privacy decisions are dynamically determined to achieve privacy risk assessment and control of data circulation.

Benefits of technology

It enables a dynamic balance between data subject privacy protection and data sharing without pre-setting access conditions, adapting to the data circulation needs of different scenarios and improving the privacy security and efficiency of data circulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939663B_ABST
    Figure CN119939663B_ABST
Patent Text Reader

Abstract

This invention provides a data circulation control method, apparatus, electronic device, and storage medium. The method includes: determining information sensitivity, the data subject's trust value towards the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario; determining privacy security conditions based on a preset set of privacy protection strategies, information sensitivity, the data audience's privacy attitude towards the data subject, and the data subject's trust value towards the data audience; determining the privacy risk of each data audience based on the information sensitivity, the data audience's privacy attitude towards the data subject, and the data subject's risk preference value in the privacy security conditions; and determining a target privacy decision for each data audience based on the privacy risk and privacy security conditions, so as to control the data to be circulated to be shared according to the target privacy decision. This achieves a balance between data subject privacy protection and data sharing based on reinforcement learning methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data flow control method, apparatus, electronic device, and storage medium. Background Technology

[0002] Existing data flow control technologies mostly rely on access control or static solutions.

[0003] Access control-based methods pre-set conditions for information flow, controlling information within predetermined scope, such as allowing information to circulate only within a specified domain or allowing access only to users with specific permissions. These methods require pre-setting data access conditions, which is cumbersome and difficult to meet the needs of different scenarios, failing to balance the data subject's privacy protection with the data sharing requirements. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a data flow control method, apparatus, electronic device, and storage medium to solve the problem in the prior art that cannot balance the protection of data subject privacy and data sharing.

[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0006] A first aspect of this invention discloses a data flow control method, the method comprising:

[0007] Acquire the data to be circulated and its corresponding data circulation scenario, wherein the data to be circulated includes the data subject and its data content;

[0008] The sensitivity of the information, the trust level of the data subject to the data audience, and the privacy attitude of the data audience towards the data subject are determined based on the data to be circulated and the data circulation scenario.

[0009] Privacy and security conditions are determined based on a set of preset privacy protection strategies, the sensitivity of the information, the data audience's attitude toward the data subject's privacy, and the data subject's trust in the data audience.

[0010] The privacy risk of each data audience is determined based on the sensitivity of the information in the privacy and security conditions, the data audience's attitude toward the privacy of the data subject, and the data subject's risk preference value.

[0011] Based on the privacy risks and privacy security conditions of each data audience, a target privacy decision is determined for each data audience in order to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0012] Optionally, determining the information sensitivity, the data subject's trust in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario includes:

[0013] The sensitivity of information is determined based on the data content in the data to be circulated and the data circulation scenario.

[0014] The data audience's attitude towards the privacy of the data subject is determined based on the data to be circulated;

[0015] The trust level of the data subject towards the data audience is determined based on the privacy attitude and the data to be circulated.

[0016] Optionally, determining the data subject's trust level in the data audience based on the privacy attitude and the data to be circulated includes:

[0017] Set the initial trust value of the data subject to the data audience;

[0018] The trust impact function of the privacy threat is calculated based on the sensitivity of the information and the number of the data audience.

[0019] The trust value of the data subject to the data audience is determined based on the initial trust value and the trust impact function of the privacy threat.

[0020] Optionally, the privacy risk of each data audience is determined based on the information sensitivity in the privacy and security conditions, the data audience's attitude towards the data subject's privacy, and the data subject's risk preference value, including:

[0021] The probability of the data audience forwarding the data subject's data is determined based on the privacy attitude.

[0022] The initial privacy risk is determined by calculating based on the sensitivity of the information and the probability that the data to be circulated will be forwarded.

[0023] The privacy risk of each data audience is determined based on the data subject's risk preference value and the initial privacy risk, wherein the risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario.

[0024] Optionally, determining the target privacy decision for each data audience based on the privacy risks and privacy security conditions for each data audience includes:

[0025] For each data audience, a privacy protection strategy for that data audience is selected from the preset privacy protection strategy set of the privacy and security conditions;

[0026] Based on the trust value of the data subject to the data audience in the privacy and security conditions, determine the benefits of the data audience to the data subject under each privacy policy;

[0027] For each data audience, the target privacy decision is determined based on the data subject's benefits under each privacy policy and the data audience's privacy risks.

[0028] Optional, also includes:

[0029] For each data subject, cumulative privacy risks are determined based on the data subject's historical privacy risks;

[0030] The corresponding historical benefits are determined based on the historical sharing utility value and historical privacy risks of the data subject under each privacy policy.

[0031] The accumulated privacy risks, historical privacy protection strategies, and historical benefits are used to determine the privacy protection objectives for the circulation of data elements.

[0032] Based on the privacy protection objectives of data element circulation, a target privacy protection strategy is determined for data subjects to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0033] A second aspect of the present invention discloses a data flow control device, the device comprising:

[0034] A determining unit is used to acquire the data to be circulated and its corresponding data circulation scenario, wherein the data to be circulated includes the data subject and the data content;

[0035] A privacy and security condition generation unit is used to determine the information sensitivity, the data subject's trust value in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario; and to determine privacy and security conditions based on a preset set of privacy protection strategies, the information sensitivity, the data audience's privacy attitude towards the data subject, and the data subject's trust value towards the data audience.

[0036] A privacy risk prediction unit is used to determine the privacy risk of each data audience based on the sensitivity of the information in the privacy security conditions, the data audience's attitude towards the privacy of the data subject, and the data subject's risk preference value.

[0037] The processing unit is configured to determine a target privacy decision for each data audience based on the privacy risks and privacy security conditions of each data audience, so as to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0038] Optionally, the privacy risk prediction unit is specifically used for:

[0039] The probability of the data audience forwarding the data subject's data is determined based on the privacy attitude.

[0040] The initial privacy risk is determined by calculating based on the sensitivity of the information and the probability that the data to be circulated will be forwarded.

[0041] The privacy risk of each data audience is determined based on the data subject's risk preference value and the initial privacy risk, wherein the risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario.

[0042] A third aspect of the present invention discloses an electronic device, the electronic device including a processor and a memory, the memory being used to store program code and data for data generation, and the processor being used to call program instructions in the memory to execute the data flow control method as described in the first aspect of the present invention.

[0043] A fourth aspect of the present invention discloses a storage medium including a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the data flow control method as described in the first aspect of the present invention.

[0044] Based on the above embodiments of the present invention, a data circulation control method, apparatus, electronic device, and storage medium are provided. The method includes: acquiring data to be circulated and its corresponding data circulation scenario; characterizing the privacy and security conditions of the data subject based on the data to be circulated and the data circulation scenario; assessing the privacy risk of each data audience based on the privacy and security conditions, that is, assessing the potential privacy risk posed by each data audience to the data subject based on prospect theory; and determining a target privacy decision for each data audience based on the privacy risk of each data audience and the privacy and security conditions, so as to control the sharing of the data to be circulated according to the target path. This method does not require pre-setting data access conditions and can achieve a balance between data subject privacy protection and data sharing based on reinforcement learning methods. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating a data flow control method according to an embodiment of the present invention;

[0047] Figure 2This is a schematic diagram illustrating privacy security conditions, privacy risks, and privacy decision generation in an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the structure of a data flow control device according to an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0051] It should be noted that the descriptions involving "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0052] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0053] See Figure 1 This invention provides a data flow control method, which includes:

[0054] Step S101: Obtain the data to be circulated and its corresponding data circulation scenario;

[0055] In the specific implementation step S101, the current data to be circulated and the data circulation scenario input by the user are obtained.

[0056] Before data can be circulated, it is necessary to assess the privacy risks faced by data subjects under the influence of their subjective privacy risk preferences based on prospect theory, in order to determine whether the data can be circulated.

[0057] It should be noted that the data to be circulated includes data content and data body, and the data content includes multiple fields.

[0058] Different data circulation scenarios correspond to different data audiences, therefore, it is necessary to obtain the data audience corresponding to the aforementioned data circulation scenarios.

[0059] Optionally, the privacy risks to data subjects vary depending on the data's circulation path or the audience it faces. Therefore, it is necessary to predict the privacy risks during data circulation from the perspective of the data subject.

[0060] Step S102: Determine the information sensitivity, the data subject's trust in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario;

[0061] The specific implementation process of step S102 includes the following steps:

[0062] Step S11: Determine the information sensitivity M based on the data content in the data to be circulated and the data circulation scenario.

[0063] The specific implementation of step S11 includes the following steps:

[0064] Step S21: Determine the level corresponding to each other data element based on the data subject;

[0065] Because data subjects have varying degrees of sensitivity to data in different data circulation scenarios, several sensitivity levels are pre-defined to accommodate different data circulation scenarios.

[0066] The levels include S, A, B, C, D, and E, with the level decreasing from left to right.

[0067] Step S22: Based on the level of each data element, traverse the preset level table to determine the sensitivity of each data element to the information of the data subject.

[0068] It should be noted that a pre-defined correspondence between different levels and information sensitivity is established, and a preset level table is constructed based on this correspondence.

[0069] In the specific implementation step S22, the preset level table is traversed to find the information sensitivity corresponding to the level of each data element.

[0070] It should be noted that information sensitivity is a number between 0 and 1, and the higher the level, the closer the corresponding information sensitivity is to 1.

[0071] For example, in a cross-institutional electronic medical record (EMR) workflow scenario, the patient is the data subject. The level of data 1 affected by this data subject in this scenario is determined to be E, and the level of data 2 is determined to be B. Then, a preset level table is traversed to find the information sensitivity level of 0.4 corresponding to level E of data 1, and the information sensitivity level of 0.7 corresponding to level B of data 2.

[0072] Among them, data element 1 is the ID card number and data element 2 is the home address.

[0073] Step S12: Determine the data audience's attitude towards the privacy of the data subject based on the data to be circulated.

[0074] In the specific implementation step S12, the data audience's attitude towards the data subject's privacy is determined based on the personal attributes of the data audience in the data to be circulated.

[0075] Specifically, the personal attributes of the data audience are input into the privacy information dissemination model so that the privacy information dissemination model maps the personal attributes to the corresponding privacy attitudes, thereby determining the privacy attitude e of each data audience towards the circulating data.

[0076] It should be noted that the privacy information propagation model is used to describe how the forwarding decisions of data audiences affect the propagation of privacy information in social networks. In other words, the privacy information propagation model is used to describe the personal attributes of data audiences and their attitudes toward the privacy of circulating data.

[0077] Privacy attitude, or the data audience's tendency to protect the privacy of the data subject, refers to the data audience's inclination in protecting the privacy of the subject, including protective attitude and non-protective attitude.

[0078] Step S13: Determine the trust value of the data subject to the data audience based on the privacy attitude and the data to be circulated.

[0079] It should be noted that the specific implementation of step S13 includes the following steps:

[0080] Step S31: Set the initial trust value of the data subject to the data audience;

[0081] In the specific implementation of step S31, let the initial trust value of the data subject towards the data audience be denoted as _____. .

[0082] Step S32: Calculate the trust impact function of the privacy threat based on the sensitivity of the information and the number of the data audience.

[0083] In the specific implementation step S32, since the privacy threat faced by the data subject is proportional to the sensitivity M of the data itself and the number of information audiences K, the product of the information sensitivity of each data element to the data subject and the number of information audiences is calculated to obtain the potential privacy threat r caused by the information forwarding behavior, i.e., r = M * K; then, the reciprocal of the potential privacy threat r caused by the information forwarding behavior is used as the trust influence function f(r) regarding the privacy threat, i.e. .

[0084] It should be noted that since the sensitivity of information is applied to each data audience, the number of trust impact functions f(r) obtained for the privacy threat is the same as the number of data audiences, and consequently, the number of trust values ​​calculated subsequently is also multiple.

[0085] Step S33: Determine the trust value of the data subject to the data audience based on the initial trust value and the trust impact function of the privacy threat.

[0086] In the specific implementation step S33, after receiving the data, the data audience becomes a data receiver. If the data receiver further forwards the data, the data subject's trust value towards the data receiver will change accordingly. Specifically, the initial trust value... The product of the trust influence function f(r) of the privacy threat updates the trust value T of the data subject to the data audience, i.e. .

[0087] Where f(r) is the trust impact function of the privacy threat.

[0088] This application can perform dynamic trust assessment of a data subject on a data recipient based on the privacy threats posed to the data subject by the data recipient's information forwarding behavior.

[0089] Step S103: Determine privacy and security conditions based on a preset set of privacy protection strategies, the sensitivity level M of the information, the privacy attitude e of the data audience towards the data subject, and the trust value T of the data subject towards the data audience.

[0090] The target privacy decision is determined based on the historical utility of the data subject's privacy risk preferences.

[0091] This invention proposes the concept of Privacy and Security Conditions (PSCs) to describe the conditions that data subjects must meet to achieve privacy protection in a given context. Based on the fundamental elements of data circulation scenarios and the relationships between these elements, the core elements of the PSC are selected without loss of generality: data sensitivity (M), the data subject's trust level in the data audience (T), and the data audience's privacy protection inclination (e). Other elements can be added according to specific scenarios and the specific privacy needs of the data subject, with the elements in the PSC determining the choice of privacy protection strategy.

[0092] In the specific implementation step S103, a set of privacy protection policies is first pre-set for the data audience during the data circulation process. This set may include shareable policies, non-shareable policies, and other privacy de-identification policies. The specific privacy protection policy for each data audience is selected from the pre-set set of privacy protection policies.

[0093] Next, based on the influence of the information sensitivity level M, the data audience's privacy attitude e towards the data subject, and the data subject's trust value T towards the data audience on the privacy protection strategy, a preset privacy protection set is established. Substituting into formula (1), we obtain the privacy and security condition PSC. In other words, the privacy and security condition PSC at this time includes a set of preset privacy protection policies. The sensitivity of the information M, the privacy attitude e of the data audience towards the data subject, and the trust value T of the data subject towards the data audience.

[0094] Formula (1):

[0095] (1)

[0096] Among them, a set of preset privacy protection policies , A set of m preset privacy protection policies that can be selected for the data subject.

[0097] It should be further noted that the privacy and security conditions (PSC) at this point are the conditions that must be met to satisfy the privacy needs of the data subject.

[0098] The degree of trust T between the data subject and the data audience, and the data audience's tendency to protect the privacy of the data subject, can both be sets or numerical ranges.

[0099] Step S104: Based on the information sensitivity M, the data audience's privacy attitude e towards the data subject, and the data subject's risk preference value in the privacy and security conditions. Determine the privacy risk g(r) for each data audience.

[0100] Specifically, based on the influencing factors (M, T, e) in PSC, the risk preference value of the data subjects is assessed based on prospect theory. Determine the impact of privacy risks g(r) on each data audience.

[0101] In this application, data to be circulated will pose different privacy risks to the data subject depending on the circulation path or the data audience. Therefore, it is necessary to predict the privacy risks during the data circulation process from the perspective of the data subject, that is, to predict the privacy risks to the data subject from different data audiences. This invention provides a privacy risk prediction method based on prospect theory, so that the prediction results can reflect the data subject's subjective privacy risk preferences.

[0102] The specific implementation process of step S104 includes the following steps:

[0103] Step S41: Determine the forwarding probability of the data subject by the data audience based on the privacy attitude.

[0104] In the specific implementation step S41, the number of data audiences with a non-protective attitude towards the privacy of the data to be circulated is calculated and used as the first number. Then, the ratio of the first number to the total data of all data audiences is calculated to obtain the probability that the data subject will be forwarded by the data audience, i.e., the forwarding probability. In other words, the privacy attitude of information audiences towards the circulated data is calculated to obtain the probability distribution of each information audience's privacy protection attitude and non-privacy protection attitude towards the data subject. Furthermore, the probability that the data audience will forward the circulated data is evaluated, i.e., the forwarding probability P(A).

[0105] Among them, P(A) is directly proportional to the data audience's tendency to protect the privacy of the data subject.

[0106] Step S42: Calculate and determine the initial privacy risk r based on the information sensitivity M and the probability P(A) of the data to be circulated being forwarded.

[0107] In the specific implementation of step S42, without considering the subjective privacy risk preferences of the data subject, the privacy risk r caused by the data receiver is the product of the sensitivity of the information M and the probability P(A) of the data to be circulated being forwarded, i.e., r=M*P(A).

[0108] Step S43: Risk preference value based on data subject The initial privacy risk r is used to determine the privacy risk g(r) for each data audience.

[0109] The risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario, i.e., subjective privacy risk preference, and the specific value can be set by the data subject.

[0110] In the specific implementation step S43, for each data audience, firstly, the risk preference value of the data audience in the data circulation scenario is obtained; then, the risk preference value of the data subject is... Substituting the initial privacy risk r into formula (2), the privacy risk g(r) of the data audience is determined.

[0111] Formula (2):

[0112] (2)

[0113] Where, parameter α>1, parameter , parameter α and parameter The values ​​are all preset.

[0114] Formula (2) expresses the impact of the data subject's subjective privacy risk preference on the real privacy risk. The data subject can select different data audience risk preferences in different data circulation scenarios, thereby obtaining the privacy risk in that scenario.

[0115] Step S105: Determine a target privacy decision for each data audience based on the privacy risks and privacy security conditions, so as to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0116] The specific implementation of step S105 includes the following steps:

[0117] Step S51: For each data audience, select a privacy protection strategy for the data audience from the preset privacy protection strategy set of the privacy and security conditions.

[0118] The privacy protection strategy includes shareable and non-shareable strategies for the data audience, and may also include other privacy desensitization strategies.

[0119] In the specific implementation of step S51, for each data audience, a preset optional privacy protection strategy is first set, including a shareable strategy and a non-shareable strategy. Privacy desensitization strategies can be added as needed so that the optimal strategy, i.e. the target strategy, can be determined later.

[0120] Step S52: Based on the trust value of the data subject to the data audience in the privacy and security conditions, determine the benefits of the data audience to the data subject under each privacy protection strategy;

[0121] Specifically, when a data subject chooses to share data, the data subject's benefit s(t) under the shareable strategy is related to the data subject's trust value T towards the data audience. In other words, the data subject's benefit s(t) under the shareable strategy is directly proportional to the data subject's trust value T towards the data audience. Based on this, the data subject's trust value T in the data audience is used to determine the data subject's benefit s(t) under the shareable strategy.

[0122] At this point, the privacy loss is consistent with the privacy risk prediction result, i.e., C. i =g(r).

[0123] When a data subject chooses not to share data with its data audience, i.e., when the data audience adopts a non-sharing strategy, the data subject's benefit under the non-sharing strategy is 0, i.e., s(t) = 0.

[0124] The privacy loss at this point corresponds to a privacy risk outcome where the data subject's privacy risk preference value is 0, i.e., C. i =g(0).

[0125] Step S53: For each data audience, based on the benefits of the data subject to the data audience under each privacy protection strategy and the privacy risk g(r) of the data audience, determine the target privacy decision for the data audience.

[0126] Specifically, a non-shareable strategy targeting the data audience, namely... =0, will =0, the data subject's benefit s(t)=0, privacy risk result with privacy risk preference value of 0, i.e. privacy risk g(0) is substituted into formula (3) for processing to obtain the utility value of the data subject under the non-shareable strategy;

[0127] Next, a shareable strategy for the data audience, namely... =1, will enable the shareable strategy =1. The data audience's benefit s(t) under the shareable strategy, and the privacy risk C. i =g(r) is substituted into formula (3) for processing to obtain the utility value of the data subject in the shareable strategy;

[0128] Then, for each other privacy policy for the data audience, the policy will be... The value of , the benefit s(t) of the data subject, and the privacy risk are substituted into formula (3) for processing to obtain the utility value of the data subject in other privacy strategies.

[0129] Finally, the strategy with the highest utility value is taken as the optimal privacy decision, i.e., the target privacy decision. In other words, the final utility of the data subject under each privacy protection strategy is compared, and the privacy protection strategy corresponding to the maximum final benefit is taken as the optimal privacy decision, i.e., the target privacy decision.

[0130] Based on the above specific process, the target privacy decision for each data audience is determined by formula (3).

[0131] Accordingly, data subjects can choose the optimal privacy decision based on formula (3):

[0132] Formula (3):

[0133] (3)

[0134] in, s represents the set of privacy protection policies for data subjects. i (t) represents the benefit corresponding to the i-th privacy protection strategy, g i (r) represents the privacy risk corresponding to the i-th privacy protection policy, and its value is negative.

[0135] In this embodiment of the invention, steps S51 to S53 are applicable to cold start scenarios, i.e., when there is limited historical interaction information between the data subject and the data audience. Furthermore, to achieve a balance between privacy protection and data sharing, this invention uses a reinforcement learning-based privacy decision-making method to determine the optimal privacy decision, thereby controlling the sharing of the data to be circulated according to the target privacy decision.

[0136] To better understand the methods illustrated in the above embodiments of the present invention, examples are provided below.

[0137] This invention validates path control in social networks, assessing how changes in privacy risks faced by data subjects when sharing data occur under the use of path control. Data disseminated on social networks encounters diverse and heterogeneous audiences.

[0138] If the trust value of the data subject V among social network users is determined to be Gaussian distribution The privacy protection preferences of data receiver W towards data subject V are uniformly distributed. .

[0139] Three different privacy scenarios were set up, namely scenario x, scenario y and scenario z, with information sensitivity levels of 0.8, 0.5 and 0.2 respectively, and privacy risk preference values ​​of data subjects of 0.8, 0.5 and 0.2 respectively.

[0140] With a data sensitivity level of M=0.8 and the data subject's privacy risk preference value... For example, let's take 0.5 as an example.

[0141] First, based on sampling the trust value T among social network users and the data receiver's privacy protection tendency e towards the data subject, the privacy and security conditions for the data subject W are determined, for example: PSC=({l},0.6,0.8), where privacy decision... , A value of 1 indicates that data is being sent. A value of 0 indicates that no data is sent.

[0142] It should be noted that the privacy decisions here include targeted privacy decisions for multiple data audiences.

[0143] Then, based on prospect theory, the privacy risk g(r) faced by data subject V is predicted.

[0144]

[0145] Where r = 0.8 * P(A), .

[0146] Then, the optimal privacy decision in this scenario is obtained using formula (3), which maximizes the utility of data sharing while satisfying privacy and security.

[0147] In this embodiment of the invention, the privacy and security conditions of data subjects are characterized based on their privacy needs, resulting in the conditions required to satisfy those needs. Then, prospect theory is used to assess the privacy risks faced by data subjects within the data circulation context. Finally, based on the data subjects' own privacy and security conditions and the privacy risks they face, a privacy protection strategy is determined for each data audience, transforming the privacy protection problem into a utility maximization problem, thus obtaining control decisions for the data circulation path. This achieves a balance between data subject privacy protection and data sharing without requiring pre-setting data access.

[0148] Optionally, based on the method shown in the above embodiments of the present invention, the present invention also shows another implementation, including the following steps:

[0149] Step S61: For each data subject, determine the cumulative privacy risk based on the data subject's historical privacy risks.

[0150] In the specific implementation of step S61, after executing step S101, for each data audience, obtain each historical privacy risk g(r) of data element circulation within the historical time period; and accumulate them to obtain the cumulative privacy risk S=∑g(r).

[0151] In practical implementation, the cumulative privacy risks S faced by data subjects in the historical circulation path of data elements are taken as the state of the Markov Decision Process (MDP).

[0152] Step S62: Determine the corresponding historical benefits based on the historical sharing utility value and historical privacy risk g(r) of the data subject in each privacy policy.

[0153] In the specific implementation step S62, the historical sharing utility value s of the data subject in each privacy policy and the historical privacy risk g(r) are substituted into formula (4) for processing to obtain the benefit for each privacy protection policy.

[0154] (4)

[0155] in, The benefit is defined as the difference between the utility of sharing data elements and the privacy risks. This represents the data sharing benefits that a data subject receives when adopting a privacy protection strategy l, assuming a trust value T for the data audience. This indicates the risk of privacy breaches faced by data subjects when adopting privacy protection strategies.

[0156] Step S63: Determine the privacy protection objectives for data element circulation based on the accumulated privacy risks, historical privacy protection strategies, and historical benefits.

[0157] First, the privacy protection decision-making process of data subjects during the flow of data elements is described based on MDP.

[0158] The states, actions, and benefits of an MDP are defined as follows:

[0159] State refers to the cumulative privacy risks faced by a data subject in the historical circulation path of data elements in the context of data element circulation, i.e., S=∑g(r).

[0160] Actions refer to the set of privacy protection strategies that data subjects can take. Specific historical privacy protection strategies .

[0161] The benefit refers to the difference between the data element sharing utility and the privacy risk, which is given that the circulation of data elements brings certain positive benefits to the data subject, but also causes the risk of privacy leakage. That is, formula (4).

[0162] Next, based on the description of MDP, the privacy protection goal of data element circulation is determined, namely, to maximize the historical reward of the data element circulation process by selecting appropriate strategies for each data audience.

[0163] Specifically, based on the description of MDP, the calculation is performed by substituting it into formula (5) to obtain the privacy protection target for data element circulation.

[0164] Formula (5):

[0165] (5)

[0166] in, and These represent the states at time t and time t+1, respectively. The relative entropy KL divergence algorithm is used to measure the loss of data utility caused by privacy protection policies. This refers to raw data, i.e., unprocessed data awaiting circulation. This represents the data to be processed and circulated after being processed by l.

[0167] It should be noted that the summation of g(r) at different time points satisfies the condition that the summation is less than a preset value. conditions .

[0168] Wherein, the preset value It was set up based on multiple experiments.

[0169] Step S64: Determine the target privacy protection strategy for the data subject based on the privacy protection objectives of data element circulation, so as to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0170] In the specific implementation step S64, the privacy protection goal of data element circulation, that is, each data audience selects an appropriate strategy to maximize the reward of the data element circulation process, is input into formula (6) for processing to find the target privacy protection strategy for each data subject, that is, the optimal privacy protection strategy.

[0171] Formula (6):

[0172] (6)

[0173] Among these, by achieving the aforementioned privacy protection goals, the optimal privacy protection strategy can be found for data subjects. That is, to maximize the long-term benefits in the process of data element circulation; ∈[0,1] is the discount factor.

[0174] The specific implementation of steps S63 and S64 can be performed in the trained deep Q-network.

[0175] It should be noted that the specific training process of a deep Q-network includes:

[0176] Step S71: Use the reinforcement learning algorithm Q-Learning to map the states and actions described in the MDP.

[0177] In the specific implementation step S71, the Q-learning reinforcement learning algorithm is used to learn the mapping relationship between the state and the action, i.e., the Q function, so that the expected reward of taking a certain action in a certain state can be predicted based on the Q function.

[0178] The Q function can be obtained based on the Bellman equation, as shown in formula (7) below:

[0179] (7)

[0180] in, This represents the mapping relationship between the state and action at the next time step. This represents the mapping relationship between the state and action at the current time step. This indicates the pre-set learning rate. This represents the mapping relationship between the current time step state and the optimal action under the influence of the discount factor. This indicates the initial privacy risks at the current time step.

[0181] Step S72: Initialize the main network and the target network.

[0182] In the specific implementation step S72, a multilayer perceptron with the same structure is constructed, that is, the main network is... and the target network is .

[0183] Step S73: Based on the mapping relationship between the state and the action, and the benefits of selecting the privacy protection strategy li under state Si.

[0184] Specifically, firstly, a privacy protection strategy is selected based on the ε-greedy principle, that is, at each time step, the privacy protection strategy with the highest benefit is selected with a probability of ε, and a random privacy protection strategy l is selected with a probability of (1-ε).

[0185] Among them, the ε-greedy strategy is a commonly used exploration strategy in reinforcement learning, which aims to balance exploration and exploitation.

[0186] In other words, if state Si chooses the privacy protection strategy li, its state Si will transform into state Si' in the corresponding circulation path. At this time, the corresponding reward can be determined as Re. i Then, the currently selected state, the privacy protection strategy li chosen in state Si, the new state Si', and the corresponding reward Re are calculated. i As a tuple, and the tuple <Si,li,Si’,Re i >Store into the replay unit; determine multiple tuples for each data audience at different time steps in the manner described above.

[0187] The replay unit is used to store and randomly sample multiple tuples from the past, thereby breaking the correlation between data and improving learning efficiency.

[0188] Step S74: Train a deep Q-network based on the main network, the target network, and the tuples to obtain the trained Q-network.

[0189] Specifically, based on the main network and target network In each training round, a certain number of round samples are randomly selected from the replay unit to update its policy, that is, gradient descent is performed by formula (8) to update the parameters in the main network, that is, policy update is performed.

[0190] Each sample is a tuple.

[0191] Formula (8):

[0192] (8)

[0193] The parameters θ of the target network are periodically copied and updated from the main network, which helps reduce instability during training.

[0194] It should be noted that after each training session, the judgment is... If the value is less than the preset value, then the target network obtained is the trained deep Q-network; otherwise, continue to the next training round.

[0195] Optionally, steps S63 and S64 can be performed based on the trained deep Q-network to obtain the optimal privacy protection strategy. .

[0196] In this embodiment of the invention, a privacy protection decision-making process for data subjects during data element circulation is described using a Methodological Discrete (MDP). The actions, states, and rewards within the MDP determine that each data audience selects an appropriate strategy to maximize the reward during data element circulation, thereby finding a target privacy protection strategy, i.e., the optimal privacy protection strategy, for each data subject. The potential privacy risks posed by each data audience to the data subject serve as the basis for controlling the data element circulation path, thus controlling the sharing of data to be circulated according to the target privacy decision. This eliminates the need to pre-set data access conditions, thereby achieving a balance between data subject privacy protection and data sharing.

[0197] Optionally, based on the data flow control method shown in the above embodiments of the present invention, a corresponding structural schematic diagram of a data flow control device is also shown in the embodiments of the present invention, such as... Figure 3 As shown, the device includes;

[0198] The determining unit 301 is used to acquire the data to be circulated and its corresponding data circulation scenario;

[0199] The privacy and security condition generation unit 302 is used to determine the information sensitivity, the data subject's trust value to the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario; and to determine the privacy and security conditions based on a preset set of privacy protection strategies, the information sensitivity, the data audience's privacy attitude towards the data subject, and the data subject's trust value towards the data audience.

[0200] The privacy risk prediction unit 303 is used to determine the privacy risk of each data audience based on the information sensitivity in the privacy security conditions, the data audience's attitude towards the data subject's privacy, and the data subject's risk preference value.

[0201] Processing unit 304 is used to determine the target privacy decision for each data audience based on the privacy risks and privacy security conditions of each data audience, so as to control the data to be circulated to be shared in accordance with the target privacy decision.

[0202] The specific principles and execution processes of each unit in the data flow control device disclosed in the above embodiments of the present invention are the same as the corresponding contents in the data flow control method provided in the above embodiments of the present invention. Please refer to the corresponding parts in the data flow control method disclosed in the above embodiments of the present invention, and they will not be repeated here.

[0203] In this embodiment of the invention, the privacy and security conditions of a data subject are characterized based on their privacy needs, resulting in the conditions required to satisfy those needs. Then, prospect theory is used to assess the privacy risks faced by the data subject within the data circulation context. Finally, based on the data subject's own privacy and security conditions and the privacy risks they face, a target non-sharing strategy is determined for each data audience, transforming the privacy protection problem into a utility maximization problem, thus obtaining control decisions for the data circulation path. This eliminates the need to pre-set data access conditions, thereby achieving a balance between data subject privacy protection and data sharing.

[0204] Optionally, based on the data circulation control device shown in the above embodiments of the present invention, the privacy and security condition generation unit 302, which determines the information sensitivity, the data subject's trust value in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario, is specifically used for:

[0205] The sensitivity of information is determined based on the data content in the data to be circulated and the data circulation scenario.

[0206] The data audience's attitude towards the privacy of the data subject is determined based on the data to be circulated;

[0207] The trust level of the data subject towards the data audience is determined based on the privacy attitude and the data to be circulated.

[0208] The determination of the data subject's trust level in the data audience based on the privacy attitude and the data to be circulated includes:

[0209] Set the initial trust value of the data subject to the data audience;

[0210] The trust impact function of the privacy threat is calculated based on the sensitivity of the information and the number of the data audience.

[0211] The trust value of the data subject to the data audience is determined based on the initial trust value and the trust impact function of the privacy threat.

[0212] Optionally, based on the data flow control device shown in the above embodiments of the present invention, the privacy risk prediction unit 303 is specifically used for:

[0213] The probability of the data audience forwarding the data subject's data is determined based on the privacy attitude.

[0214] The initial privacy risk is determined by calculating based on the sensitivity of the information and the probability that the data to be circulated will be forwarded.

[0215] The privacy risk of each data audience is determined based on the data subject's risk preference value and the initial privacy risk, wherein the risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario.

[0216] Optionally, based on the data flow control device shown in the above embodiments of the present invention, the processing unit 304 is specifically used for:

[0217] For each data audience, select a privacy protection strategy for that data audience from the preset privacy protection strategy set of the privacy and security conditions;

[0218] Based on the trust value of the data subject to the data audience in the privacy and security conditions, determine the benefits of the data audience to the data subject under each privacy policy;

[0219] For each data audience, the target privacy decision is determined based on the data subject's benefits under each privacy policy and the data audience's privacy risks.

[0220] Optionally, based on the data flow control device shown in the above embodiments of the present invention, the processing unit 304 is further configured to:

[0221] For each data subject, cumulative privacy risks are determined based on the data subject's historical privacy risks;

[0222] The corresponding historical benefits are determined based on the historical sharing utility value and historical privacy risks of the data subject under each privacy policy.

[0223] The accumulated privacy risks, historical privacy protection strategies, and historical benefits are used to determine the privacy protection objectives for the circulation of data elements.

[0224] Based on the privacy protection objectives of data element circulation, a target privacy protection strategy is determined for data subjects to control the sharing of the data to be circulated in accordance with the target privacy decision.

[0225] This application provides an electronic device, which includes a processor and a memory. The memory is used to store data flow control program code and data, and the processor is used to call the program instructions in the memory to execute the steps shown in the data flow control method in the above embodiments.

[0226] This invention provides a storage medium, which includes the electronic device provided in the above-described embodiments of this application. The electronic device is used to execute the data flow control method disclosed in the embodiments of this application.

[0227] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0228] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0229] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data flow control method, characterized in that, The method includes: Acquire the data to be circulated and its corresponding data circulation scenario, wherein the data to be circulated includes the data subject and its data content; The sensitivity of information is determined based on the data content in the data to be circulated and the data circulation scenario. The data audience's attitude towards the privacy of the data subject is determined based on the data to be circulated; Set the initial trust value of the data subject to the data audience; Calculate the trust impact function of privacy threats based on the sensitivity of the information and the number of the data audience; The trust value of the data subject to the data audience is determined based on the initial trust value and the trust impact function of the privacy threat. Privacy and security conditions are determined based on a set of preset privacy protection strategies, the sensitivity of the information, the data audience's attitude toward the data subject's privacy, and the data subject's trust in the data audience. The probability of the data audience forwarding the data subject's data is determined based on the privacy attitude. The initial privacy risk is determined by calculating based on the sensitivity of the information and the probability that the data to be circulated will be forwarded. The privacy risk of each data audience is determined based on the data subject's risk preference value and the initial privacy risk, wherein the risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario; For each data audience, a privacy protection strategy for that data audience is selected from the preset privacy protection strategy set of the privacy and security conditions; Based on the trust value of the data subject to the data audience in the privacy and security conditions, determine the benefits of the data audience to the data subject under each privacy policy; For each data audience, based on the data subject's benefits under each privacy policy and the data audience's privacy risks, a target privacy decision is determined for the data audience to control the sharing of the data to be circulated in accordance with the target privacy decision.

2. The method according to claim 1, characterized in that, Also includes: For each data subject, cumulative privacy risks are determined based on the data subject's historical privacy risks; The corresponding historical benefits are determined based on the historical sharing utility value and historical privacy risks of the data subject under each privacy policy. The accumulated privacy risks, historical privacy protection strategies, and historical benefits are used to determine the privacy protection objectives for the circulation of data elements. Based on the privacy protection objectives of data element circulation, a target privacy protection strategy is determined for data subjects to control the sharing of the data to be circulated in accordance with the target privacy decision.

3. A data flow control device, characterized in that, The device includes: A determining unit is used to acquire the data to be circulated and its corresponding data circulation scenario, wherein the data to be circulated includes the data subject and the data content; A privacy and security condition generation unit is used to determine the information sensitivity, the data subject's trust value in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario; and to determine privacy and security conditions based on a preset set of privacy protection strategies, the information sensitivity, the data audience's privacy attitude towards the data subject, and the data subject's trust value towards the data audience. A privacy risk prediction unit is used to determine the privacy risk of each data audience based on the sensitivity of the information in the privacy security conditions, the data audience's attitude towards the privacy of the data subject, and the data subject's risk preference value. The processing unit is configured to determine a target privacy decision for each data audience based on the privacy risks and privacy security conditions of each data audience, so as to control the sharing of the data to be circulated in accordance with the target privacy decision; The privacy and security condition generation unit, which determines the information sensitivity, the data subject's trust in the data audience, and the data audience's privacy attitude towards the data subject based on the data to be circulated and the data circulation scenario, is specifically used for: The sensitivity of information is determined based on the data content in the data to be circulated and the data circulation scenario. The data audience's attitude towards the privacy of the data subject is determined based on the data to be circulated; The trust level of the data subject to the data audience is determined based on the privacy attitude and the data to be circulated. Determining the data subject's trust level in the data audience based on the privacy attitude and the data to be circulated includes: Set the initial trust value of the data subject to the data audience; Calculate the trust impact function of privacy threats based on the sensitivity of the information and the number of the data audience; The trust value of the data subject to the data audience is determined based on the initial trust value and the trust impact function of the privacy threat. The privacy risk prediction unit is specifically used for: The probability of the data audience forwarding the data subject's data is determined based on the privacy attitude. The initial privacy risk is determined by calculating based on the sensitivity of the information and the probability that the data to be circulated will be forwarded. The privacy risk of each data audience is determined based on the data subject's risk preference value and the initial privacy risk, wherein the risk preference value refers to the risk preference value selected by each data audience in the data circulation scenario; The processing unit is specifically used for: For each data audience, select a privacy protection strategy for that data audience from the preset privacy protection strategy set of the privacy and security conditions; Based on the trust value of the data subject to the data audience in the privacy and security conditions, determine the benefits of the data audience to the data subject under each privacy policy; For each data audience, the target privacy decision is determined based on the data subject's benefits under each privacy policy and the data audience's privacy risks.

4. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory being used to store program code and data for data generation, and the processor being used to call program instructions in the memory to execute the data flow control method as described in any one of claims 1-2.

5. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the data flow control method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Data security risk assessment method and system based on privacy calculation

    CN118940291A

  • Information circulation method, device and system

    WO2020087876A1