A Privacy Protection Method for Data Sharing
By dynamically evaluating the sensitive type of data to be accessed during data sharing and the risks of accessing users, combined with complaint information and geographical location information, the refined protection of highly sensitive data is achieved, and the problem of lagging response and lack of flexibility in traditional methods is solved.
Patent Information
- Application Number
- CN202411708495.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Traditional data sharing privacy protection methods are difficult to effectively respond to complex privacy protection needs, cannot dynamically evaluate user access risks, and fail to make full use of complaint information and geographical location information, resulting in a lagging response to privacy protection measures and lack of flexibility.
By obtaining the number of complaint information involving the data to be accessed in the complaint information, combining the data existence time and geographical location information, dynamically determine the sensitive type of data to be accessed, and assessing their access rights based on the risk of the access user, so as to achieve refined protection of highly sensitive data.
It improves the recognition rate of highly sensitive data, enhances the flexibility and accuracy of data privacy protection, dynamically adjusts user access rights, and reduces the possibility of data breaches.
Smart Images

Figure CN119203247B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of privacy protection for data sharing, and specifically to a privacy protection method for data sharing. Background Art
[0002] Carbon footprint data may contain sensitive information such as an enterprise's production process, energy usage, and supply chain management. If this information is leaked to competitors, it may weaken the enterprise's market position and competitive advantage. Therefore, protecting the privacy of carbon footprint data helps maintain these cooperative relationships and prevent trust crises caused by the leakage of sensitive information. Additionally, carbon footprint data may involve an enterprise's technological innovation and intellectual property rights. Protecting these data can prevent competitors from obtaining the enterprise's core technologies, thereby protecting the enterprise's innovation achievements. Therefore, a privacy protection method for carbon footprint data is particularly important.
[0003] The risk assessment of traditional data sharing privacy protection methods often relies on a single dimension. For example, access risks are determined only based on access frequency, user identity, or data sensitivity. This single evaluation model cannot effectively address complex privacy protection requirements because data leakage risks are usually the result of the superposition of multiple factors. The single-dimensional evaluation method is difficult to capture the specific risks of access behaviors. For example, the risks brought by considering the user's IP address, complaint records, and accessed data simultaneously. Therefore, the limitations of this evaluation method reduce the accuracy and effectiveness of privacy protection.
[0004] During the data sharing process, users' complaint information is an important feedback data source. Traditional methods rarely use complaint information to adjust access policies in privacy protection. A large number of complaint information for specific data may indicate privacy risks in the data, but traditional methods do not effectively utilize these feedbacks and fail to associate complaint information with sensitivity settings, user permission management, etc. This defect results in a relatively lagged adjustment of privacy protection measures, making it difficult to respond promptly to the risks feedback by users and lacking flexibility and responsiveness.
[0005] Traditional privacy protection methods for data sharing generally determine access permissions based on user identity or permission level and are difficult to dynamically evaluate users' access risks according to user behavior characteristics. For example, the number of complaints received by a user may change over time, affecting the judgment of their access risk. If the access risks of users are not dynamically updated according to these changes, the system may allow high-risk users to access sensitive data, increasing the likelihood of data leakage.
[0006] In the scenario of data sharing, the geographical location information of users can provide important basis for risk judgment. Traditional privacy protection methods for data sharing usually do not consider the impact of users' geographical locations on data access security, and only take access permissions as the core determining factor. The geographical location distribution helps to reveal the potential threats of data circulation paths and the spread of sensitive information. For example, if a large number of access requests come from a specific area, it may indicate abnormal cluster access behavior. Ignoring this information will lead to blind spots in security protection. At the same time, this defect may cause cross-regional privacy risks faced by data sharing to be ignored. Summary of the Invention
[0007] The present invention provides a privacy protection method for data sharing, which is used to facilitate the solution of the problems mentioned in the above background technology.
[0008] The present invention provides the following technical solutions: Optionally, a privacy protection method for data sharing, characterized in that it includes,
[0009] Denote the user applying for access to data as the accessing user, and denote the owner of the data applied for access by the accessing user as the holding user;
[0010] Denote the data applied for access by the accessing user as the data to be accessed;
[0011] Obtain the number of complaint information involving the data to be accessed among all complaint information;
[0012] The complaint information includes the data involved in the complaint and the time of the complaint;
[0013] Judge whether the data to be accessed is highly sensitive data or low-sensitive data according to the number of complaint information involving the data to be accessed;
[0014] Establish a first location set;
[0015] If the data to be accessed is low-sensitive data, obtain the IP address of the accessing user applying for access to the data to be accessed, obtain the geographical location and longitude and latitude coordinates corresponding to each IP address, form a pairing relationship between the IP address and the geographical location and longitude and latitude coordinates corresponding to it, and list all pairing relationships in the first location set;
[0016] Judge whether the data to be accessed has the risk of being transformed from low-sensitive data to high-sensitive data according to the pairing relationship in the first location set;
[0017] Obtain the access risk of the accessing user;
[0018] Combine the access risk of the accessing user and the type of the data to be accessed to judge whether the accessing user can access the data to be accessed.
[0019] Optionally, determine the sensitive type of the data to be accessed according to the number of complaint information related to the data to be accessed in the complaint information, specifically as follows:
[0020] Obtain the number of complaint information related to the data to be accessed in the complaint information, and denote the number as A;
[0021] Set the number threshold Z 1 ;
[0022] If A≥Z 1 , record the data to be accessed as highly sensitive data;
[0023] If A<Z 1 , obtain the existence time of the data to be accessed, denoted as T days;
[0024] Set the data existence time threshold T 1 , with the unit set as days, and set the ratio threshold Z 2 ;
[0025] If T≤T 1 , obtain the ratio of A to T, calculated as A÷T;
[0026] If (A÷T)≥Z 2 , record the data to be accessed as highly sensitive data;
[0027] If (A÷T)<Z 2 , record the data to be accessed as low-sensitive data;
[0028] If T>T 1 , obtain the moment when the data to be accessed is uploaded completely, denoted as the upload moment;
[0029] Denote the moment after the upload moment and at an interval of (2 / 3)T from the upload moment as the front-end moment, and denote the moment after the upload moment and at an interval of T from the upload moment as the back-end moment;
[0030] Denote the time period between the upload moment and the front-end moment as the front-end time period, and denote the time period between the front-end moment and the back-end moment as the back-end time period;
[0031] Obtain the number of complaint information related to the data to be accessed in the complaint information within the front-end time period, denoted as A 1 , obtain the number of complaint information related to the data to be accessed in the complaint information within the back-end time period, denoted as A - A 1 ;
[0032] If {A 1 ÷[(2 / 3)T]}<{(A - A 1 )÷[(1 / 3)T]}, record the data to be accessed as highly sensitive data;
[0033] If {A 1 ÷[(2 / 3)T]} ≥ {(A - A 1 ) ÷ [(1 / 3)T]}, mark the data to be accessed as low-sensitivity data.
[0034] Optionally, determine whether the data to be accessed has the risk of being converted from low-sensitivity data to high-sensitivity data according to the pairing relationship within the first location set, specifically:
[0035] Obtain the number of pairing relationships within the first location set, denoted as D;
[0036] Mark the area corresponding to each geographical location in the D pairing relationships within the first location set on the map;
[0037] Set the threshold Z 3 ;
[0038] Obtain the number of times each area is marked;
[0039] For each marked area, if the ratio of the number of times the area is marked to D ≥ Z 3 , then mark the area as a risk area;
[0040] If the ratio of the number of times the area is marked to D < Z 3 , then do not process;
[0041] Obtain the number of times the risk area is marked, denoted as M;
[0042] Mark the geographical locations corresponding to the M IP addresses in the risk area;
[0043] Obtain the coordinate values of the longitude and latitude corresponding to the M geographical locations marked in the risk area, denoted as (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M );
[0044] Respectively obtain the average values of the longitude values and latitude values, denoted as E and F, calculated as E = (E 1 + E 2 + E 3 ... E M ) ÷ M, F = (F 1 + F 2 + F 3 ... F M ) ÷ M;
[0045] Record the longitude and latitude coordinates (E, F) as the coordinates of the first center;
[0046] Obtain the distances from (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M ) to the coordinates of the first center, and record them as G 1 , G 1 , G 1 ... G M ;
[0047] Obtain the average value of G 1 , G 1 , G 1 ... G M , and record the average value as G AVG , and calculate it as G AVG = (G 1 + G 2 + G 3 ... G M ) ÷ M;
[0048] Draw a circle with the coordinates of the first center as the center and a radius of G AVG , and obtain the number of geographical locations corresponding to M longitude and latitude coordinates of (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M ) that are located within the area covered by the drawn circle, and record the number as G.
[0049] Optionally, the step of judging whether the data to be accessed has the risk of being converted from low-sensitivity data to high-sensitivity data according to the pairing relationship in the first location set further includes:
[0050] Set the threshold Z 4 for the proportion of the number of geographical locations, and set the threshold Z 5 for the proportion of the number of risk users;
[0051] If G ÷ M ≥ Z 4 , obtain whether there are risk users among the access users corresponding to the IP addresses within the area covered by the drawn circle. If there are risk users, obtain the number of risk users and record it as B;
[0052] The risk users are specifically:
[0053] If the cumulative number of complaints against a user is greater than 0, the user will be marked as a risky user;
[0054] If the ratio of B to G (B÷G) ≥ Z 5 , the area covered by the drawn circle will be marked as a potential risk area, and the data to be accessed has the risk of being transformed from low-sensitivity data to high-sensitivity data;
[0055] If the ratio of B to G (B÷G) < Z 5 , there is no risky user among the accessing users corresponding to the IP addresses within the area covered by the drawn circle or G÷M < Z 4 , then the data to be accessed does not have the risk of being transformed from low-sensitivity data to high-sensitivity data.
[0056] Optionally, the obtaining of the access risk of the accessing user is specifically as follows:
[0057] For the accessing user who has been complained against:
[0058] If the cumulative number of complaints against the accessing user = 1, the accessed account will be blocked for 30 days;
[0059] If the cumulative number of complaints against the accessing user = 2, the user's account will be blocked for 180 days;
[0060] If the cumulative number of complaints against the accessing user = 3, the user's account will be blocked for 365 days;
[0061] If the cumulative number of complaints against the accessing user = 4, the user's account will be permanently blocked;
[0062] Obtain the cumulative number of blocks of the accessing user. If the cumulative number of blocks of the user is greater than 0 but less than or equal to 2, the accessing user has a type-I access risk, and the type-I access risk value is denoted as P 1 ;
[0063] If the cumulative number of blocks of the accessing user is greater than 2, the accessing user has a type-I access risk, and the type-I access risk value is denoted as P 2 , P 1 < P 2 ;
[0064] If the cumulative number of blocks of the accessing user is equal to 0, the accessing user does not have a type-I access risk.
[0065] Optionally, the obtaining of the access risk of the accessing user further includes:
[0066] Set a time threshold T for the existence of the accessing user 2 ;
[0067] Obtain the existence time of the accessing user, denoted as C;
[0068] Obtain the moment when the access user completes registration, denoted as the registration moment;
[0069] Denote the moment before the registration moment and seven days apart from the registration moment as the pre-registration moment;
[0070] Denote the time period between the pre-registration moment and the registration moment as the pre-registration time period;
[0071] If the existence time C of the access user < T 2 , obtain the IP address of the access user, denoted as the target IP address;
[0072] Denote the users with the same IP address as the access user's IP address as a type of user;
[0073] If there is a type of user within the pre-registration time period and this type of user is banned, then the access user has a secondary access risk, and the secondary access risk value is denoted as Q;
[0074] If there is no type of user banned within the pre-registration time period, then the access user has no secondary access risk;
[0075] Set the distance threshold Z 5 ;
[0076] If the geographical location corresponding to the target IP address is in the potential risk area, then the access user has a tertiary access risk, and the tertiary access risk value is denoted as R 2 ;
[0077] If the geographical location corresponding to the target IP address is not in the potential risk area, and the distance between the geographical location corresponding to the target IP address and the first center coordinate ≤ Z 5 , then the access user has a tertiary access risk, and the tertiary access risk value is denoted as R 1 , R 1 < R 2 ;
[0078] If the distance between the geographical location corresponding to the target IP address and the first center coordinate > the distance threshold Z 5 , then the access user has no tertiary access risk.
[0079] Optionally, combining the access risk of the access user and the type of data to be accessed, determine whether the access user can access the data to be accessed, specifically:
[0080] Obtain the total access risk value of the access user, denoted as U;
[0081] The total access risk value is specifically the sum of the primary access risk value, the secondary access risk value, and the tertiary access risk value;
[0082] If the accessing user has a type I access risk and a type III access risk, and the type I access risk value is P 2 , and the type III access risk value is R 3 , then the total access risk value U of the accessing user = P 2 +R 3 ;
[0083] Set the threshold value Z for the total risk value 6 ;
[0084] If the total access risk value U of the accessing user ≥ Z 6 , mark the accessing user as a high-risk user;
[0085] If the total access risk value U of the accessing user < Z 6 , mark the accessing user as a low-risk user;
[0086] When the data to be accessed is low-sensitive data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed;
[0087] When the data to be accessed is low-sensitive data and the accessing user is a high-risk user, the accessing user can directly access the data to be accessed after obtaining the approval of the holding user;
[0088] When the data to be accessed is high-sensitive data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed after obtaining the approval of the holding user;
[0089] When the data to be accessed is high-sensitive data and the accessing user is a high-risk user, access is not allowed.
[0090] The present invention has the following beneficial effects:
[0091] 1. The privacy protection method for data sharing determines the sensitivity category of the data by the number of complaint messages related to the data to be accessed in the complaint messages; if the number of complaint messages related to the data to be accessed reaches the set threshold, it is marked as high-sensitive data; if the number of complaint messages related to the data to be accessed does not reach the threshold, the sensitive type is further determined by combining the data existence time and the number of complaint messages related to the data to be accessed; this solution can dynamically update the sensitive type of the data according to the complaint messages of the holding user, improve the recognition rate of high-sensitive data, and enhance the flexibility of data privacy protection.
[0092] 2. For the privacy protection method for data sharing, when determining the sensitive type of the data to be accessed, if the existence time of the data to be accessed is short, obtain the ratio of the number of complaint messages to the existence time. If the ratio is greater than the set threshold, mark the data to be accessed as highly sensitive data; if the ratio is less than the set threshold, mark the data to be accessed as low-sensitive data. If the existence time of the data to be accessed is long, divide the upload time into two different stages, respectively count the number of complaint messages in the two different stages, and conduct a comparative analysis. This method can identify whether the number of complaint messages related to the data to be accessed surges in a short period of time. If there is a surge, mark the data to be accessed as highly sensitive data; if there is no surge, mark the data to be accessed as low-sensitive data. This solution avoids making a single judgment on the data throughout its existence time, thereby making a refined judgment on the data sensitivity type. This method significantly improves the accuracy of privacy protection and makes the data protection measures more in line with actual needs.
[0093] 3. For the privacy protection method for data sharing, for an access user with a short account existence time, obtain the IP address of the access user and the ban status of the account under this IP address. If an account has been banned under this IP address recently, it may be that the access user of the banned account registered a new account again after being banned for access. Therefore, the access risk is high. Therefore, once a banned account is found, the newly registered account can be immediately monitored closely to detect and prevent malicious behaviors in a timely manner. By checking the ban records under the IP address, this repeated attack behavior can be effectively prevented and the data security can be protected.
[0094] 4. For the privacy protection method for data sharing, by counting the number of times each area in the first location set is marked, define the area where the number of marked times exceeds the proportion threshold as a risk area, and continue to mark the address locations of each access user within the risk area to determine whether there is a concentrated distribution of users within the risk area. If there is a concentrated distribution of users, obtain whether there are risk users in the area where the users are concentrated. If there are risk users and the proportion of risk users is high, mark the area where the users are concentrated as a potential risk area, then the data to be accessed has the risk of being converted from low-sensitive data to high-sensitive data. This solution can identify and lock the potential risk area. If the IP address of the access user is located in the potential risk area, the access risk of the access user is increased. This method makes the privacy protection targeted and can perform more strict access control on users in a specific area.
[0095] 5. For the privacy protection method for data sharing, the risk value of an accessing user is evaluated based on multiple factors, including the number of complaints and IP addresses, etc., which can dynamically evaluate the access risk of the user, and then calculate the total risk value of the accessing user, and determine the access permission of the accessing user in combination with the sensitive type of the data. If the user risk value exceeds the set threshold, the system will regard it as a high-risk user and thus restrict its access permission. This solution is based on the sensitive type of the data and the total user risk value to dynamically control the access permission of the user, improving the accuracy and security of access control. Compared with a single permission control method, this method can adjust the user permission according to the access risk value, effectively reducing the possibility of the data being maliciously accessed or misused. In this way, the solution achieves hierarchical and refined privacy protection, making data sharing more secure. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 It is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0098] Embodiment 1. A privacy protection method for data sharing, characterized by including:
[0099] For a data sharing platform for carbon footprint disclosure data, all the shared data are carbon footprint disclosure data. After each enterprise uploads its carbon footprint disclosure information to the sharing platform, the user who applies to access the data is recorded as the accessing user, and the owner of the data applied to be accessed by the accessing user is recorded as the holding user, and the holding users are each enterprise.
[0100] The data applied to be accessed by the accessing user is recorded as the data to be accessed.
[0101] Obtain the number of complaint information involving the data to be accessed in all the complaint information. When the accessing user leaks the data or makes improper operations, such as when the shared disclosed supply chain carbon footprint data is misused, the holding user, that is, the enterprise, can file a complaint with the data sharing platform to ensure that the data it holds will no longer be leaked and further spread. After the holding user files a complaint, the accessed user who is complained about will be banned, and the banned accessed user cannot apply to access the data within the ban time.
[0102] The complaint information includes the data involved in the complaint and the time of the complaint.
[0103] Determine whether the data to be accessed is highly sensitive data or low - sensitive data based on the number of complaint information involving the data to be accessed in the complaint information;
[0104] Establish a first - location set;
[0105] If the data to be accessed is low - sensitive data, obtain the IP address of the access user who applies to access the data to be accessed, obtain the geographical location and longitude - latitude coordinates corresponding to each IP address, form a pairing relationship between the IP address and the geographical location and longitude - latitude coordinates corresponding to it, and list all the pairing relationships in the first - location set;
[0106] Based on the pairing relationship in the first - location set, determine whether the data to be accessed has the risk of being transformed from low - sensitive data to highly sensitive data;
[0107] Obtain the access risk of the access user;
[0108] Combine the access risk of the access user and the type of the data to be accessed to determine whether the access user can access the data to be accessed.
[0109] The determination of the sensitive type of the data to be accessed according to the number of complaint information involving the data to be accessed in the complaint information is specifically as follows:
[0110] Obtain the number of complaint information involving the data to be accessed in the complaint information, and record the number as A;
[0111] Set a number threshold Z 1 , Z 1 The role of Z is to judge the sensitive type of the data to be accessed. If the number of complaint information involving the data to be accessed is greater than or equal to the number threshold Z 1 , it indicates that the risk of the data to be accessed being leaked is relatively high, then mark the data to be accessed as highly sensitive data. If the number of complaint information involving the data to be accessed is less than the number threshold Z 1 , it indicates that the risk of the data to be accessed being leaked is relatively low, but it is necessary to further judge the sensitive type of the data to be accessed;
[0112] If A≥Z 1 , it indicates that the risk of the data to be accessed being leaked is relatively high, and record the data to be accessed as highly sensitive data;
[0113] If A<Z 1 , it indicates that the risk of the data to be accessed being leaked is relatively low, obtain the existence time of the data to be accessed, and record it as T days;
[0114] Set a data existence - time threshold T 1 , when the data to be accessed is marked as low - sensitive data according to the number of complaint information involving the data to be accessed in the complaint information, it is necessary to compare its existence time with the existence - time threshold T1 The size relationship is further used to determine the sensitive type of the data to be accessed. The unit is set to days, and a ratio threshold Z is set. 2 When the existence time of the data to be accessed is less than the existence time threshold T 1 , the sensitive type of the data to be accessed is determined according to the ratio of the number of complaint information related to the data to be accessed in the complaint information to the registration time. When the ratio is greater than or equal to the ratio threshold Z 2 It indicates that the number of complaint information related to the data to be accessed in the complaint information is less than the number threshold Z 1 The reason is that the existence time of the data to be accessed is short. Therefore, the data to be accessed is marked as highly sensitive data. When the ratio is less than the ratio threshold Z 2 It indicates that the number of complaint information related to the data to be accessed in the complaint information is less than the number threshold Z 1 The reason is that the number of complaint information related to the data to be accessed in the complaint information is small. Therefore, the data to be accessed remains marked as low-sensitive data;
[0115] If T ≤ T 1 , obtain the ratio of A to T, calculated as A ÷ T;
[0116] If (A ÷ T) ≥ Z 2 It indicates that the number of complaint information related to the data to be accessed in the complaint information is less than the number threshold Z 1 The reason is that the existence time of the data to be accessed is short, and the data to be accessed is recorded as highly sensitive data;
[0117] If (A ÷ T) < Z 2 It indicates that the number of complaint information related to the data to be accessed in the complaint information is less than the number threshold Z 1 The reason is that the sensitivity of the data to be accessed is low, and the data to be accessed is recorded as low-sensitive data;
[0118] If T > T 1 It indicates that the existence time of the data to be accessed is long. Therefore, it is necessary to determine whether there is a sharp increase in the number of complaint information related to the data to be accessed. Obtain the moment when the data to be accessed is uploaded and recorded as the upload moment;
[0119] The moment after the upload moment and at an interval of (2 / 3)T from the upload moment is recorded as the front-end moment, and the moment after the upload moment and at an interval of T from the upload moment is recorded as the back-end moment;
[0120] The time period between the upload moment and the front-end moment is recorded as the front-end time period, and the time period between the front-end moment and the back-end moment is recorded as the back-end time period;
[0121] Obtain the number of complaint information related to the data to be accessed in the complaint information within the front-end time period and record it as A1 Obtain the number of complaint information involving the data to be accessed in the complaint information within the backend time period, denoted as A - A 1 ;
[0122] If {A 1 ÷[(2 / 3)T]} < {(A - A 1 )÷[(1 / 3)T]}, it indicates that the number of complaint information involving the data to be accessed has increased significantly recently, and the data to be accessed is denoted as highly sensitive data;
[0123] If {A 1 ÷[(2 / 3)T]} ≥ {(A - A 1 )÷[(1 / 3)T]}, it indicates that the number of complaint information involving the data to be accessed has not increased significantly recently, and the data to be accessed is denoted as low - sensitive data.
[0124] Determine whether the data to be accessed has the risk of being transformed from low - sensitive data to high - sensitive data according to the pairing relationship within the first - location set, specifically:
[0125] Obtain the number of pairing relationships within the first - location set, denoted as D;
[0126] Mark the regions corresponding to each geographical location in the D pairing relationships within the first - location set on the map;
[0127] Set the marking - times proportion threshold Z 3 , by counting the number of times each region within the first - location set is marked and comparing it with the marking - times proportion threshold Z 3 as the condition for judging whether the region is a risk region. When the number of marked times exceeds the proportion threshold, it indicates that the region has the condition to become a risk region. When the number of marked times does not exceed the proportion threshold, it indicates that the region does not have the condition to become a risk region;
[0128] Obtain the number of times each region is marked;
[0129] For each marked region, if the ratio of the number of times the region is marked to D ≥ Z 3 , it indicates that the region has the condition to become a risk region, then mark the region as a risk region;
[0130] If the ratio of the number of times the region is marked to D < Z 3 , it indicates that the region does not have the condition to become a risk region, then do not process it;
[0131] Obtain the number of times the risk region is marked, and denote the number as M;
[0132] Mark the geographical locations corresponding to M IP addresses in the risk area respectively, continue to mark the address locations of each accessing user within the risk area, and determine whether there is a concentrated distribution of users within the risk area;
[0133] Obtain the coordinate values of the longitude and latitude corresponding to the M geographical locations marked within the risk area, and denote them as (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M ). Perform algebraic processing based on the longitude and latitude coordinates corresponding to each IP address to determine whether there is a concentrated distribution phenomenon among the accessing users corresponding to the M IP addresses;
[0134] Obtain the average values of the longitude values and latitude values respectively, and denote them as E and F. Calculate as E = (E 1 + E 2 + E 3 ... E M ) ÷ M, F = (F 1 + F 2 + F 3 ... F M ) ÷ M;
[0135] Record the longitude and latitude coordinates (E, F) as the first center coordinates, and the first center coordinates are the geometric center of the M longitude and latitude coordinates;
[0136] Obtain the distances from (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M ) to the first center coordinates, and denote them as G 1 , G 1 , G 1 ... G M ;
[0137] Obtain the average value of G 1 , G 1 , G 1 ... G M , and denote the average value as G AVG . Calculate as G AVG = (G 1 + G 2 + G 3 ... G M ) ÷ M;
[0138] Draw a circle with the center coordinate of the first center as the center and a radius of G AVG , and obtain the number of geographical locations corresponding to the M longitude and latitude coordinates of (E 1 , F 1 ), (E 2 , F 2 ), (E 3 , F 3 )... (E M , F M ) that are located within the area covered by the drawn circle. The number is denoted as G.
[0139] Judging whether the data to be accessed has the risk of being transformed from low-sensitivity data to high-sensitivity data according to the pairing relationship in the first location set further includes:
[0140] Set the threshold Z for the proportion of the number of geographical locations 4 , and the function of Z 4 is to judge whether there is a concentrated distribution of users in the risk area. When the proportion of the number of geographical locations within the area covered by the drawn circle is greater than or equal to the threshold Z for the proportion of the number of geographical locations 4 , it indicates that there is a concentrated distribution situation. Therefore, it is necessary to further judge whether the number of risk users within the area covered by the drawn circle will cause the risk of the data to be accessed being transformed from low-sensitivity data to high-sensitivity data. Set the threshold Z for the proportion of the number of risk users 5 , and the function of Z 5 is to judge whether the number of risk users within the area covered by the drawn circle will cause the risk of the data to be accessed being transformed from low-sensitivity data to high-sensitivity data. When the proportion of risk users within the area covered by the drawn circle is greater than or equal to the threshold Z for the proportion of the number of risk users 5 , then the number of risk users within the area covered by the drawn circle will cause the risk of the data to be accessed being transformed from low-sensitivity data to high-sensitivity data. When the proportion of risk users within the area covered by the drawn circle is less than the threshold Z for the proportion of the number of risk users 5 , then the number of risk users within the area covered by the drawn circle will not cause the risk of the data to be accessed being transformed from low-sensitivity data to high-sensitivity data;;
[0141] If G÷M≥Z 4 , it indicates that there is a concentrated distribution of users in the risk area, and it is necessary to further judge whether the number of risk users within the area covered by the drawn circle will cause the risk of the data to be accessed being transformed from low-sensitivity data to high-sensitivity data. Obtain whether there are risk users among the access users corresponding to the IP addresses within the area covered by the drawn circle. If there are risk users, obtain the number of risk users, denoted as B;
[0142] The specific risk users are:
[0143] If the cumulative number of complaints against a user is greater than 0, the user is marked as a risky user;
[0144] If the ratio of B to G (B÷G) ≥ Z 5 , it indicates that the proportion of risky users in the area covered by the circle drawn is relatively large, which will cause the risk of converting the data to be accessed from low-sensitivity data to high-sensitivity data. The area covered by the circle drawn is marked as a potential risk area, and the data to be accessed has the risk of being converted from low-sensitivity data to high-sensitivity data;
[0145] If the ratio of B to G (B÷G) < Z 5 and there are no risky users among the accessing users corresponding to the IP addresses in the area covered by the circle drawn or G÷M < Z 4 , then the data to be accessed does not have the risk of being converted from low-sensitivity data to high-sensitivity data and no processing is performed; By counting the number of times each area in the first position set is marked, the area where the number of marks exceeds the proportion threshold is defined as a risk area, and the address locations of each accessing user are further marked within the risk area. It is judged whether there is a concentrated distribution of users in the risk area. If there is a concentrated distribution of users, then it is obtained whether there are risky users in the area where the users are concentrated. If there are risky users and the proportion of risky users is relatively high, then the area where the users are concentrated is marked as a potential risk area, and the data to be accessed has the risk of being converted from low-sensitivity data to high-sensitivity data; This solution can identify and lock the potential risk area. If the IP address of the accessing user is located in the potential risk area, the access risk of the accessing user is increased; This method makes privacy protection targeted and can perform more strict access control on users in specific areas.
[0146] The obtaining of the access risk of the accessing user is specifically as follows:
[0147] For the accessing user who has been complained:
[0148] If the cumulative number of complaints against the accessing user = 1, the accessed account is blocked for 30 days;
[0149] If the cumulative number of complaints against the accessing user = 2, the user's account is blocked for 180 days;
[0150] If the cumulative number of complaints against the accessing user = 3, the user's account is blocked for 365 days;
[0151] If the cumulative number of complaints against the accessing user = 4, the user's account is permanently blocked;
[0152] Obtain the cumulative number of bans of the accessing user. If the cumulative number of bans of the accessing user is greater than 0 but less than or equal to 2, then the accessing user has a type-I access risk, and the type-I access risk value is denoted as P 1 ;
[0153] If the cumulative number of bans of the accessing user is greater than 2, then the accessing user has a type-I access risk, and the type-I access risk value is denoted as P 2 , P 1 <P 2 ;
[0154] If the cumulative number of bans of the accessing user is equal to 0, then the accessing user does not have a type-I access risk.
[0155] The obtaining of the access risk of the accessing user further includes:
[0156] Set a time threshold T for the existence of the accessing user 2 ;
[0157] Obtain the existence time of the accessing user, denoted as C;
[0158] Obtain the moment when the accessing user completes registration, denoted as the registration moment;
[0159] Denote the moment seven days before the registration moment and at an interval of seven days from the registration moment as the pre-registration moment;
[0160] Denote the time period between the pre-registration moment and the registration moment as the pre-registration time period;
[0161] If the existence time C of the accessing user < T 2 , it indicates that the existence time of the accessing user's account is short. Obtain the IP address of the accessing user, denoted as the target IP address, and the access situation of the supply chain carbon footprint disclosure data shared through the IP address can be traced;
[0162] Denote the users with the same IP address as the accessing user as type-I users;
[0163] If there are type-I users within the pre-registration time period and such type-I users are banned, then the accessing user has a type-II access risk, and the type-II access risk value is denoted as Q; For accessing users with a short account existence time, obtain the IP address of the accessing user, and obtain the ban status of the account under this IP address. If there is an account banned under this IP address recently, it may be that the accessing user of the banned account registers a new account again after being banned for access. Therefore, the access risk is relatively high. Therefore, once a banned account is found, the newly registered account can be immediately monitored closely to discover and prevent malicious behaviors in a timely manner; By checking the ban records under the IP address, this kind of repeated attack behavior can be effectively prevented and the security of the data can be protected;
[0164] If there is no user of a certain type banned during the registration front-end time period, then the accessing user has no secondary access risk;
[0165] Set the distance threshold Z 5 , Z 5 The function of Z is to determine whether the distance between the accessing user and the first center coordinate will cause access risk to the accessing user. When the distance between the accessing user and the first center coordinate is less than the radius of the potential risk area, it indicates that the accessing user is within the potential risk area and has a relatively high access risk. When the distance between the accessing user and the first center coordinate is greater than the radius of the potential risk area but less than Z 5 it indicates that the accessing user is near the potential risk area and has a relatively low access risk. When the distance between the accessing user and the first center coordinate is greater than Z 5 it indicates that the accessing user is not near the potential risk area, so there is no access risk;
[0166] If the geographical location corresponding to the target IP address is within the potential risk area, it indicates that the accessing user has a very high access risk, then the accessing user has a tertiary access risk, and the tertiary access risk value is denoted as R 2 ;
[0167] If the geographical location corresponding to the target IP address is not within the potential risk area and the distance between the geographical location corresponding to the target IP address and the first center coordinate ≤ Z 5 , it indicates that the accessing user is near the potential risk area and has a relatively high access risk, then the accessing user has a tertiary access risk, and the tertiary access risk value is denoted as R 1 , R 1 < R 2 ;
[0168] If the distance between the geographical location corresponding to the target IP address and the first center coordinate > the distance threshold Z 5 , then the accessing user has no tertiary access risk.
[0169] Combining the access risk of the accessing user and the type of data to be accessed, determine whether the accessing user can access the data to be accessed. Specifically:
[0170] Obtain the total access risk value of the accessing user, denoted as U;
[0171] The total access risk value is specifically the sum of the primary access risk value, secondary access risk value, and tertiary access risk value;
[0172] If the accessing user has a primary access risk and a tertiary access risk, and the primary access risk value is P 2 , and the tertiary access risk value is R 3 , then the total access risk value U of the accessing user = P 2 + R3 ;
[0173] Set the risk total value threshold Z 6 , Z 6 The function of Z is to judge the risk type of the accessing user. When the access risk total value of the accessing user is greater than or equal to the risk total value threshold Z 6 , it indicates that the access risk of the accessing user is relatively high. Therefore, the accessing user is recorded as a high-risk user; when the access risk total value of the accessing user is less than the risk total value threshold Z 6 , it indicates that the access risk of the accessing user is relatively low. Therefore, the accessing user is recorded as a low-risk user;
[0174] The access risk total value U of the accessing user ≥ Z 6 , it indicates that the access risk of the accessing user is relatively high, and the accessing user is recorded as a high-risk user;
[0175] The access risk total value U of the accessing user < Z 6 , it indicates that the access risk of the accessing user is relatively low, and the accessing user is recorded as a low-risk user;
[0176] When the data to be accessed is low-sensitive data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed;
[0177] When the data to be accessed is low-sensitive data and the accessing user is a high-risk user, the accessing user can directly access the data to be accessed after obtaining user approval;
[0178] When the data to be accessed is high-sensitive data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed after obtaining user approval;
[0179] When the data to be accessed is high-sensitive data and the accessing user is a high-risk user, access is not allowed; the risk value of the accessing user is evaluated based on multiple factors, including the number of complaints and IP addresses, etc. The access risk of the user can be dynamically evaluated, and then the risk total value of the accessing user can be calculated. And combined with the sensitive type of the data, the access permission of the accessing user is determined; if the user risk value exceeds the set threshold, the system will regard it as a high-risk user, thus restricting its access permission; this solution is based on the data sensitive type and the user risk total value, dynamically controls the access permission of the user, and improves the accuracy and security of access control; compared with a single permission control method, this method can adjust the user permission according to the access risk value, effectively reducing the possibility of malicious access or improper use of data; in this way, the solution achieves hierarchical and refined privacy protection, making data sharing more secure.
[0180] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0181] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A privacy protection method for data sharing, characterized by: include, The user who applies for access to the data is recorded as the access user, and the owner of the data that the access user applies to access is recorded as the holding user; Record the data requested by the accessing user as data to be accessed; Obtain the number of complaint information involving the data to be accessed among all complaint information; The complaint information includes the data involved in the complaint and the time of the complaint; Judging whether the data to be accessed is high-sensitivity data or low-sensitivity data according to the number of complaint information involving the data to be accessed in the complaint information; Establish position set number one; If the data to be accessed is low-sensitivity data, obtain the IP address of the user who applies to access the data to be accessed, obtain the geographical location and longitude and latitude coordinates corresponding to each IP address, form a pairing relationship between the IP address and the geographical location and longitude and latitude coordinates corresponding to the IP address, and include all pairing relationships in the first location set; Based on the pairing relationship in the first position set, determine whether the data to be accessed has a risk of being converted from low-sensitivity data to high-sensitivity data; Get the access risk of the access user; Based on the access risk of the accessing user and the type of the data to be accessed, determine whether the accessing user can access the data to be accessed.
2. A privacy protection method for data sharing according to claim 1, characterized in that: The method of determining the sensitive type of the data to be accessed according to the number of complaint information related to the data to be accessed in the complaint information is specifically: Obtain the number of complaint information involving the data to be accessed in the complaint information, the number is recorded as A; Set the number threshold Z1; If A ≥ Z1, the data to be accessed is recorded as highly sensitive data; If A<Z1, obtain the existence time of the data to be accessed, recorded as T days; Set the data existence time threshold T1, the unit is set to day, and set the ratio threshold Z2; If T≤T1, obtain the ratio of A to T and calculate it as A÷T; If (A÷T) ≥ Z2, the data to be accessed is recorded as highly sensitive data; If (A÷T) < Z2, the data to be accessed is recorded as low-sensitivity data; If T>T1, obtain the time when the data to be accessed is uploaded and record it as the upload time; The time after the upload time and with an interval of (2 / 3) T from the upload time is recorded as the front-end time, and the time after the upload time and with an interval of T from the upload time is recorded as the back-end time; The time period between the upload time and the front-end time is recorded as the front-end time period, and the time period between the front-end time and the back-end time is recorded as the back-end time period; The number of complaint information involving the data to be accessed in the complaint information in the front-end time period is obtained, recorded as A1, and the number of complaint information involving the data to be accessed in the complaint information in the back-end time period is obtained, recorded as A-A1; If {A1÷[(2 / 3)T]}<{(A-A1)÷[(1 / 3)T]}, the data to be accessed is recorded as highly sensitive data; If {A1÷[(2 / 3)T]}≥{(A-A1)÷[(1 / 3)T]}, the data to be accessed is recorded as low-sensitivity data.
3. A privacy protection method for data sharing according to claim 1, characterized in that: The step of judging whether the data to be accessed has a risk of being converted from low-sensitivity data to high-sensitivity data based on the pairing relationship in the first position set is as follows: Get the number of pairing relationships in the first position set, denoted as D; Mark the area corresponding to each geographical location in the D pairing relationships in the first location set on the map; Set the marking times percentage threshold Z3; Get the number of times each area is marked; For each marked area, if the ratio of the number of times the area is marked to D is ≥ Z3, the area is recorded as a risk area; If the ratio of the number of times a region is marked to D is less than Z3, no processing is performed; Get the number of times the risk area is marked, recorded as M; Mark the geographical locations corresponding to M IP addresses in the risk area; Obtain the coordinate values of the longitude and latitude corresponding to the M geographical locations marked in the risk area, which are recorded as (E1, F1), (E2, F2), (E3, F3)...(E M , F M ); Get the average of the longitude and latitude values, record them as E and F respectively, and calculate them as E=(E1+E2+E3……E M )÷M, F=(F1+F2+F3……F M )÷M; Record the longitude and latitude coordinates (E, F) as the center coordinates of point 1; Get (E1, F1), (E2, F2), (E3, F3)... (E M , F M ) to the center coordinate of No. 1, which is recorded as G1, G1, G1...G M ; Get G1, G1, G1...G M The average value is recorded as G AVG , calculated as G AVG =(G1+G2+G3……G M )÷M; Take the center coordinate of No. 1 as the center and make a circle with radius G AVG circle, get (E1, F1), (E2, F2), (E3, F3)... (E M , F M ) Among the geographical locations corresponding to the M longitude and latitude coordinates, the number of geographical locations located in the area covered by the circle is denoted by G.
4. A privacy protection method for data sharing according to claim 1, characterized in that: The step of judging whether the data to be accessed has a risk of being converted from low-sensitivity data to high-sensitivity data according to the pairing relationship in the first position set further includes: Set the geographical location ratio threshold Z4 and the risk user ratio threshold Z5; If G÷M≥Z4, obtain whether there are risky users among the visiting users corresponding to the IP addresses in the circle coverage area. If there are risky users, obtain the number of risky users, recorded as B; The risky users are specifically: If the cumulative number of complaints against a user is greater than 0, the user will be marked as a risky user; If the ratio of B to G (B÷G) ≥ Z5, the area covered by the circle is recorded as a potential risk area, and the data to be accessed has the risk of being converted from low-sensitivity data to high-sensitivity data; If the ratio of B to G (B÷G) is less than Z5, there are no risky users among the accessing users corresponding to the IP addresses within the circle coverage area, or G÷M is less than Z4, then the data to be accessed does not have the risk of being converted from low-sensitivity data to high-sensitivity data.
5. The privacy protection method for data sharing according to claim 1, characterized in that: The access risk of obtaining the access user is specifically: For users who are complained about: If the cumulative number of complaints against the accessing user = 1, the accessing account will be banned for 30 days; If the cumulative number of complaints against the visiting user = 2, the user's account will be banned for 180 days; If the cumulative number of complaints against the visiting user = 3, the user's account will be banned for 365 days; If the cumulative number of complaints against the visiting user = 4, the user's account will be permanently banned; Get the cumulative number of bans for the access user. If the cumulative number of bans for the access user is greater than 0 but less than or equal to 2, the access user has a Class I access risk, and the Class I access risk value is recorded as P1. If the cumulative number of bans for the access user is greater than 2, the access user has a first-class access risk, and the first-class access risk value is recorded as P2, P1<P2; If the cumulative number of bans for the accessing user is equal to 0, then the accessing user does not have a Class I access risk.
6. A privacy protection method for data sharing according to claim 1, characterized in that: The obtaining of the access risk of the access user further includes: Set the access user existence time threshold T2; Get the existence time of the accessing user, recorded as C; Get the time when the access user completes registration, which is recorded as the registration time; The time before the registration time and seven days after the registration time is recorded as the registration front end time; The time period between the registration front-end time and the registration time is recorded as the registration front-end time period; If the access user's existence time C < T2, obtain the access user's IP address and record it as the target IP address; Users with the same IP address as the visiting user are recorded as one type of user; If there is a type of user in the registration front-end time period and this type of user is banned, the accessing user has a type 2 access risk, and the type 2 access risk value is recorded as Q; If there is no first-class user banned during the registration front-end time period, the accessing user does not have the second-class access risk; Set distance threshold Z5; If the geographical location corresponding to the target IP address is in a potential risk area, the access user has three types of access risks, and the three types of access risk values are recorded as R2; If the geographical location corresponding to the target IP address is not located in the potential risk area, and the distance between the geographical location corresponding to the target IP address and the center coordinate of No. 1 is ≤ Z5, then the access user has three types of access risks, and the three types of access risk values are recorded as R1, R1<R2; If the distance between the geographical location corresponding to the target IP address and the center coordinates of No. 1 is greater than the distance threshold Z5, then there is no third-level access risk for the accessing user.
7. A privacy protection method for data sharing according to claim 1, characterized in that: The determination of whether the accessing user can access the data to be accessed is based on the access risk of the accessing user and the type of the data to be accessed, specifically: Get the total access risk value of the access user, denoted as U; The total access risk value is specifically the sum of the first-category access risk value, the second-category access risk value, and the third-category access risk value; If the access user has one type of access risk and three types of access risks, and the value of the first type of access risk is P2, and the value of the third type of access risk is R3, then the total access risk value of the access user is U=P2+R3; Set the total risk value threshold Z6; The total access risk value of the access user is U≥Z6, and the access user is recorded as a high-risk user; The total access risk value of the access user U<Z6, and the access user is recorded as a low-risk user; When the data to be accessed is low-sensitivity data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed; When the data to be accessed is low-sensitivity data and the accessing user is a high-risk user, the accessing user can directly access the data to be accessed after the user approves it; When the data to be accessed is highly sensitive data and the accessing user is a low-risk user, the accessing user can directly access the data to be accessed after the user approves it; When the data to be accessed is highly sensitive data and the accessing user is a high-risk user, access will be denied.
Citation Information
Patent Citations
Novel information security access control system and method based on user risk assessment
CN112685711A
Access authority management method and equipment based on access behavior
CN116992411A