A data security agent system for a trusted data space

By dividing the data system in the trusted data space and calculating the privacy level sequence, combined with visitor behavior analysis, precise control of data access is achieved, solving the leakage risk and low efficiency problems caused by improper authority allocation in data sharing, and achieving a balance between data security and utilization.

CN120342776BActive Publication Date: 2025-10-10BEIJING ZHONGAN NEBULA SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510765237.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-10
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the data sharing process, existing trusted data spaces only allocate permissions based on visitor identity information, ignoring visitors' behavioral habits and actual access needs, leading to data leakage risks or low business efficiency.

Method used

Through the data classification module, data is divided into multiple data systems, privacy keywords are set and privacy scores are calculated. Combined with the visitor's basic identity information and historical access records, the access keywords and data association values ​​are determined to achieve precise control of data openness.

Benefits of technology

It achieves precise protection of data privacy during the data opening process, avoids leakage caused by excessive opening, and at the same time does not affect the value utilization of data, realizing the organic unity of data security and open sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342776B_ABST
    Figure CN120342776B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data security, in particular to a data security proxy system of a trusted data space, which comprises the following steps: dividing the trusted data space into multiple data systems, calculating a privacy score for each single data set in each data system, arranging the single data sets according to the privacy levels through the privacy scores to obtain a privacy level sequence, identifying the basic identity information of a target visitor, determining a reference visitor, analyzing the historical access records of the reference visitor to obtain the access keywords of the target visitor, calculating the data correlation value between the system keywords and the access keywords, comparing the data correlation value with the privacy level sequence, determining the open information and displaying the open information to the target visitor, so that the privacy risk of data is fully considered in the data opening process, and data leakage caused by excessive opening is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security technology, and in particular to a data security proxy system for a trusted data space. Background Art

[0002] The trusted data space is different from the traditional data circulation platform. Its essence is to provide a "controllable", "manageable" and "preventable" safe and trusted environment for the generation, circulation, sharing and use of data.

[0003] The prior art CN110266687A discloses a design method for an IoT security proxy data sharing module using blockchain technology, including integrating a blockchain network whose processing nodes act as proxy servers. When a user is a registered member of the network, the user can access the data after verification through the blockchain network. The proxy also re-encrypts the data by converting the policy set during the data sharing process. The blockchain network works in conjunction with the cloud server to ensure anti-collusion scheme. The present invention also implements fine-grained access control to the data. Experimental results show that proxy re-encryption increases latency, but the use of blockchain records all interactions between entities, eliminating dependence on trusted third parties.

[0004] However, in the process of data sharing, trusted data spaces only allocate permissions based on visitor identity information, ignoring the visitor's behavioral habits and actual access needs. On the one hand, this will grant visitors too many permissions, leading to the risk of data leakage; on the other hand, too many restrictions will affect the effective use of data and reduce business efficiency. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems in the background technology and to propose a data security proxy system for a trusted data space.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A data security proxy system for a trusted data space, comprising:

[0008] The data classification module is used to divide the data in the trusted data space into multiple data systems, and then subdivide the data in the data system according to data type to obtain individual data sets;

[0009] The level establishment module is used to set privacy keywords for individual data sets, obtain public cases corresponding to the privacy keywords, perform decision evaluation and calculation on the public cases, obtain the privacy score of each privacy keyword, and sort the privacy scores in a data system in ascending order to obtain the privacy level sequence of the data system;

[0010] A visitor identification module is configured to identify a target visitor of a trusted data space and obtain basic identity information of the target visitor.

[0011] A visitor analysis module is configured to determine a reference visitor based on the basic identity information, obtain a historical access record of the reference visitor, calculate an inertia access frequency in the historical access record, determine a commonly used information group, extract a keyword in the commonly used information group, and obtain an access keyword of the target visitor.

[0012] An integrated analysis module is configured to obtain a system keyword of a data system and calculate a data correlation value between the system keyword and the access keyword.

[0013] A security opening module is configured to compare the data correlation value with a privacy level sequence, determine opening information, and display information content corresponding to the opening information to the target visitor.

[0014] As a further scheme of the present application, the data system refers to data stored under different data systems, the single data set refers to a collection of a type of data in the data system, and the data system is composed of a plurality of single data sets.

[0015] As a further scheme of the present application, the method for obtaining the privacy score of the privacy keyword includes:

[0016] S1: arbitrarily selecting a data system as a designated analysis object, obtaining a single data set in the data system, identifying a data type of each single data set, and marking an information word corresponding to the data type as a privacy keyword of the single data set, wherein one single data set corresponds to one privacy keyword;

[0017] S2: setting a decision influencing factor, the decision influencing factor referring to a necessary related factor in privacy analysis, the decision influencing factor including information sensitivity, data publicity, and information correlation;

[0018] arbitrarily selecting one privacy keyword and marking it as a key word, and arbitrarily selecting one decision influencing factor and marking it as a target factor, first obtaining an open case of the key word, using natural semantic processing technology to perform semantic analysis on the open case of the key word, and obtaining opening influence information of the key word, wherein the open case refers to a case in which the key word is illegally disclosed to cause disturbance or harm to a privacy user;

[0019] the opening influence information of the key word is respectively evaluated according to the decision influencing factor, and a plurality of evaluation values are obtained;

[0020] All evaluation values ​​of a decision-making influencing factor are averaged, and the output value after the average processing is marked as the decision-making influencing value Ci of this decision-making influencing factor, where i represents different decision-making influencing values, and i=1, 2, and 3 represent the information sensitivity, data disclosure, and information relevance of the decision-making influencing factor respectively;

[0021] S3: Use the formula DF=C1×a1+C2×a2+C3×a3 to obtain the privacy score DF of the key vocabulary, where a1, a2, and a3 are proportional coefficients, and a1+a2+a3=1.

[0022] As a further solution of the present invention, a method for obtaining a privacy level sequence of a data system includes:

[0023] The privacy keywords of all individual data sets in the data system are taken as key words in turn, and the privacy score DF of each privacy keyword is calculated. Then, the privacy scores DF of all privacy keywords in a data system are obtained, and the privacy scores DF are arranged in ascending order. At the same time, the arranged combination is marked as a privacy level sequence.

[0024] As a further solution of the present invention, damage assessment refers to setting different score levels based on the impact caused by illegal disclosure, and each score level corresponds to an assessment value, wherein the score levels include level one, level two, level three, level four and level five, corresponding to assessment values ​​of 1 point, 2 points, 3 points, 4 points and 5 points respectively, wherein, for the degree of damage affected: level five > level four > level three > level two > level one.

[0025] As a further solution of the present invention, a method for obtaining the inertial access frequency includes:

[0026] SS1: Based on the target visitor's basic identity information, obtain other visitors with the same historical identity information in the trusted database and mark them as reference visitors. The same identity information refers to the same occupation range and the same occupation level.

[0027] Obtain historical visit records of reference visitors, identify and summarize the visit information in the historical visit records, that is, integrate the same visit information to obtain multiple groups of information data, where each group of information data corresponds to an information point;

[0028] SS2: Count the data access duration in each group of information data to obtain the inertia duration Tj, where j represents a different information data group and j∈[1, J], indicating that there are J information data groups in total;

[0029] The inertial access frequency FPj of each information data group is obtained using the formula FPj=Tj÷Tz, where Tz is the total visit duration of the reference visitor, that is, .

[0030] As a further solution of the present invention, a method for obtaining target visitor's access keywords includes:

[0031] Compare the inertial access frequency FPj with the frequency threshold Fy. If FPj < Fy, the corresponding information data group is marked as a limited-use data group. Otherwise, if FPj ≥ Fy, the corresponding information data group is marked as a common-use information group.

[0032] Extract all common information groups, use the TextRank algorithm to calculate the keywords in the common information groups, and mark the words with keyword scores exceeding X1 as the access keywords of the target visitors.

[0033] As a further solution of the present invention, a method for calculating the data association value includes:

[0034] Obtain all data systems in the trusted data space, obtain the subject name corresponding to each data system, and mark the subject name as the system keyword of the corresponding data system;

[0035] Using the cosine similarity algorithm, the target visitor's access keywords are respectively associated with the system keywords of each data system to obtain the data association value of each data system, where one data system corresponds to one data association value.

[0036] As a further embodiment of the present invention, a method for determining open information includes:

[0037] Obtain the target visitor's real-time access traces, identify the data system corresponding to the real-time access traces, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence;

[0038] Obtain the total number M of private keywords in the target sequence, and then use the formula W = M × GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence;

[0039] Then, the integer portion of the public range value W is taken and marked as the public positioning value. The target sequence is obtained again, starting with the position corresponding to the first private keyword in the target sequence. The private keywords in the public positioning value are marked as public information in order.

[0040] The security opening module then identifies the open information and displays the information content corresponding to the open information to the target visitor.

[0041] Compared with the existing technology, the advantages of the present invention are:

[0042] This invention divides the trusted data space into multiple data systems, calculates privacy scores for individual data sets in each data system, and uses the privacy scores to rank the individual data sets according to their privacy levels, resulting in a privacy level sequence. During the data openness process, based on this privacy level sequence, precise control of data access can be achieved, effectively balancing the relationship between data value utilization and privacy protection.

[0043] The present invention identifies the basic identity information of the target visitor, determines the reference visitor, analyzes its historical visit records, obtains the target visitor's access keywords, calculates the data association value between the system keyword and the access keyword, and compares it with the privacy level sequence to determine the open information and display it to the target visitor. In this way, during the data opening process, the privacy risk of the data is fully considered to avoid data leakage caused by excessive opening. At the same time, the value utilization of the data will not be affected by excessive restrictions, thus realizing the organic unity of data security and open sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0046] Reference Figure 1 ,A data security agent system for a trusted data space, including a data classification module, a level establishment module, a visitor identification module, a visitor analysis module, an integrated analysis module and a security opening module;

[0047] The data classification module is used to identify data information in the trusted data space and classify the data in the trusted data space according to the information type of the data information, thereby obtaining multiple data systems. Each data system is then divided into details to obtain multiple individual data sets. A data system is composed of multiple individual data sets. Thereafter, a one-way communication connection is established between the data classification module and the level establishment module.

[0048] It should be further explained that the data system in the trusted data space refers to the data stored in different data systems. For example, in the trusted data space of the medical system, the data system includes the electronic medical record system, laboratory information management system, medical insurance information system, and medical imaging storage system. A single data set refers to a collection of one type of data in the data system. For example, in the electronic medical record system, there are multiple single data sets such as patient name, symptoms, test results, diagnosis conclusion, and diagnostic process.

[0049] The level establishment module is used to obtain individual data sets in the data system and set a privacy level sequence for each individual data set in the data system. The setting method of the privacy level sequence includes:

[0050] S1: Arbitrarily select a data system as a designated analysis object. Taking this designated analysis object as an example, obtain individual data sets in this data system, identify the data type of each individual data set, and mark the information word corresponding to this data type as the privacy keyword of this individual data set. One individual data set corresponds to one privacy keyword. For example, for an individual data set whose data type is patient name, use the patient name as the privacy keyword of this individual data set.

[0051] S2: Setting decision-making influencing factors. Decision-making influencing factors refer to relevant factors necessary for privacy analysis. Specific decision-making influencing factors are set by those skilled in the art based on big data experience. In this embodiment, decision-making influencing factors include information sensitivity, data publicity, and information relevance.

[0052] Randomly select a privacy keyword and mark it as a key word. At the same time, randomly select a decision-making influencing factor and mark it as a target factor. First, obtain public cases of the key word. Then, use natural semantic processing technology to perform semantic analysis on the public cases of the key word to obtain the open impact information of the key word. Among them, public cases refer to cases where the key word is illegally disclosed and causes trouble or harm to privacy users. Furthermore, the public cases are obtained through big data network retrieval.

[0053] The open impact information of key words is evaluated for damage according to the decision-making influencing factors to obtain multiple evaluation values;

[0054] All evaluation values ​​in a decision-making influencing factor are averaged, and the output value after the average processing is marked as the decision-making influencing value Ci of this decision-making influencing factor, where i represents different decision-making influencing values, and i=1, 2, and 3 represent the information sensitivity, data disclosure, and information relevance in the decision-making influencing factor, respectively. That is, there is one decision-making influencing value corresponding to information sensitivity, one decision-making influencing value corresponding to data disclosure, and one decision-making influencing value corresponding to information relevance in the decision-making influencing factor;

[0055] Furthermore, the damage assessment refers to setting different score levels according to the impact caused by illegal disclosure, and each score level corresponds to an assessment value. In this embodiment, the score levels include level 1, level 2, level 3, level 4 and level 5, corresponding to assessment values ​​of 1 point, 2 points, 3 points, 4 points and 5 points respectively. Among them, the degree of damage is: level 5 > level 4 > level 3 > level 2 > level 1;

[0056] It should be further explained that the number of evaluation values ​​involved in the mean calculation needs to reach the minimum sample size, and the minimum sample size is the threshold. The specific value is set by those skilled in the art based on big data experience;

[0057] S3: Obtain the privacy score DF of the key word using the formula DF = C1 × a1 + C2 × a2 + C3 × a3, where a1, a2, and a3 are proportional coefficients, and a1 + a2 + a3 = 1. Furthermore, the specific values ​​of a1, a2, and a3 are obtained by those skilled in the art through big data calculations.

[0058] S4: Take the privacy keywords of all individual data sets in the data system as key words in turn, and calculate the privacy score DF for each privacy keyword according to the above method. Then, obtain the privacy scores DF of all privacy keywords in the data system, arrange the privacy scores DF in ascending order, and mark the arranged combination as a privacy level sequence;

[0059] After that, there is a one-way communication connection between the level establishment module and the security opening module;

[0060] The visitor identification module is used to detect and identify users accessing the trusted data space. When a user is identified, the corresponding user is marked as a target visitor. At the same time, the basic identity information of the target visitor is obtained and transmitted to the visitor analysis module. The basic identity information includes the target visitor's age, occupation range, and occupation level.

[0061] The visitor analysis module is used to obtain the basic identity information of the target visitor. Based on the basic identity information, it also obtains other visitors of the same type and their visit information, performs inertia analysis on the visit information, and determines the visit keywords of the target visitor. The specific methods for determining the visit keywords include:

[0062] SS1: Based on the target visitor's basic identity information, obtain other visitors with the same historical identity information in the trusted database and mark them as reference visitors. The same identity information refers to the same occupation range and the same occupation level.

[0063] Obtain historical visit records of reference visitors, identify and summarize the visit information in the historical visit records, that is, integrate the same visit information to obtain multiple groups of information data, where each group of information data corresponds to an information point;

[0064] SS2: Count the data access duration in each group of information data to obtain the inertia duration Tj, where j represents a different information data group and j∈[1, J], indicating that there are J information data groups in total;

[0065] The inertial access frequency FPj of each information data group is obtained using the formula FPj=Tj÷Tz, where Tz is the total visit duration of the reference visitor, that is, ;

[0066] SS3: Compare the inertial access frequency FPj with the frequency threshold Fy. If FPj < Fy, mark the corresponding information data group as a restricted data group. Conversely, if FPj ≥ Fy, mark the corresponding information data group as a common information group. The specific value of the frequency threshold Fy is obtained by those skilled in the art through big data calculation.

[0067] SS4: Extract all common information groups, calculate the keywords in the common information groups using the TextRank algorithm, and mark the words with a keyword score exceeding X1 as the target visitor's access keywords. The TextRank algorithm is a prior art and will not be described in detail here. X1 is the score threshold, and the specific value is obtained by those skilled in the art through big data calculations.

[0068] The visitor analysis module then transmits the target visitor's access keywords to the integrated analysis module;

[0069] The integrated analysis module is used to obtain the target visitor's access keywords, perform correlation processing on the access keywords, and determine the data association value. The specific method for determining the data association value includes:

[0070] Obtain all data systems in the trusted data space, obtain the subject name corresponding to each data system, and mark the subject name as the system keyword of the corresponding data system;

[0071] Using the cosine similarity algorithm, the target visitor's access keyword is processed with the system keyword of each data system for similarity association processing, and the data association value of each data system is obtained, wherein one data system corresponds to one data association value. Furthermore, the cosine similarity algorithm belongs to the prior art and will not be described in detail here.

[0072] The integrated analysis module then transmits the data association value to the security and openness module;

[0073] The security and openness module is used to compare the data association value with the privacy level sequence to determine the open information of a single data set in the data system. The specific method for determining the open information includes:

[0074] Obtain the target visitor's real-time access traces, identify the data system corresponding to the real-time access traces, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence;

[0075] Obtain the total number M of private keywords in the target sequence, and then use the formula W = M × GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence;

[0076] Then, the integer portion of the public range value W is taken and marked as the public positioning value. The target sequence is obtained again, starting with the position corresponding to the first private keyword in the target sequence. The private keywords in the public positioning value are marked as public information in order.

[0077] The security opening module then identifies the open information and displays the information content corresponding to the open information to the target visitor.

[0078] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A data security proxy system for a trusted data space, characterized in that: include: The data classification module is used to divide the data in the trusted data space into multiple data systems, and then subdivide the data in the data system according to data type to obtain individual data sets; The level establishment module is used to set privacy keywords for individual data sets, obtain public cases corresponding to the privacy keywords, perform decision evaluation and calculation on the public cases, obtain the privacy score of each privacy keyword, and sort the privacy scores in a data system in ascending order to obtain the privacy level sequence of the data system; Visitor identification module, used to identify target visitors in the trusted data space and obtain basic identity information of target visitors; The visitor analysis module is used to identify the reference visitor through basic identity information, obtain the historical visit records of the reference visitor, calculate the inertial visit frequency in the historical visit records, determine the commonly used information groups, and then extract the keywords in the commonly used information groups to obtain the access keywords of the target visitor; An integrated analysis module is used to obtain system keywords of the data system and calculate the data association value between the system keywords and the access keywords; The calculation method of data association value includes: Obtain all data systems in the trusted data space, obtain the subject name corresponding to each data system, and mark the subject name as the system keyword of the corresponding data system; Using the cosine similarity algorithm, the target visitor's access keywords are respectively associated with the system keywords of each data system to obtain the data association value of each data system, where one data system corresponds to one data association value; The security opening module is used to compare the data association value with the privacy level sequence, determine the open information, and display the information content corresponding to the open information to the target visitor.

2. A data security proxy system for a trusted data space according to claim 1, characterized in that: A data system refers to the data stored in different data systems. A single data set refers to a collection of one type of data in a data system, and a data system consists of multiple single data sets.

3. The data security proxy system of a trusted data space according to claim 1, characterized in that: Methods for obtaining privacy scores for privacy keywords include: S1: Arbitrarily select a data system as the designated analysis object, obtain individual data sets in this data system, identify the data type of each individual data set, and mark the information word corresponding to this data type as the privacy keyword of this individual data set, where one individual data set corresponds to one privacy keyword; S2: Set decision-making factors. Decision-making factors refer to relevant factors necessary for privacy analysis, including information sensitivity, data disclosure, and information relevance. Randomly select a privacy keyword and mark it as a key word. At the same time, randomly select a decision-making factor and mark it as a target factor. First, obtain public cases of the key word. Use natural semantic processing technology to perform semantic analysis on the public cases of the key word to obtain the open impact information of the key word. Among them, public cases refer to cases where the illegal disclosure of the key word causes trouble or harm to the privacy user. The open impact information of key words is evaluated for damage according to the decision-making influencing factors to obtain multiple evaluation values; All evaluation values ​​of a decision-making influencing factor are averaged, and the output value after the average processing is marked as the decision-making influencing value Ci of this decision-making influencing factor, where i represents different decision-making influencing values, and i=1, 2, and 3 represent the information sensitivity, data disclosure, and information relevance of the decision-making influencing factor respectively; S3: Use the formula DF=C1×a1+C2×a2+C3×a3 to obtain the privacy score DF of the key vocabulary, where a1, a2, and a3 are proportional coefficients, and a1+a2+a3=1.

4. A data security proxy system for a trusted data space according to claim 3, characterized in that: Methods for obtaining the privacy level sequence of a data system include: The privacy keywords of all individual data sets in the data system are taken as key words in turn, and the privacy score DF of each privacy keyword is calculated. Then, the privacy scores DF of all privacy keywords in a data system are obtained, and the privacy scores DF are arranged in ascending order. At the same time, the arranged combination is marked as a privacy level sequence.

5. The data security proxy system of a trusted data space according to claim 3, characterized in that: Harm assessment refers to setting different score levels based on the impact caused by illegal disclosure. Each score level corresponds to an assessment value. The score levels include Level 1, Level 2, Level 3, Level 4 and Level 5, corresponding to assessment values ​​of 1 point, 2 points, 3 points, 4 points and 5 points respectively. Among them, the degree of harm is: Level 5 > Level 4 > Level 3 > Level 2 > Level 1.

6. The data security proxy system of a trusted data space according to claim 1, characterized in that: Methods for obtaining inertial access frequency include: SS1: Based on the target visitor's basic identity information, obtain other visitors with the same historical identity information in the trusted database and mark them as reference visitors. The same identity information refers to the same occupation range and the same occupation level. Obtain historical visit records of reference visitors, identify and summarize the visit information in the historical visit records, that is, integrate the same visit information to obtain multiple groups of information data, where each group of information data corresponds to an information point; SS2: Count the data access duration in each group of information data to obtain the inertia duration Tj, where j represents a different information data group and j∈[1, J], indicating that there are J information data groups in total; The inertial access frequency FPj of each information data group is obtained using the formula FPj=Tj÷Tz, where Tz is the total visit duration of the reference visitor, that is, .

7. The data security proxy system of a trusted data space according to claim 6, characterized in that: Methods for obtaining target visitors' access keywords include: Compare the inertial access frequency FPj with the frequency threshold Fy. If FPj < Fy, the corresponding information data group is marked as a limited-use data group. Otherwise, if FPj ≥ Fy, the corresponding information data group is marked as a common-use information group. Extract all common information groups, use the TextRank algorithm to calculate the keywords in the common information groups, and mark the words with keyword scores exceeding X1 as the target visitor's access keywords, where X1 is the score threshold.

8. The data security proxy system of a trusted data space according to claim 1, characterized in that: Methods for determining open information include: Obtain the target visitor's real-time access traces, identify the data system corresponding to the real-time access traces, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence; Obtain the total number M of private keywords in the target sequence, and then use the formula W = M × GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence; Then, the integer portion of the public range value W is taken and marked as the public positioning value. The target sequence is obtained again, starting with the position corresponding to the first private keyword in the target sequence. The private keywords in the public positioning value are marked as public information in order. The security opening module then identifies the open information and displays the information content corresponding to the open information to the target visitor.

Citation Information

Patent Citations

  • Internet of Things security agent data sharing module design method adopting a block chain technology

    CN110266687A

  • Privacy data hierarchical protection method based on big data

    CN112100670A

  • System and method for determining risk of group privacy leakage

    CN113709090A