Data security agent system of trusted data space
By dividing the data system in a trusted data space and calculating the privacy hierarchy sequence, and combining with the analysis of visitor behavior, data open information is determined, and the problem of improper permission allocation in data sharing is solved, and the organic unity of data security and open sharing is achieved.
Patent Information
- Application Number
- CN202510765237.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-10
AI Technical Summary
During the data sharing process, trusted data space is only allocated based on the guest's identity information, ignoring the visitors' behavioral habits and actual access needs, resulting in excessive permissions leading to the risk of data leakage or excessive restrictions affecting data utilization efficiency.
The trusted data space is divided into multiple data systems, the privacy scores of the single data set of each data system are calculated, and data access is accurately controlled through the privacy level sequence. Combined with the visitor's basic identity information and historical access records, the data correlation value is calculated to determine the open information.
In the process of data opening, we can balance data value utilization and privacy protection, avoid data leakage, and improve data utilization efficiency.
Smart Images

Figure CN120342776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and in particular to a data security proxy system for a trusted data space. Background Art
[0002] A trusted data space is different from traditional data circulation platforms. Its essence lies in providing a'secure and trusted environment that is 'controllable', 'governable' and 'defendable' for the generation, circulation, sharing and use of data.
[0003] The prior art CN110266687A discloses a design method for an Internet of Things security proxy data sharing module using blockchain technology, including integrating a blockchain network, whose processing nodes act as proxy servers. When a user is a registered member of the network, they are verified through the blockchain network and can access data. The proxy also re-encrypts the data by converting the policy set during the process of sharing data. The blockchain network works in cooperation with a cloud server to ensure an anti-collusion scheme. The present invention also realizes fine-grained access control of data. Experimental results show that proxy re-encryption increases latency, but the use of blockchain records all interactions between entities and eliminates the dependence on a trusted third party.
[0004] However, during the process of data sharing in a trusted data space, only the visitor's identity information is used for permission allocation, ignoring the visitor's behavior habits and actual access needs. On the one hand, it will grant visitors too many permissions, leading to a risk of data leakage; on the other hand, excessive restrictions will affect the effective utilization of data and reduce business efficiency. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems in the background art, and a data security proxy system for a trusted data space is proposed.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A data security proxy system for a trusted data space, including: A data classification module, used to divide the data in the trusted data space into multiple data systems, and then further divide the data in the data system according to the data type to obtain single-item data sets; A level establishment module, used to set privacy keywords for the single-item data sets, obtain public cases corresponding to the privacy keywords, conduct decision evaluation and calculation on the public cases to obtain the privacy scores of each privacy keyword, and sort the privacy scores in a data system in ascending order to obtain the privacy level sequence of the data system; A visitor identification module, used to identify the target visitor in the trusted data space and obtain the basic identity information of the target visitor; A visitor analysis module, which is used to determine a reference visitor through basic identity information, obtain the historical access records of the reference visitor, calculate the inertial access frequency in the historical access records, determine the common information group at the same time, and then extract the keywords in the common information group to obtain the access keywords of the target visitor; An integrated analysis module, which is used to obtain the system keywords of the data system and calculate the data association value between the system keywords and the access keywords; A security and openness module, which is used to compare the data association value with the privacy level sequence to determine the open information, and display the information content corresponding to the open information to the target visitor.
[0007] As a further solution of the present invention, the data system refers to the data stored under different data systems. A single data set refers to a collection of a type of data in the data system, and a data system is composed of multiple single data sets.
[0008] As a further solution of the present invention, the method for obtaining the privacy score of privacy keywords includes: S1: Arbitrarily select a data system as the specified analysis object, obtain the single data sets in this data system, identify the data types of each single data set, and mark the information words corresponding to this data type as the privacy keywords of this single data set. Among them, a single data set corresponds to one privacy keyword; S2: Set decision influencing factors. The decision influencing factors refer to the necessary relevant factors in privacy analysis. The decision influencing factors include information sensitivity, data publicity, and information relevance; Arbitrarily select a privacy keyword and mark it as the key word. At the same time, arbitrarily select a decision influencing factor and mark it as the target factor. First, obtain the public cases of the key word, and use natural semantic processing technology to perform semantic analysis on the public cases of the key word to obtain the open influence information of the key word. Among them, the public case refers to the case that causes trouble or harm to the privacy user after the key word is illegally made public; Evaluate the harm of the open influence information of the key word according to the decision influencing factors respectively to obtain multiple evaluation values; Perform mean processing on all the evaluation values in a decision influencing factor, and mark the result value output after mean processing as the decision influence value Ci of this decision influencing factor. i represents different decision influence values, and i = 1, 2, 3, which are the information sensitivity, data publicity, and information relevance in the decision influencing factors respectively; S3: Use the formula DF = C1×a1 + C2×a2 + C3×a3 to obtain the privacy score DF of the key word, where a1, a2, and a3 are proportionality coefficients respectively, and a1 + a2 + a3 = 1.
[0009] As a further solution of the present invention, the method for obtaining the privacy level sequence of the data system includes: Take the privacy keywords of all individual data sets in the data system as key words in turn, and for each privacy keyword, obtain the privacy score DF. Then, obtain the privacy scores DF of all privacy keywords in a data system, arrange the privacy scores DF in ascending order, and mark the arranged combination as the privacy level sequence.
[0010] As a further solution of the present invention, the harm assessment refers to setting different score levels according to the impact caused by illegal disclosure, and each score level corresponds to an evaluation value. Among them, the score levels include level one, level two, level three, level four, and level five, corresponding to the evaluation values of 1 point, 2 points, 3 points, 4 points, and 5 points respectively. Among them, for the degree of harm of the impact: level five > level four > level three > level two > level one.
[0011] As a further solution of the present invention, the method for obtaining the inertial access frequency includes: SS1: Based on the basic identity information of the target visitor, obtain other visitors with the same identity information in the trusted database and mark them as reference visitors. Among them, the same identity information refers to the same occupation range and the same occupation level; Obtain the historical access records of the reference visitors, identify and summarize the access information in the historical access records, that is, integrate the same access information to obtain multiple groups of information data. Among them, each group of information data corresponds to an information point; SS2: Statistically analyze the data access duration in each group of information data to obtain the inertial duration Tj, where j represents different information data groups, and j ∈ [1, J], indicating that there are a total of J information data groups; Use the formula FPj = Tj ÷ Tz to obtain the inertial access frequency FPj of each information data group, where Tz is the total access duration of the reference visitor, that is .
[0012] As a further solution of the present invention, the method for obtaining the access keywords of the target visitor includes: Compare the inertial access frequency FPj with the frequency threshold Fy. If FPj < Fy, mark the corresponding information data group as a restricted use data group. Conversely, if FPj ≥ Fy, mark the corresponding information data group as a common information group; Extract all common information groups, use the TextRank algorithm to calculate the keywords in the common information groups, and mark the words with a keyword score exceeding X1 as the access keywords of the target visitor.
[0013] As a further solution of the present invention, the method for calculating the data association value includes: Obtain all data systems in the trusted data space, obtain the topic names corresponding to each data system, and mark the topic names as the system keywords of the corresponding data systems; Using the cosine similarity algorithm, perform similarity association processing on the access keywords of the target visitor with the system keywords of each data system respectively to obtain the data association value of each data system, where one data system corresponds to one data association value.
[0014] As a further solution of the present invention, the method for determining open information includes: Obtain the real-time access trace of the target visitor, identify the data system corresponding to the real-time access trace, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence; Obtain the total number M of privacy keywords in the target sequence, and then use the formula W = M×GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence; After that, take the integer part of the public range value W, and mark the taken integer part as the public positioning value. Obtain the target sequence again, starting from the position corresponding to the first privacy keyword in the target sequence, and in sequence, mark the privacy keywords within the public positioning value as open information; After that, the security open module identifies the open information and displays the information content corresponding to the open information to the target visitor.
[0015] Compared with the existing technology, the advantages of the present invention are: The present invention divides the trusted data space into multiple data systems, calculates the privacy scores for individual data sets in each data system, arranges the positions of the privacy levels for the individual data sets through the privacy scores to obtain the privacy level sequence. During the data opening process, based on this privacy level sequence, precise control of data access can be achieved, effectively balancing the relationship between data value utilization and privacy protection; The present invention identifies the basic identity information of the target visitor, determines the reference visitor, analyzes its historical access records to obtain the access keywords of the target visitor, calculates the data association value between the system keywords and the access keywords, and compares it with the privacy level sequence to determine the open information and display it to the target visitor, enabling full consideration of the privacy risks of the data during the data opening process, avoiding data leakage caused by over-opening, and at the same time, not affecting the data value utilization due to excessive restrictions, achieving the organic unity of data security and open sharing. Brief Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the system structure of the present invention. Detailed Embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0018] Referring to Figure 1 , a data security proxy system for a trusted data space, including a data classification module, a level establishment module, a visitor identification module, a visitor analysis module, an integrated analysis module, and a security opening module; Among them, the data classification module is used to identify the data information in the trusted data space, and at the same time, classify the data in the trusted data space according to the information type of the data information, so as to obtain multiple data systems. Then, each data system is further divided into details to obtain multiple single-item data sets. Among them, one data system is composed of multiple single-item data sets. After that, there is a one-way communication connection between the data classification module and the level establishment module; It should be further noted that the data system in the trusted data space refers to the data stored under different data systems. For example, in the trusted data space of the medical system, the data system includes an electronic medical record system, a laboratory information management system, a medical insurance information system, and a medical image storage system, etc. The single-item data set refers to the set of a certain type of data in the data system. For example, for the electronic medical record system, there are multiple single-item data sets such as patient name, symptom manifestation, examination result, diagnosis conclusion, and diagnosis process; The level establishment module is used to obtain the single-item data sets in the data system and set a privacy level sequence for the single-item data sets in each data system. Among them, the setting method of the privacy level sequence includes: S1: Arbitrarily select a data system as the specified analysis object. Taking this specified analysis object as an example, obtain the single-item data sets in this data system, identify the data type of each single-item data set, and mark the information word corresponding to this data type as the privacy keyword of this single-item data set. Among them, one single-item data set corresponds to one privacy keyword. For example, for the single-item data set with the data type of patient name, mark the patient name as the privacy keyword of this single-item data set; S2: Set decision influence factors. The decision influence factors refer to the necessary relevant factors in privacy analysis. The specific decision influence factors are set by those skilled in the art according to big data experience. In this embodiment, the decision influence factors include information sensitivity, data publicity, and information relevance; Arbitrarily select a privacy keyword and mark it as a key term. At the same time, arbitrarily select a decision - influencing factor and mark it as the target factor. First, obtain the public cases of the key term, and then use natural language processing technology to perform semantic analysis on the public cases of the key term to obtain the open - influence information of the key term. Among them, the public case refers to a case that causes trouble or harm to privacy users after the illegal public disclosure of the key term. Further, the acquisition method of the public case is obtained through big - data network retrieval; Conduct harm assessments on the open - influence information of the key term according to the decision - influencing factors respectively to obtain multiple evaluation values; Perform mean - value processing on all the evaluation values in a decision - influencing factor, and mark the result value output after the mean - value processing as the decision - influence value Ci of this decision - influencing factor. i represents different decision - influence values, and i = 1, 2, 3, corresponding to the information sensitivity, data publicity, and information relevance in the decision - influencing factor respectively. That is, there is a decision - influence value corresponding to the information sensitivity in the decision - influencing factor, a decision - influence value corresponding to the data publicity, and a decision - influence value corresponding to the information relevance; Further, the harm assessment means setting different score levels according to the impact caused by the illegal public disclosure. Each score level corresponds to an evaluation value. In this embodiment, the score levels include level one, level two, level three, level four, and level five, corresponding to the evaluation values of 1 point, 2 points, 3 points, 4 points, and 5 points respectively. Among them, for the degree of harm of the impact: level five > level four > level three > level two > level one; It should be further noted that the number of evaluation values participating in the mean - value calculation needs to reach the minimum sample size, and the minimum sample size is the threshold. The specific value is set by those skilled in the art according to big - data experience; S3: Use the formula DF = C1×a1 + C2×a2 + C3×a3 to obtain the privacy score DF of the key term, where a1, a2, and a3 are proportionality coefficients respectively, and a1 + a2 + a3 = 1. Further, the specific values of a1, a2, and a3 are obtained by those skilled in the art through big - data operations; S4: Take the privacy keywords of all single - data sets in the data system as key terms in turn, and calculate the privacy score DF of each privacy keyword according to the above method. Then obtain the privacy scores DF of all privacy keywords in a data system, and arrange the privacy scores DF in ascending order of position, and mark the arranged combination as the privacy - level sequence; After that, there is a one - way communication connection between the level - establishment module and the security - opening module; The visitor identification module is used to detect and identify the access users in the trusted data space. When the access user is identified, the corresponding user is marked as the target visitor. At the same time, the basic identity information of the target visitor is obtained and transmitted to the visitor analysis module. Among them, the basic identity information includes the age, occupation range and occupation level of the target visitor; The visitor analysis module is used to obtain the basic identity information of the target visitor. At the same time, according to the basic identity information, other visitors of the same type and access information are obtained, and the access information is analyzed for inertia to determine the access keywords of the target visitor. The specific method for determining the access keywords includes: SS1: Based on the basic identity information of the target visitor, other visitors with the same historical identity information in the trusted database are obtained and marked as reference visitors. Among them, the same identity information refers to the same occupation range and the same occupation level; Obtain the historical access records of the reference visitors, identify and summarize the access information in the historical access records, that is, integrate the same access information to obtain multiple groups of information data. Among them, each group of information data corresponds to an information point; SS2: Statistically analyze the data access duration in each group of information data to obtain the inertia duration Tj, where j represents different information data groups, and j ∈ [1, J], indicating that there are a total of J information data groups; Use the formula FPj = Tj ÷ Tz to obtain the inertia access frequency FPj of each information data group, where Tz is the total access duration of the reference visitor, that is ; SS3: Compare the inertia access frequency FPj with the frequency threshold Fy. If FPj < Fy, the corresponding information data group is marked as a limited-use data group. On the contrary, if FPj ≥ Fy, the corresponding information data group is marked as a common information group. Among them, the specific value of the frequency threshold Fy is obtained by those skilled in the art through big data operations; SS4: Extract all common information groups, use the TextRank algorithm to calculate the keywords in the common information groups, and mark the words with a keyword score exceeding X1 as the access keywords of the target visitor. Among them, the TextRank algorithm belongs to the prior art and will not be elaborated here. X1 is the score threshold, and the specific value is obtained by those skilled in the art through big data operations; After that, the visitor analysis module transmits the access keywords of the target visitor to the integrated analysis module; The integrated analysis module is used to obtain the access keywords of the target visitor and perform correlation processing on the access keywords to determine the data association value. The specific method for determining the data association value includes: Obtain all data systems in the trusted data space, obtain the corresponding theme names of each data system, and mark the theme names as the system keywords of the corresponding data systems; Using the cosine similarity algorithm, the access keywords of the target visitor are respectively subjected to similarity association processing with the system keywords of each data system to obtain the data association values of each data system. Among them, one data system corresponds to one data association value. Further, the cosine similarity algorithm belongs to the prior art and will not be elaborated here; After that, the integrated analysis module transmits the data association value to the security and openness module; The security and openness module is used to compare the data association value with the privacy level sequence to determine the openness information of the individual data sets in the data system. The specific method for determining the openness information includes: Obtain the real-time access trace of the target visitor, identify the data system corresponding to the real-time access trace, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence; Obtain the total number M of privacy keywords in the target sequence. Then, use the formula W = M × GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence; After that, take the integer part of the public range value W and mark the taken integer part as the public positioning value. Obtain the target sequence again. Starting from the position corresponding to the first privacy keyword in the target sequence, and in sequence, mark the privacy keywords within the public positioning value as the openness information; After that, the security and openness module identifies the openness information and displays the information content corresponding to the openness information to the target visitor.
[0019] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A data security proxy system for a trusted data space, characterized in that Including: A data classification module, which is used to classify the data in the trusted data space into multiple data systems, and then further subdivide the data in the data system according to the data type to obtain individual data sets; A level establishment module, which is used to set privacy keywords for the individual data sets, obtain the public cases corresponding to the privacy keywords, conduct decision-making evaluation and calculation on the public cases to obtain the privacy scores of each privacy keyword, and sort the privacy scores in a data system in ascending order to obtain the privacy level sequence of the data system; A visitor identification module, which is used to identify the target visitor in the trusted data space and obtain the basic identity information of the target visitor; A visitor analysis module, which is used to determine the reference visitor through the basic identity information, obtain the historical access records of the reference visitor, calculate the inertial access frequency in the historical access records, and at the same time determine the common information group, and then extract the keywords in the common information group to obtain the access keywords of the target visitor; An integrated analysis module, which is used to obtain the system keywords of the data system and calculate the data association value between the system keywords and the access keywords; A security and openness module, which is used to compare the data association value with the privacy level sequence to determine the open information, and display the information content corresponding to the open information to the target visitor.
2. The data security proxy system for a trusted data space according to claim 1, wherein The data system refers to the data stored under different data systems. The individual data set refers to the set of data of one type in the data system, and a data system is composed of multiple individual data sets.
3. The data security proxy system for a trusted data space according to claim 1, characterized in that, The method for obtaining the privacy score of the privacy keyword includes: S1: Arbitrarily select a data system as the specified analysis object, obtain the individual data sets in this data system, identify the data types of each individual data set, and mark the information words corresponding to this data type as the privacy keywords of this individual data set. Among them, one individual data set corresponds to one privacy keyword; S2: Set decision-making influencing factors. The decision-making influencing factors refer to the necessary relevant factors in privacy analysis. The decision-making influencing factors include information sensitivity, data publicity, and information relevance; Arbitrarily select a privacy keyword and mark it as the key word, and at the same time arbitrarily select a decision-making influencing factor and mark it as the target factor. First, obtain the public cases of the key word, and use natural semantic processing technology to perform semantic parsing on the public cases of the key word to obtain the open influence information of the key word. Among them, the public case refers to the case that causes trouble or harm to the privacy user after the key word is illegally made public; Conduct harm assessment on the open influence information of the key word according to the decision-making influencing factors respectively to obtain multiple evaluation values; Perform mean processing on all the evaluation values in a decision-making influencing factor, and mark the result value output after mean processing as the decision-making influence value Ci of this decision-making influencing factor. i represents different decision-making influence values, and i = 1, 2, 3, which are the information sensitivity, data publicity, and information relevance in the decision-making influencing factors respectively; S3: Use the formula DF = C1×a1 + C2×a2 + C3×a3 to obtain the privacy score DF of the key word, where a1, a2, and a3 are proportionality coefficients respectively, and a1 + a2 + a3 = 1.
4. The data security proxy system for a trusted data space according to claim 3, characterized in that, The method for obtaining the privacy level sequence of a data system includes: Take the privacy keywords of all individual data sets in the data system as key words in turn, and the privacy score DF of each privacy keyword. Then obtain the privacy scores DF of all privacy keywords in a data system, and arrange the privacy scores DF in ascending order. At the same time, mark the arranged combination as the privacy level sequence.
5. The data security proxy system for a trusted data space according to claim 3, characterized in that, Harm assessment refers to setting different score levels according to the impact caused by illegal disclosure, and each score level corresponds to an assessment value. Among them, the score levels include level one, level two, level three, level four, and level five, corresponding to the assessment values of 1 point, 2 points, 3 points, 4 points, and 5 points respectively. Among them, for the degree of harm of the impact: level five > level four > level three > level two > level one.
6. The data security proxy system for a trusted data space according to claim 1, characterized in that, The method for obtaining the inertial access frequency includes: SS1: Based on the basic identity information of the target visitor, obtain other visitors with the same identity information in the trusted database and mark them as reference visitors. Among them, the same identity information refers to the same occupation range and the same occupation level; Obtain the historical access records of the reference visitors, identify and summarize the access information in the historical access records, that is, integrate the same access information to obtain multiple groups of information data. Among them, each group of information data corresponds to an information point; SS2: Statistically analyze the data access duration in each group of information data to obtain the inertial duration Tj, where j represents different information data groups, and j ∈ [1, J], indicating that there are a total of J information data groups; The inertia access frequency FPj of each information data group is obtained by using the formula FPj = Tj ÷ Tz, where Tz is the total access duration of the reference visitor, that is .
7. The data security proxy system for a trusted data space according to claim 6, characterized in that, The method for obtaining the access keywords of the target visitor includes: Compare the inertial access frequency FPj with the frequency threshold Fy. If FPj < Fy, mark the corresponding information data group as a restricted use data group. Otherwise, if FPj ≥ Fy, mark the corresponding information data group as a common information group; Extract all common information groups, use the TextRank algorithm to calculate the keywords in the common information groups, and mark the words with a keyword score exceeding X1 as the access keywords of the target visitor.
8. The data security proxy system for a trusted data space according to claim 1, characterized in that, The calculation method of the data association value includes: Obtain all data systems in the trusted data space, obtain the corresponding theme name of each data system, and mark the theme name as the system keyword of the corresponding data system; Use the cosine similarity algorithm to perform similarity association processing on the access keywords of the target visitor and the system keywords of each data system respectively to obtain the data association value of each data system. Among them, one data system corresponds to one data association value.
9. The data security proxy system for a trusted data space according to claim 1, characterized in that, The method for determining open information includes: Obtain the real-time access trace of the target visitor, identify the data system corresponding to the real-time access trace, obtain the privacy level sequence corresponding to this data system, and mark it as the target sequence; Obtain the total number M of privacy keywords in the target sequence, and then use the formula W = M × GL to obtain the public range value W, where GL represents the data association value of the data system corresponding to the target sequence; After that, take the integer part of the public range value W, and mark the taken integer part as the public positioning value. Obtain the target sequence again, take the position corresponding to the first privacy keyword in the target sequence as the starting position, and in sequence, mark the privacy keywords within the public positioning value as open information; After that, the security opening module identifies the open information and displays the information content corresponding to the open information to the target visitor.
Citation Information
Patent Citations
Internet of Things security agent data sharing module design method adopting a block chain technology
CN110266687A
Privacy data hierarchical protection method based on big data
CN112100670A
Public open data-based safety privacy integrated protection system and method thereof
CN112528272A
System and method for determining risk of group privacy leakage
CN113709090A
Multi-layer and multi-level personalized local differential privacy method for classifying and grading data frequency estimation
CN117390683A
Cited By
Data exchange security verification method based on trusted data space
CN120811722A
A data exchange security verification method based on a trusted data space
CN120811722B