A Dynamic Sensitive Data Masking Method and Related Devices in a Cloud Environment
By generating task chain permission dependency graphs and dynamically determining desensitization rules in the cloud environment, the problem of low efficiency in sensitive data access permission management and data desensitization processing in the cloud environment is solved, and the system response speed and data security are improved.
Patent Information
- Application Number
- CN202411955450.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-28
AI Technical Summary
In the cloud environment, the existing technology has low access rights management and data desensitization processing efficiency for sensitive data in multi-role collaborative working scenarios, resulting in slow system response and affecting business processing efficiency.
By generating a task chain permission dependency graph in the cloud environment, predicting the user's next task operation and preloading the corresponding permissions, dynamically determining the desensitization rules, and desensitizing processing in real time during data access.
It improves the accuracy of system response speed and permission management, ensures that users efficiently access and process sensitive data under the premise of legality and compliance, and enhances the security and controllability of data.
Smart Images

Figure CN119358039B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data desensitization, and particularly to a method and related device for dynamically desensitizing sensitive data in a cloud environment. Background Art
[0002] With the rapid development of cloud computing technology, more and more enterprises choose to migrate their business systems to the cloud environment. In the cloud environment, the security protection of enterprise sensitive data faces greater challenges. Especially in the scenario of multi-role collaborative work, the access rights of different roles to sensitive data need to be managed and controlled with fine-grained.
[0003] In the related art, the sensitive data protection scheme usually adopts a static data desensitization method, that is, fixed masking or replacement processing is performed on sensitive fields when data is stored. At the same time, the access rights of users to sensitive data are managed through a pre-set role permission configuration table, and users need to perform identity authentication and permission verification when accessing data.
[0004] However, when users need to process cross-system related services, they need to re-perform permission authentication and data desensitization rule matching every time they access a new service module. This mechanism will cause the system processing response to be slow and affect the business processing efficiency. Summary of the Invention
[0005] This application provides a method and related device for dynamically desensitizing sensitive data in a cloud environment, which is used to improve the processing efficiency of data desensitization.
[0006] In a first aspect, this application provides a method for dynamically desensitizing sensitive data in a cloud environment, which is applied to a system for dynamically desensitizing sensitive data in a cloud environment. The method includes: generating a task chain permission dependency graph according to the role permission combination, where the task chain permission dependency graph includes the permission inheritance and transfer relationships between each task node, and the role permission combination includes the viewing permission level, data operation permission, and timeliness constraint; predicting the next task operation of the target user in the task chain permission dependency graph according to the position of the task node where the target user is currently located, and preloading the target role permission corresponding to the next task operation; determining a preset dynamic desensitization rule for the target sensitive data according to the target role permission; and desensitizing the target sensitive data according to the preset dynamic desensitization rule during the data access process of the target sensitive data.
[0007] By adopting the above technical solution, a task chain permission dependency graph is generated according to the role permission combination, which includes the permission inheritance and transfer relationships between task nodes. The role permission combination includes the viewing permission level, data operation permission, and timeliness constraint. When a user performs a task operation, based on this graph, the next operation can be accurately predicted and the corresponding permissions can be pre-loaded, thereby improving the system's response speed and the accuracy of permission management, ensuring that users can efficiently access and process sensitive data under the premise of being legal and compliant, and enhancing the security and controllability of sensitive data in the cloud environment.
[0008] Combined with some embodiments of the first aspect, in some embodiments, before the step of generating the task chain permission dependency graph according to the role permission combination, the method further includes: obtaining the operation timing data of the target user, and extracting the standard work task chain of the target user according to the operation timing data; labeling each task node of the standard work task chain with the role permission combination corresponding to the task node.
[0009] By adopting the above technical solution, before generating the task chain permission dependency graph, first obtain the operation timing data of the target user and extract the standard work task chain, and then label each task node with the role permission combination. Through the analysis and processing of the operation timing data, the user's work mode and habits can be accurately captured, making the generated standard work task chain more in line with the actual work situation. The labeling of the role permission combination for each task node further refines the permission management, enabling the system to perform more accurate permission control and data desensitization processing according to the specific permission requirements of the user at different task nodes.
[0010] Combined with some embodiments of the first aspect, in some embodiments, the step of obtaining the operation timing data of the target user and extracting the standard work task chain of the target user according to the operation timing data specifically includes: performing time series analysis on the operation timing data to identify the repeated operation patterns in the operation timing data; determining the operation sequences with frequencies exceeding a preset threshold in the repeated operation patterns as candidate task chains; calculating the task similarity of the candidate task chains, and merging the candidate task chains with similarities greater than a preset similarity threshold to obtain the standard work task chain of the target user, where the task similarity is determined by calculating the matching degrees of the operation types, operation objects, and operation times of the task nodes in the candidate task chains.
[0011] By adopting the above technical solution, when extracting the standard work task chain according to the operation timing data, time series analysis is performed on the data to identify repeated operation patterns. The operation sequence with a frequency exceeding the preset threshold is determined as the candidate task chain, and then the standard work task chain is obtained through calculating the task similarity and merging. This method makes full use of the historical data of user operations, and mines out the representative work processes through scientific analysis methods, so that the generated standard work task chain can more objectively and accurately reflect the actual work situation of users.
[0012] Combined with some embodiments of the first aspect, in some embodiments, the step of generating the task chain permission dependency graph according to the role permission combination specifically includes: obtaining the role permission combination corresponding to each task node in the standard work task chain; analyzing the role permission combinations of adjacent task nodes in the standard work task chain to generate a permission correlation matrix; determining the permission inheritance relationship between task nodes according to the permission correlation matrix, and generating an initial permission dependency graph; analyzing the permission transfer relationship between non-adjacent task nodes according to the initial permission dependency graph, where the permission transfer relationship indicates that the permission can be transferred through intermediate task nodes; updating the permission transfer relationship to the initial permission dependency graph to form the task chain permission dependency graph.
[0013] By adopting the above technical solution, in the process of generating the task chain permission dependency graph according to the role permission combination, first obtain the role permission combination corresponding to each task node in the standard work task chain, then analyze the role permission combinations of adjacent task nodes to generate a permission correlation matrix, and then determine the permission inheritance relationship between task nodes and generate an initial permission dependency graph, and then analyze the permission transfer relationship between non-adjacent task nodes and update it to the initial permission dependency graph, so that the task chain permission dependency graph can comprehensively and accurately reflect the complex permission relationships between task nodes, including direct permission inheritance and indirect permission transfer.
[0014] Combined with some embodiments of the first aspect, in some embodiments, the step of predicting the next task operation of the target user in the task chain permission dependency graph according to the position of the task node where the target user is currently located specifically includes: obtaining the historical operation trajectory of the target user in the task chain permission dependency graph; analyzing the association pattern between task nodes in the historical operation trajectory to obtain the association relationship between task nodes; determining the set of predicted task operations that the target user may execute according to the association relationship and the position of the task node where the target user is currently located; based on the historical execution frequency of each task node in the set of predicted task operations, selecting the task node with the highest execution probability as the next task operation.
[0015] By adopting the above technical solution, according to the position of the current task node where the target user is located, the next task operation of the target user is predicted in the task chain permission dependency graph. By obtaining the historical operation trajectory of the user in the task chain permission dependency graph, the association relationship is obtained by analyzing the association pattern between task nodes, and then the predicted task operation set is determined in combination with the current task node position. Based on the historical execution frequency, the task node with the highest execution probability is selected as the next task operation, which can anticipate the user's behavior in advance and make corresponding preparations for the system in advance, such as preloading the permissions of the target role, etc.
[0016] Combined with some embodiments of the first aspect, in some embodiments, after the step of desensitizing the target sensitive data according to the preset dynamic desensitization rule during the data access process of the target sensitive data, the method further includes: obtaining the historical audit log data of the target user within a preset time period, where the historical audit log data includes the access time, access frequency, access location, data operation type of the target user to the target sensitive data, and the task node conversion record in the task chain permission dependency graph; training a machine learning model based on the historical audit log data to obtain the normal access behavior benchmark model of the target user; evaluating the real-time access behavior of the target user according to the normal access behavior benchmark model, and when the deviation between the real-time access behavior and the normal access behavior exceeds a preset threshold, adjusting the desensitization level in the preset dynamic desensitization rule to a preset level.
[0017] By adopting the above technical solution, after desensitizing the target sensitive data, the historical audit log data of the target user within a preset time period is obtained, and a machine learning model is trained based on this to obtain the normal access behavior benchmark model, and then the real-time access behavior of the target user is evaluated according to this model. When the deviation between the real-time access behavior and the normal access behavior exceeds a preset threshold, the desensitization level in the preset dynamic desensitization rule is adjusted to a preset level, realizing the real-time monitoring of the user's access behavior and dynamically adjusting the desensitization level, and being able to detect abnormal access behavior in time and take corresponding protection measures.
[0018] Combined with some embodiments of the first aspect, in some embodiments, after the step of training a machine learning model based on the historical audit log data to obtain the normal data access behavior benchmark model of the target user, the method further includes: real-time collecting the current access data of the target user and inputting it into the normal data access behavior benchmark model for calculation to obtain the abnormal degree score of the current access behavior; adjusting the access permission of the target sensitive data according to the abnormal degree score.
[0019] By adopting the above technical solution, after training a machine learning model based on historical audit log data to obtain a benchmark model of the normal data access behavior of the target user, the current access data of the target user is collected in real time and input into the model to calculate the anomaly degree score, and then the access permission of the target sensitive data is adjusted according to the anomaly degree score. When detecting abnormal user behavior, by reducing or restricting their access permission, possible malicious operations or misoperations can be effectively prevented, further strengthening the protection of sensitive data.
[0020] In a second aspect, an embodiment of the present application provides a sensitive data dynamic desensitization system in a cloud environment. The sensitive data dynamic desensitization system in the cloud environment includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the sensitive data dynamic desensitization system in the cloud environment to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product containing instructions. When the above computer program product runs on a sensitive data dynamic desensitization system in a cloud environment, it enables the sensitive data dynamic desensitization system in the cloud environment to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium including instructions. When the above instructions run on a sensitive data dynamic desensitization system in a cloud environment, it enables the sensitive data dynamic desensitization system in the cloud environment to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0023] It can be understood that the sensitive data dynamic desensitization system in the cloud environment provided in the second aspect above, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, and will not be elaborated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0025] 1. This application generates a task chain permission dependency graph based on role permission combinations, which includes permission inheritance and transfer relationships between task nodes. The role permission combinations include viewing permission levels, data operation permissions, and timeliness constraints. When a user performs a task operation, this graph can be used to accurately predict the next operation and preload the corresponding permissions, thereby improving the system's response speed and the accuracy of permission management, ensuring that users can efficiently access and process sensitive data under the premise of being legal and compliant, and enhancing the security and controllability of sensitive data in the cloud environment.
[0026] 2. When this application extracts the standard work task chain from the operation time series data, it performs time series analysis on the data to identify repeated operation patterns, determines the operation sequences with frequencies exceeding a preset threshold as candidate task chains, and then merges them through calculating task similarity to obtain the standard work task chain. This method makes full use of the historical data of user operations and uses scientific analysis methods to mine representative work processes, making the generated standard work task chain more objectively and accurately reflect the actual work situation of users.
[0027] 3. This application predicts the next task operation of the target user in the task chain permission dependency graph according to the current task node position of the target user. By obtaining the historical operation trajectory of the user in the task chain permission dependency graph, analyzing the association patterns between task nodes to obtain the association relationship, then combining the current task node position to determine the set of predicted task operations, and selecting the task node with the highest execution probability based on the historical execution frequency as the next task operation. It can anticipate the user's behavior in advance and make corresponding preparations for the system in advance, such as preloading the target role permissions, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a schematic flowchart of a method for dynamically desensitizing sensitive data in a cloud environment according to an embodiment of the present application;
[0029] Figure 2 is another schematic flowchart of a method for dynamically desensitizing sensitive data in a cloud environment according to an embodiment of the present application;
[0030] Figure 3 is still another schematic flowchart of a method for dynamically desensitizing sensitive data in a cloud environment according to an embodiment of the present application;
[0031] Figure 4 is a schematic structural diagram of an entity device of a system for dynamically desensitizing sensitive data in a cloud environment according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the above", "the", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations of one or more of the listed items.
[0033] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0034] For ease of understanding, the method provided in this embodiment is described in terms of a process below. Please refer to Figure 1 , which is a schematic flowchart of a method for dynamically desensitizing sensitive data in a cloud environment in an embodiment of the present application.
[0035] S101. Generate a task chain permission dependency graph according to the role permission combination. The task chain permission dependency graph includes the permission inheritance and transfer relationships between each task node, and the role permission combination includes a view permission level, data operation permissions, and timeliness constraints.
[0036] Among them, the role permission combination represents the set of permissions granted to a user in a specific task, including attributes such as data access level, operation type, and validity period. The task chain permission dependency graph refers to a directed graph structure that describes the permission relationships between task nodes and is used to represent the inheritance and transfer rules of permissions. A task node represents an independent operation unit in a workflow. The permission inheritance relationship refers to the subordinate relationship in which a high-level task node includes the permissions of a low-level task node. The permission transfer relationship represents the mechanism by which permissions are indirectly transferred through intermediate nodes. The view permission level refers to the grading standard for the data visible range. The data operation permission represents the authorization range for operations such as adding, deleting, modifying, and querying data. The timeliness constraint is used to limit the valid time period of permissions.
[0037] During the system initialization phase, it is necessary to establish a task chain permission dependency graph as the basic framework for permission management. Specifically, the system first obtains the permission configuration information of all task nodes, including the permission level, operation permission, and time limit constraint of each node. Then, it analyzes the permission relationships between adjacent task nodes to determine the inheritance direction and transfer rules of permissions. For permission inheritance, the system establishes a parent-child node relationship to ensure that child nodes inherit the basic permissions of parent nodes. For permission transfer, the system defines the transfer path and transfer conditions to achieve cross-node permission sharing. Finally, all permission relationships are integrated into a complete dependency graph structure.
[0038] In some embodiments, the generation of the task chain permission dependency graph can be achieved in various ways. Optionally, first construct a permission attribute table for task nodes to define the basic permission configuration of each node; then analyze the business relationships between nodes to determine the permission inheritance and transfer rules; finally, use a graph data structure to store all permission relationships. Optionally, adopt a bottom-up approach, first define the permissions of the most basic task nodes; then gradually construct the permission hierarchy through permission extension rules; finally, form a complete permission dependency network. It can be understood that other methods can also be used to generate the task chain permission dependency graph, such as methods based on historical data analysis, expert rule definition, etc., which are not limited here.
[0039] S102. According to the current task node position of the target user, predict the next task operation of the target user in the task chain permission dependency graph, and preload the target role permissions corresponding to the next task operation.
[0040] Among them, the target user refers to the system user who currently needs to perform permission prediction. The task node position refers to the work process node where the user is currently executing. The next task operation refers to the subsequent tasks that the user may perform after completing the current node. The target role permissions refer to the permission configuration required for predicting the task node. Permission preloading means preparing the permission data required for the next operation in advance. The task chain permission dependency graph is a directed graph model that describes the permission relationships between task nodes. The prediction process is a calculation process that infers the user's next behavior based on historical data and the current state.
[0041] During the user's task execution process, the system needs to predict the user's next operation in real time and prepare the required permissions in advance. Specifically, the system first obtains the task node information where the user is currently located, including the node type, execution status, and context data. Then, according to the node relationships in the task chain permission dependency graph and combined with the user's historical operation patterns, it predicts the most likely next operation. For the predicted task node, the system preloads the corresponding permission data from the permission configuration library in advance, including access permissions, operation permissions, and validity period restrictions.
[0042] In some embodiments, task operation prediction and permission preloading can be achieved in various ways. Optionally, machine learning methods are used to analyze the user's historical operation records to establish a behavior prediction model; the most likely next operation is selected according to the model prediction result; the corresponding permission configuration is preloaded into the memory cache. Optionally, task flow rules are defined based on a rule engine; the execution conditions and business parameters of the current node are analyzed; the next operation is determined according to rule matching and the permission is preloaded. It can be understood that other methods can also be used to achieve task operation prediction and permission preloading, such as statistical probability models, expert systems, etc., which are not limited here.
[0043] S103. Determine the preset dynamic desensitization rule of the target sensitive data according to the target role permission.
[0044] Among them, the target role permission represents the permission configuration granted to the user in a specific task node. The target sensitive data refers to the business data that needs to be protected and desensitized. The preset dynamic desensitization rule represents the data desensitization policy automatically generated according to the permission level. Dynamic desensitization refers to the mechanism of adjusting the desensitization policy in real time according to the access scenario and permission level. The desensitization rule is used to define the processing method of the data field. The rule determination process refers to the mapping process of converting the permission requirements into specific desensitization policies. The sensitive data field represents the data attribute that needs to be protected. The field processing method refers to the specific method of transforming or hiding the sensitive data.
[0045] When the system receives a data access request, it is necessary to determine an appropriate desensitization policy according to the user's role permission. Specifically, the system first parses the data access level, operation permission, and time limit constraint in the target role permission. Then, a matching basic rule set is selected from the desensitization rule template library. For each sensitive data field, the system selects the corresponding desensitization method according to the permission level, such as masking, replacement, encryption, etc. At the same time, considering the special requirements of the business scenario, the desensitization parameters and exception rules are adjusted. Finally, a complete dynamic desensitization rule configuration is generated, including field mapping, processing method, and parameter setting.
[0046] In some embodiments, the determination of the dynamic desensitization rule can be achieved in various ways: Optionally, the permission configuration information is read through a rule engine, and the predefined desensitization rule template is matched; the desensitization parameters are adjusted according to the permission level and business requirements; a rule execution script is generated and deployed to the data access layer. Optionally, in the configuration mapping method, a correspondence table between the permission level and the desensitization policy is established; the user permission attributes are parsed to select an appropriate policy combination; the dynamic desensitization rule configuration is generated and cached in the memory. It can be understood that other methods can also be used to determine the dynamic desensitization rule, such as machine learning methods, expert system reasoning, etc., which are not limited here.
[0047] S104. During the data access process of the target sensitive data, desensitize the target sensitive data according to the preset dynamic desensitization rule.
[0048] Among them, the data access process refers to the complete process by which the system reads and returns sensitive data. The preset dynamic desensitization rule refers to a set of predetermined data processing strategies. Desensitization processing is used to represent the operation of transforming or hiding sensitive data. The target sensitive data refers to the original data content that needs to be protected. The processing process represents the specific steps of executing the desensitization rule. Data transformation refers to the method of changing the data display form through a specific algorithm. Data hiding represents the technology of replacing sensitive information with placeholders. Rule execution refers to the process of applying the desensitization strategy to specific data.
[0049] When a user requests access to sensitive data, the system needs to perform real-time desensitization processing before returning the data. Specifically, the system first intercepts the query request at the data access layer and loads the desensitization rule configuration of the current session. Then it parses the query result set to identify the data fields that need to be desensitized. For each sensitive field, the system calls the corresponding desensitization processor to perform data conversion operations. The processor selects an appropriate algorithm according to the rule configuration, such as regular expression replacement, encryption transformation, randomization processing, etc. Finally, the desensitized data is reassembled into a result set and returned to the user.
[0050] In some embodiments, the desensitization processing of sensitive data can be implemented in various ways: Optionally, obtain the query result through a data access interceptor, parse the fields that need to be desensitized; call the corresponding desensitization algorithm for data conversion; repackage and return the processed data. Optionally, use a parallel processing framework to slice and process the data set concurrently; apply the corresponding desensitization rule to each slice; merge the processing results and perform consistency verification. It can be understood that other ways can also be adopted to implement the desensitization processing of sensitive data, such as streaming processing, asynchronous processing and other technical solutions, which are not limited here.
[0051] The following further describes the more specific process of the method provided in this embodiment. Please refer to Figure 2 , which is another process schematic diagram of the dynamic desensitization method for sensitive data in the cloud environment in the embodiment of the present application.
[0052] S201. Obtain the operation timing data of the target user, perform time series analysis on the operation timing data, and identify the repeated operation patterns in the operation timing data.
[0053] The target user refers to a specific user account that needs to access sensitive data in the cloud environment. Operational timing data refers to the recorded data of all operation behaviors of users in the system and their occurrence times, including operations such as logging in to the system, viewing files, editing data, submitting approvals, etc., and the specific timestamp information of their occurrence. Time series analysis is a data analysis method that performs statistics and pattern recognition on data arranged in chronological order. Repeated operation patterns refer to fixed operation combinations that appear multiple times in the user's operation sequence.
[0054] First, the system continuously records each operation of the user through the log collection module. The recorded content includes information such as operation ID, operation type, operation object, operation time, etc. For the collected operation log data, the system uses the sliding window method for sequence pattern mining. The window size can be set to 8 hours (one working day), and the sliding step is 1 hour. Within each window, a sequence pattern mining algorithm (such as the PrefixSpan algorithm) is used to identify frequently occurring operation sequences. During the identification process, the operation sequences are sorted according to the timestamps, and by calculating features such as operation interval time and operation type conversion rules, operation combinations with temporal correlations are extracted.
[0055] S202. Determine the operation sequences with frequencies exceeding the preset threshold in the repeated operation pattern as candidate task chains.
[0056] The repeated operation pattern refers to a fixed operation combination that appears multiple times in the timing data. The occurrence frequency refers to the proportion of the number of times a specific operation sequence appears in the overall operation records. The preset threshold is a frequency determination criterion predefined by the system for screening high-frequency operation sequences. The candidate task chain refers to a combination of operation sequences that may represent the standard work process of the user.
[0057] The system performs frequency statistics on the identified repeated operation patterns, calculates the number of times each operation sequence appears in all operation records, and divides by the total number of working days to obtain the daily average frequency. The preset threshold can be set to 0.3, that is, if the average daily occurrence frequency of an operation sequence exceeds 30%, it is determined as a candidate task chain. During the statistical process, the system also considers the time span of the operation sequence, requiring that the operation sequence must maintain a high frequency for more than 3 consecutive months to ensure that the selected ones are stable work patterns rather than temporary operation combinations.
[0058] S203. Calculate the task similarity of the candidate task chain, and merge the candidate task chains with similarity greater than the preset similarity threshold to obtain the standard work task chain of the target user. The task similarity is determined by calculating the matching degree of the operation types, operation objects, and operation times of each task node in the candidate task chain.
[0059] Task similarity is an indicator to measure the similarity degree between different candidate task chains. The operation type refers to the specific operation method for data, such as viewing, modifying, deleting, etc. The operation object refers to the specific data item or functional module to be operated. The operation time refers to the time points when each operation occurs and their timing relationships. The preset similarity threshold is the standard value for judging whether to merge task chains. The standard work task chain is the standardized work process obtained after merging.
[0060] S204. Label each task node of the standard work task chain with the role - permission combination corresponding to this task node.
[0061] The standard work task chain refers to the standardized work - flow sequence after identification and merging. A task node refers to each independent operation step in the work process. The role - permission combination includes three dimensions: the viewing permission level (determining the visible range of data), the data operation permission (such as reading, modifying, deleting, etc.), and the timeliness constraint (the expiration period of the permission). The labeling process is a process of establishing a mapping relationship between role - permission information and task nodes. For example, for the task node of "customer information review", its corresponding role - permission combination is {viewing level: visible for second - level sensitive data, operation permission: read - only, timeliness: 9:00 - 18:00 on weekdays}.
[0062] The system realizes permission labeling through a rule engine. First, establish a task - permission mapping rule library, which includes a basic permission rule set and a scenario - permission rule set. The basic permission rule set defines the standard permission requirements for different types of task nodes. For example, the task of viewing customer information requires the customer basic information viewing permission. The scenario - permission rule set defines the additional permission requirements in specific business scenarios. For example, the temporary authorization rule in the emergency approval scenario. For each task node, the system sequentially performs permission rule matching to determine the minimum permission set required for this node. The permission - labeling result is stored in the permission configuration table, which includes fields such as task - node ID, permission - combination ID, effective time, expiration time, etc.
[0063] S205. Obtain the role - permission combination corresponding to each task node in the standard work task chain.
[0064] A role - permission combination is a set of defined permissions, including permission types, permission levels, and constraints. Permission types define the ways of operating on data, such as reading, modifying, deleting, etc. Permission levels determine the visible scope of data, such as fully visible, partially visible, etc. Constraints include time limits, geographical location limits, etc. Task nodes in the standard work task chain are the basic operation units in the work process, and each node requires a specific permission combination to execute. The system reads the permission information of task nodes from the permission configuration database. By querying the permission configuration table through the task node ID, all permission records associated with this node are obtained. Permission records include fields such as permission ID, permission type code, permission level code, constraints, etc. The system parses the queried permission records into permission objects, which include permission attributes and constraint rules. For each task node, the system constructs a permission combination object, which contains all the permission information required for this node. The permission combination object is stored in a tree structure for subsequent analysis of permission inheritance and transmission.
[0065] S206. Analyze the role - permission combinations of adjacent task nodes in the standard work task chain to generate a permission correlation matrix.
[0066] Adjacent task nodes refer to nodes that have a direct before - and - after relationship in the work process. A role - permission combination is the set of permissions required for each node. Permission correlation is an indicator that measures the degree of closeness between two permission combinations. The permission correlation matrix is a two - dimensional matrix that records the permission correlation degrees between all nodes in the task chain. For example, for two adjacent nodes, "data entry" and "data review", their permission combinations have a relatively high correlation because the permissions of the review node include some of the permissions of the entry node.
[0067] The system uses the permission feature vector method to calculate the permission correlation. First, each permission combination is converted into a permission feature vector, and the vector dimensions include features such as data access level, operation type, time constraint, etc. For each pair of adjacent nodes in the task chain, the correlation score of their permission feature vectors is calculated. The correlation calculation uses the weighted summation method, and the weights of different permission features are: data access level 0.4, operation type 0.3, time constraint 0.3. The calculated correlation is filled into the N×N correlation matrix, where N is the number of nodes in the task chain. Each element aij in the matrix represents the permission correlation from node i to node j, and the value range is [0, 1].
[0068] S207. Determine the permission inheritance relationship between task nodes according to the permission correlation matrix and generate an initial permission dependency graph.
[0069] The permission correlation matrix is an N×N matrix that records the numerical values of the permission correlation degrees between task nodes. The permission inheritance relationship refers to the subordinate relationship where the permissions of one task node contain those of another node. The initial permission dependency graph is a directed graph structure, with nodes representing tasks and edges representing permission inheritance relationships. For example, in the loan approval process, the "Senior Approval" node inherits all the permissions of the "Junior Approval" node, forming an inheritance relationship.
[0070] The system constructs the initial permission dependency graph through the following steps: First, set the permission inheritance threshold to 0.8. Traverse each element aij in the permission correlation matrix. When aij is greater than the threshold, it is determined that there is a permission inheritance relationship between node i and node j. Use the graph data structure to store the permission dependency relationship. Each graph node contains attributes such as task ID and permission combination. The graph edges represent the inheritance relationship, and the weight of the edge is the correlation value. The system uses an adjacency list to store the graph structure for subsequent graph traversal and analysis. Perform a transitive closure calculation on the identified inheritance relationships to ensure the integrity of the inheritance relationships. Finally, serialize and store the constructed graph structure in the graph database.
[0071] S208. Analyze the permission transfer relationship between non-adjacent task nodes based on this initial permission dependency graph. This permission transfer relationship indicates that permissions can be transferred through intermediate task nodes.
[0072] The initial permission dependency graph is a directed graph representing direct permission inheritance relationships. Non-adjacent task nodes refer to nodes that are not directly connected in the task chain. The permission transfer relationship refers to the chain of relationships where permissions are indirectly transferred through intermediate nodes. Intermediate task nodes are the transitional nodes in the permission transfer process. For example, in the process of "Data Entry - Data Review - Final Approval", the final approval node can indirectly obtain some of the permissions of the data entry node through the data review node.
[0073] The system uses a graph traversal algorithm to analyze the permission transfer relationship. First, use the Floyd-Warshall algorithm to calculate the shortest path between any two nodes. The path length uses the reciprocal of the permission correlation as the edge weight. For each pair of non-adjacent nodes (i, j), if there exists a path with a product of permission correlations greater than 0.6, it is considered that there is a permission transfer relationship between node i and j. The system records all the intermediate nodes on the transfer path and calculates the actual permission range after transfer. The permission transfer calculation uses the method of taking the intersection of permissions to ensure that the permission range does not expand during the transfer process. Store the analyzed transfer relationship in the transfer relationship table, which contains information such as the source node, target node, intermediate node sequence, and transferred permission range.
[0074] S209. Update the permission transfer relationship to the initial permission dependency graph to form a task chain permission dependency graph. The task chain permission dependency graph contains the permission inheritance and transfer relationships between each task node. The role permission combination includes the view permission level, data operation permission, and timeliness constraint.
[0075] The permission transfer relationship is an indirect permission sharing mechanism implemented through intermediate nodes. The task chain permission dependency graph is a complete graph structure that includes all direct inheritance and indirect transfer permission relationships. The view permission level defines the hierarchical visibility scope of the data. The data operation permission specifies the specific operation types on the data. The timeliness constraint specifies the valid time range of the permission. For example, the permission combination of a certain node is {view level: level 3, operation permission: read / write, timeliness: valid within 24 hours}.
[0076] The system integrates the permission transfer relationship into the initial permission dependency graph. First, convert the records in the transfer relationship table into edge relationships of the graph. The attributes of the edge include information such as transfer type, transfer path, and the permission scope after transfer. For each transfer relationship, add a dashed edge from the source node to the target node in the graph to represent the indirect permission relationship. The updated graph structure is stored using an attribute graph database. The node attributes include task information and permission combination, and the edge attributes include relationship type (direct inheritance / indirect transfer) and permission transfer rules. The system attaches a permission calculator to each edge for dynamically calculating the actual available permission scope. The final task chain permission dependency graph supports functions such as permission inheritance relationship query, permission transfer path analysis, and permission scope verification.
[0077] S210. Obtain the historical operation track of the target user in the task chain permission dependency graph.
[0078] The target user refers to a specific user account for which permission prediction is required. The task chain permission dependency graph is a directed graph structure that describes the permission relationships between task nodes. The historical operation track is a record of the task sequence executed by the user in the system, including information such as task type, execution time, and execution order. For example, the historical operation track of a credit approval officer includes a sequence of data such as {time 1: receive application, time 2: review materials, time 3: risk assessment, time 4: approval decision}.
[0079] The system extracts the operation records of the target user in the recent 90 days from the operation log database. The operation records include fields such as operation ID, task node ID, operation timestamp, operation status, etc. The system first sorts the operation records by timestamp to construct a time-series operation sequence. For each operation record, it queries the task chain permission dependency graph to obtain the complete information of the corresponding node, including node attributes and permission information. The system uses the sliding window method to group consecutive operation records into operation sessions, and each session represents a complete business processing flow. Finally, all operation sessions are integrated into the user's historical operation trajectory and stored as an ordered sequence of task nodes.
[0080] S211. Analyze the association patterns between task nodes in the historical operation trajectory to obtain the association relationships between these task nodes.
[0081] The historical operation trajectory is a time-series sequence of task nodes executed by the user. The association pattern refers to the regular combination methods that appear between task nodes. The association relationships between task nodes describe the sequential dependencies and conditional trigger relationships between nodes. For example, in the loan approval process, the "risk assessment" node always executes after the "data review" node and must be started within 4 hours after the completion of the "data review", which constitutes an association pattern.
[0082] The system uses the sequence pattern mining algorithm to analyze the operation trajectory. First, it uses the SPADE (Sequential Pattern Discovery using Equivalence classes) algorithm to identify frequently occurring task sequence patterns, and the support threshold is set to 0.3. For each identified sequence pattern, it calculates the time interval distribution between task nodes and uses the kernel density estimation method to obtain the optimal time window. The system constructs a task transition probability matrix, and the matrix elements represent the transition probabilities between nodes. It uses the association rule mining algorithm to analyze the conditional dependency relationships between nodes to obtain a rule set in the form of "if task A is executed and condition C holds, then the probability of executing task B within the time window T is P". Finally, the transition probabilities and conditional rules are combined to form a complete association relationship model.
[0083] S212. Determine the set of predicted task operations that the target user may perform according to the association relationship and the current task node position of the target user.
[0084] The association relationship is a model that describes the transition rules between task nodes. The current task node position refers to the task node that the user is currently executing. The set of predicted task operations is a list of task nodes that the system predicts the user may perform next. For example, when the user completes the "loan material review" task, the set of predicted task operations includes {risk assessment (probability 0.7), return for supplementary materials (probability 0.2), directly reject (probability 0.1)}.
[0085] The system predicts the next operation through a Markov decision process. First, it obtains the status information of the user's current task node, including task completion rate, execution duration, relevant business parameters, etc. It substitutes the current status information into the association relationship model to calculate the set of subsequent task nodes that meet the trigger conditions. For each candidate task node, the system calculates the transfer probability score, and the score calculation formula is: Final score = historical transfer probability × 0.6 + time window matching degree × 0.2 + conditional rule matching degree × 0.2. The system selects the task nodes with a score exceeding 0.4 to form the predicted task operation set, and each task node in the set contains the prediction probability value and trigger condition information. The prediction results are sorted in descending order of probability for subsequent permission preloading processing.
[0086] S213. Based on the historical execution frequency of each task node in the predicted task operation set, select the task node with the highest execution probability as the next task operation, and preload the target role permissions corresponding to the next task operation.
[0087] The predicted task operation set is a list of possible subsequent task nodes predicted by the system. The historical execution frequency refers to the statistical number of times a task node is executed in the historical record. The execution probability is the likelihood of a task node being selected calculated based on historical data. The target role permission is the permission configuration required for a task node to execute. Preloading refers to the process of preparing and caching permission data in advance. For example, in the credit approval process, if the historical execution frequency of the "risk assessment" node is 10 times per day, accounting for 70% of all subsequent operations, then its probability of being selected as the next task operation is the highest.
[0088] The system uses a weighted scoring method to select the optimal task node. For each node in the predicted task operation set, it calculates the average daily execution frequency and execution time distribution within the last 30 days. It constructs a scoring function: Node score = historical execution frequency weight (0.5) × normalized frequency value + time coincidence degree weight (0.3) × time matching score + business rule weight (0.2) × rule matching score. Select the task node with the highest score as the next operation. Preload the permission configuration of the selected node, including: obtaining the complete permission definition from the permission configuration library, constructing a permission cache object, warming up the data access interface, and preparing the desensitization rule configuration. Store the preloaded permission data in the distributed cache system and set a validity period of 15 minutes.
[0089] S214. Determine the preset dynamic desensitization rule for the target sensitive data according to the target role permission.
[0090] The target role permissions are a set of permission configurations required for specific task nodes. The target sensitive data is the data object that needs access control and desensitization processing. The preset dynamic desensitization rule is a data desensitization policy dynamically generated according to the permission level. The desensitization rule defines the processing method of data fields, such as masking, replacement, encryption, etc. For example, for the customer's mobile phone number, different permission levels correspond to different desensitization rules: fully visible, hide the last 4 digits, only display the first 3 and the last 4 digits, etc.
[0091] The system determines the desensitization policy through the rule engine. First, it loads the desensitization rule template library, which includes the basic rule set and the business rule set. The basic rule set defines the desensitization methods for common data types, such as ID card numbers, mobile phone numbers, bank card numbers, etc. The business rule set defines the desensitization requirements for specific business scenarios. The system analyzes three dimensions of the target role permissions: data access level, operation permission type, and timeliness constraint, and generates a desensitization rule decision tree. According to the decision tree, it matches the appropriate desensitization rule combination, and the rule combination includes: field desensitization method, desensitization parameter configuration, and exception handling conditions. Serialize the generated desensitization rule into a configuration object and store it in the rule cache.
[0092] S215. During the data access process of the target sensitive data, desensitize the target sensitive data according to the preset dynamic desensitization rule.
[0093] The data access process refers to the process by which the system reads and returns sensitive data. The preset dynamic desensitization rule is a set of pre-defined data processing policies. The desensitization process is the process of transforming or hiding sensitive data. The target sensitive data is the original data object that needs to be protected. For example, when accessing customer information, the system desensitizes the mobile phone number "13812345678" to "138****5678" and the ID card number "440123199001011234" to "440123********1234" according to the rules.
[0094] The system implements real-time desensitization processing at the data access layer. First, it constructs a data access interceptor to intercept all query requests for sensitive data. The interceptor obtains the desensitization rule configuration of the current session from the context. Process each row of the query result set and apply the corresponding desensitization method to each data field. The implementation of the desensitization method includes: regular expression replacement, AES encryption, character masking, randomization processing, etc. The system uses a parallel stream processing framework to improve the desensitization efficiency and processes the data set in parallel by sharding. After processing, reassemble the desensitized data into a result set and return it. At the same time, record the desensitization operation log, including information such as the desensitization rule ID, processing time, and data volume. The processing delay of the entire desensitization process is controlled within 50 milliseconds.
[0095] The following further describes the more specific process of the method provided in this embodiment. Please refer to Figure 3, which is another process schematic diagram of the sensitive data dynamic desensitization method in the cloud environment in the embodiments of this application.
[0096] S301. Obtain the historical audit log data of the target user within a preset time period. The historical audit log data includes the access time, access frequency, access location, data operation type of the target user to the target sensitive data, and the task node conversion record in the task chain permission dependency graph.
[0097] The target user is a specific user account that needs to conduct behavior analysis in the system. The preset time period is a fixed time range for analysis, usually the last 90 days. The historical audit log data is all the operation records of the user in the system. The access time records the specific timestamp of each data access. The access frequency counts the number of accesses per unit time. The access location includes the IP address and geographical location information. The data operation type includes operation classifications such as query, modification, and deletion. The task node conversion record describes the node jump sequence of the user in the task process. For example, the audit log of a certain credit approval officer contains: {timestamp: 2024-01-01 10:00:00, IP: 192.168.1.100, location: headquarters building, operation: query customer information, task node: risk assessment}.
[0098] The system obtains historical data through the audit log collection module. First, determine the data collection time range, which is set to 90 days before the current time. Construct the log query conditions: user ID, time range, operation type list. Batch read the log records that meet the conditions from the distributed log storage system. Clean and format the original log data: parse the JSON format log content, extract the key fields, unify the time format to UTC timestamp, standardize the IP address format, and convert the geographical location encoding. Sort the log records in chronological order and calculate the access frequency metrics: number of accesses per hour, number of accesses per day, access distribution on weekdays and non-weekdays. Store the processed log data in the analysis database and establish an index to support fast query.
[0099] S302. Train a machine learning model based on the historical audit log data to obtain the normal access behavior benchmark model of the target user.
[0100] The historical audit log data is a set of time-series records of user behavior. The machine learning model is a behavior pattern recognizer obtained through data training. The normal access behavior benchmark model defines the standard behavior characteristics of the user. The training process is a calculation process of learning behavior patterns from historical data. For example, the model learns that the user usually accesses the system from a fixed office location during 9:00 - 18:00 on weekdays, processes an average of 10 transactions per hour, and 90% of the operations follow the standard task process.
[0101] The system adopts a multi-level behavior modeling method. First, perform feature engineering: extract time features (hour distribution, week distribution, holiday pattern), spatial features (IP address segment, geographical location set), operation features (operation type distribution, data access pattern), and task features (node transition sequence, processing duration distribution). Construct a behavior vector space, use One-Hot encoding to process categorical features, and use normalization to process numerical features. Adopt an ensemble learning method to train the model: use IsolationForest to detect outliers, use LSTM network to learn time series patterns, and use Gaussian Mixture Model (GMM) to establish behavior distributions. The model training adopts a cross-validation method, using 80% of the data for training and 20% for validation. Save the trained model parameters as a benchmark model file.
[0102] Specifically, the model can compare and analyze the real-time access behavior data of the input target user with the normal behavior patterns learned from historical audit log data, and output an evaluation result indicating whether the current real-time access behavior conforms to the normal behavior pattern. This evaluation result can be in the form of a score, a category label (such as normal or abnormal), or a probability value (representing the probability of normal behavior), etc. For example, the model may output a score between 0 and 1. When this score is close to 1, it indicates that the real-time access behavior is very similar to the normal behavior pattern, and close to 0 indicates a large difference. When this score is lower than a preset threshold (such as 0.8), it is considered that the deviation of the real-time access behavior from the normal behavior exceeds the preset threshold, thereby triggering the operation of adjusting the desensitization level in the preset dynamic desensitization rule to the preset level. The input data of the normal access behavior benchmark model is the historical audit log data of the target user within a preset time period (usually the recent 90 days), which details various access information of the target user to the target sensitive data, including access time (specific to the timestamp), access frequency (statistics of the number of accesses per hour, per day, etc.), access location (IP address and corresponding geographical location), data operation type (such as query, modification, deletion, etc.), and task node transition records in the task chain permission dependency graph. The training data is the historical audit log data after cleaning and formatting. Specific cleaning and formatting operations include: parsing the JSON format log content to extract key information such as user ID, time range, operation type, etc.; uniformly converting the time format to UTC timestamp for subsequent time series analysis; standardizing the IP address to make its format uniform; converting the geographical location information to a specific encoding format, etc. The input data is used to determine whether the target user conforms to the normal behavior pattern and can output an evaluation result, such as a score or a category label (such as normal or abnormal).
[0103] S303. Evaluate the real-time access behavior of the target user according to the normal access behavior benchmark model. When the deviation between the real-time access behavior and the normal access behavior exceeds the preset threshold, adjust the desensitization level in the preset dynamic desensitization rule to the preset level.
[0104] The normal access behavior benchmark model is a mathematical model that describes the standard behavior characteristics of users. The real-time access behavior is the operation characteristics that the user is currently performing. The evaluation process is to calculate the degree of difference between the current behavior and the benchmark model. The preset threshold is the standard value for judging abnormal behavior. The desensitization level is the intensity level of data protection, which is divided into multiple levels from low to high. The preset level is the protection level enabled after an abnormality is detected. For example, when a user accesses sensitive data in a large amount at an unusual time, the system detects the abnormality and increases the desensitization intensity.
[0105] The system monitors the user behavior in real time and conducts evaluations. Build a real-time feature extractor to calculate the behavior feature vector of the last 5 minutes every 60 seconds. Input the feature vector into the benchmark model to calculate the anomaly score: the time dimension score (the access probability at the current time point), the space dimension score (the location matching degree), the operation dimension score (the deviation of the operation frequency), and the task dimension score (the process compliance). Use the weighted average method to calculate the comprehensive anomaly score, and the weights are 0.3, 0.2, 0.3, and 0.2 respectively. Set the anomaly threshold to 0.8, and trigger the desensitization level adjustment when the comprehensive score exceeds the threshold. The desensitization level adjustment follows the preset rules: the desensitization level is increased by one level in the interval [0.8 - 0.9], increased by two levels in the interval [0.9 - 1.0], and increased to the highest level when it exceeds 1.0. The adjusted desensitization rule takes effect immediately, and the duration is set to 30 minutes.
[0106] S304. Real-time collect the current access data of the target user and input it into the normal data access behavior benchmark model for calculation to obtain the anomaly degree score of the current access behavior.
[0107] The target user is a specific user account that needs to be monitored in the system. The current access data is the system operation information that the user is performing, including real-time data such as access time, access address, and operation type. The normal data access behavior benchmark model is a user standard behavior model obtained through machine learning training. The anomaly degree score is a numerical index that measures the degree to which the current behavior deviates from the normal mode, and the value range is [0, 1]. For example, a user downloads a large amount of customer data from an unknown IP address at 3 am, and the system calculates an anomaly degree score of 0.95, indicating that the behavior is highly abnormal.
[0108] The system obtains user behavior data through a real-time data collection pipeline. The collection module triggers data collection every 5 seconds and obtains the following information through the application program interface (API): user session information (session ID, login time, terminal type), access data (request URL, HTTP method, request parameters), and system information (CPU usage, memory usage, network traffic). Feature extraction is performed on the collected data: short-term features (operation frequency and data access volume in the last minute), medium-term features (operation type distribution and access mode in the last 10 minutes), and session features (accumulated number of operations and duration of the current session) are calculated. The extracted feature vector is input into the benchmark model to calculate the anomaly scores of four dimensions: time dimension (operation probability at the current time point), spatial dimension (credibility of access location), behavioral dimension (standardization of operation sequence), and resource dimension (rationality of resource use). The weighted summation method is used to calculate the final anomaly score: anomaly score = Σ (dimension weight × dimension score), with weights of 0.3, 0.2, 0.3, and 0.2 respectively.
[0109] S305: Adjust access rights to the target sensitive data according to the abnormality level score.
[0110] The abnormality score is a quantitative behavioral risk indicator. Target sensitive data is important business data that needs to be protected. Access rights are a set of rules that define the scope of user operations on data. Permission adjustment is the process of dynamically changing data access control policies based on risk levels. For example, when the abnormality score reaches 0.8, the system reduces the user's access rights to customer sensitive information from "full access" to "read-only access" and increases the intensity of field desensitization.
[0111] The system implements a risk-based adaptive access control mechanism. First, define the abnormality level classification: normal ([0-0.6]), mild abnormality ([0.6-0.8]), moderate abnormality ([0.8-0.9]), severe abnormality ([0.9-1.0]). For each abnormality level, set the corresponding permission adjustment strategy: enable operation auditing in case of mild abnormality and record detailed operation logs; reduce data access permissions in case of moderate abnormality, reduce write permissions to read-only permissions, and increase the field desensitization level; temporarily freeze sensitive operation permissions in case of severe abnormality, retain only basic query permissions, and use the highest level of desensitization for all sensitive fields. Permission adjustment execution process: obtain the user's current permission configuration from the permission configuration library, calculate new permission rules based on the abnormality level, update the permission configuration through the distributed cache, and set the permission effective time to 30 minutes. At the same time, start the permission recovery timer, and automatically restore the original permission settings after the user behavior returns to normal (the abnormality score is lower than 0.6 for 10 minutes). All permission adjustment operations are recorded in the audit log, including adjustment time, trigger reason, permission change details and other information.
[0112] It is understandable that steps S301 to S305 can be executed after step S215.
[0113] The following describes the sensitive data dynamic desensitization system in the cloud environment in the embodiments of the present invention application from the perspective of hardware processing. Please refer to Figure 4 which is a schematic structural diagram of an entity device of the sensitive data dynamic desensitization system in the cloud environment in the embodiments of the present application.
[0114] It should be noted that Figure 4 the structure of the sensitive data dynamic desensitization system in the cloud environment shown is only an example and should not bring any restrictions to the functions and usage scope of the embodiments of the present invention.
[0115] As Figure 4 shown, the sensitive data dynamic desensitization system in the cloud environment includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403, such as executing the method described in the above embodiments. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0116] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. The drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that the computer program read from it can be installed into the storage section 408 as needed.
[0117] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, various functions defined in the present invention are executed.
[0118] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings.
[0120] Specifically, the sensitive data dynamic desensitization system in the cloud environment of this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the sensitive data dynamic desensitization method provided in the above embodiment is implemented.
[0121] As another aspect, the present invention also provides a computer-readable storage medium. This storage medium can be included in the sensitive data dynamic desensitization system in the cloud environment described in the above embodiments; or it can exist alone without being assembled into the sensitive data dynamic desensitization system in the cloud environment. The above storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the sensitive data dynamic desensitization system in the cloud environment, the sensitive data dynamic desensitization system in the cloud environment realizes the sensitive data dynamic desensitization method provided in the above embodiments.
[0122] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0123] As used in the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0124] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store program codes.
Claims
1. A method for dynamic desensitization of sensitive data in a cloud environment, characterized in that: A sensitive data dynamic desensitization system applied in a cloud environment, the method comprising: Acquire operation time series data of a target user, perform time series analysis on the operation time series data, and identify repeated operation patterns in the operation time series data; Determine an operation sequence in the repeated operation mode whose occurrence frequency exceeds a preset threshold as a candidate task chain; Calculating task similarity for the candidate task chains, merging candidate task chains with similarity greater than a preset similarity threshold, and obtaining a standard work task chain for the target user, wherein the task similarity is determined by calculating the matching degree of the operation type, operation object, and operation time of each task node in the candidate task chain; Marking each task node of the standard work task chain with a role authority combination corresponding to the task node; Obtain the role authority combination corresponding to each task node in the standard work task chain; Analyze the role authority combination of adjacent task nodes in the standard work task chain to generate an authority correlation matrix; Determine the permission inheritance relationship between task nodes according to the permission association matrix, and generate an initial permission dependency graph; Analyzing the permission transfer relationship between non-adjacent task nodes according to the initial permission dependency graph, wherein the permission transfer relationship indicates that permissions can be transferred through intermediate task nodes; The permission transfer relationship is updated to the initial permission dependency graph to form a task chain permission dependency graph, wherein the task chain permission dependency graph includes the permission inheritance and transfer relationship between each task node, and the role permission combination includes the viewing permission level, data operation permission and timeliness constraint; According to the current task node position of the target user, the next task operation of the target user is predicted in the task chain permission dependency graph, and the target role permission corresponding to the next task operation is preloaded; Determine preset dynamic desensitization rules for target sensitive data according to the target role permissions; During the data access process of the target sensitive data, the target sensitive data is desensitized according to the preset dynamic desensitization rules.
2. The method according to claim 1, characterized in that The step of predicting the next task operation of the target user in the task chain authority dependency graph according to the current task node position of the target user specifically includes: Obtain the historical operation track of the target user in the task chain permission dependency graph; Analyze the association pattern between the task nodes in the historical operation trajectory to obtain the association relationship between the task nodes; Determine a set of predicted task operations that the target user may perform according to the association relationship and the task node position where the target user is currently located; Based on the historical execution frequency of each of the task nodes in the predicted task operation set, the task node with the highest execution probability is selected as the next task operation.
3. The method according to claim 1, characterized in that After the step of desensitizing the target sensitive data according to the preset dynamic desensitization rule during the data access process of the target sensitive data, the method further includes: Obtaining historical audit log data of the target user within a preset time period, the historical audit log data including the target user's access time, access frequency, access location, data operation type to the target sensitive data, and task node conversion records in the task chain permission dependency graph; Training a machine learning model based on the historical audit log data to obtain a normal access behavior benchmark model of the target user; The real-time access behavior of the target user is evaluated according to the normal access behavior benchmark model. When the deviation between the real-time access behavior and the normal access behavior exceeds a preset threshold, the desensitization level in the preset dynamic desensitization rule is adjusted to a preset level.
4. The method according to claim 3, characterized in that After the step of training the machine learning model based on the historical audit log data to obtain a normal data access behavior benchmark model of the target user, the method further includes: Collecting the current access data of the target user in real time, and inputting it into the normal data access behavior benchmark model for calculation to obtain the abnormality score of the current access behavior; The access permission to the target sensitive data is adjusted according to the abnormality degree score.
5. A dynamic desensitization system for sensitive data in a cloud environment, characterized in that: The sensitive data dynamic desensitization system in the cloud environment includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the sensitive data dynamic desensitization system in the cloud environment to execute the method described in any one of claims 1-4.
6. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a dynamic desensitizing system for sensitive data in a cloud environment, the dynamic desensitizing system for sensitive data in a cloud environment executes the method as described in any one of claims 1-4.
7. A computer program product, characterized in that When the computer program product runs on a dynamic desensitization system for sensitive data in a cloud environment, the dynamic desensitization system for sensitive data in a cloud environment executes the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Real-time data migration method and computer readable medium
CN118193502A