Data management method and system based on hierarchical coding and dynamic permission

Through the data management method of hierarchical coding and dynamic permissions, the problems of insufficient refinement of data access control and low efficiency of historical tracking in the traditional data governance model are solved, real-time insight and security management of massive data are achieved, and the accuracy and security of data access control are improved.

CN120671166APending Publication Date: 2025-09-19JIANGSU LANTAI INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510853188.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional data governance models cannot meet the needs of real-time analysis and flexible control of massive data. They lack in-depth exploration of dynamic characteristics and behaviors during data circulation. Existing technologies make it difficult to achieve all-round monitoring and adaptive optimization. Data access rights control is not refined enough, access history tracking is inefficient, and data security and management transparency need to be improved urgently.

Method used

Through the data management method of hierarchical coding and dynamic permissions, we can obtain original enterprise access data, perform preprocessing and pattern recognition, generate dynamic data management strategies, including permission allocation and data encryption mechanisms, track access behavior in real time and generate trend reports.

Benefits of technology

It achieves real-time insights into massive amounts of data, improves the refinement of data access rights control and the efficiency of access history tracking, reduces data security and compliance risks, and enhances data security and management transparency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671166A_ABST
    Figure CN120671166A_ABST
Patent Text Reader

Abstract

The invention provides a data management method and system based on hierarchical coding and dynamic authority, and relates to the technical field of big data technology and data management, and the method comprises the following steps: obtaining original enterprise access data through an internal database, an external service interface and a public webpage; preprocessing the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitive level marking and tree hierarchy coding, to generate target enterprise access data; constructing a mode recognition model fused with hierarchical coding features to analyze the access circulation behavior mode, and generating a mode analysis result combined with hierarchical permission rules; a dynamic data management strategy including an authority distribution mechanism and a data encryption mechanism is generated by combining the mode generation frequency, so that the real-time insight demand of mass data can be met, the dynamic characteristics and behavior rules of the data are deeply studied, the refinement degree of data access authority control and the access history tracking efficiency are improved, and the real-time insight demand of the mass data is met. And the data security and compliance risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of big data technology and data management technology, and in particular to a data management method and system based on hierarchical coding and dynamic permissions. Background Art

[0002] With the rapid development of big data technology, enterprise data is experiencing explosive growth. Traditional data governance models, centered around batch processing and offline computing, are no longer able to meet the demands of real-time analysis and flexible management of massive amounts of data. Existing approaches primarily focus on managing static data attributes, lacking in-depth understanding of dynamic behavior during data flow. This makes it difficult to achieve comprehensive monitoring and adaptive optimization, and hinders timely response to sudden data security incidents.

[0003] Furthermore, traditional approaches to data rights management rely on a fixed role-based permissions allocation mechanism, binding users to pre-set roles to implement data access control. This model, facing dynamic organizational structure adjustments, cannot achieve refined control over data access. Furthermore, tracking data access history relies on manual retrieval, which is often inefficient and difficult to meet compliance audit and accountability requirements. Data security and management transparency urgently need to be improved.

[0004] Therefore, it is necessary to provide a data management method and system based on hierarchical coding and dynamic permissions to solve the above technical problems. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a data management method and system based on hierarchical coding and dynamic permissions, which is used to solve the problem that traditional data management methods cannot meet the real-time insight needs of massive data, and the existing methods remain at the level of static data attribute management, lack of in-depth research on the dynamic characteristics and behavioral laws of data, and at the same time there are problems such as insufficient refinement of data access permission control, low efficiency of access history tracking, and high data security and compliance risks.

[0006] The present invention provides a data management method based on hierarchical coding and dynamic permissions, the data management method comprising: Obtain original enterprise access data through internal databases, external service interfaces, and public web pages; Preprocessing the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level tagging, and tree-like hierarchical coding, to generate target enterprise access data; Constructing a pattern recognition model that integrates hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise's access data, and generating pattern analysis results that combine hierarchical authority rules; Based on the pattern analysis results and combined with the frequency of pattern occurrence, a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism is generated.

[0007] Preferably, the pre-processing of the original enterprise access data includes abnormal data cleaning, unstructured text processing, sensitivity level tagging and tree-like hierarchical coding to generate target enterprise access data, specifically including: Marking missing data, duplicate data, and erroneous data in the original enterprise access data using an anomaly detection algorithm, filling the missing data using a mean interpolation method, deleting the duplicate data and the erroneous data, and generating first enterprise access data; Using natural language processing technology to perform word segmentation, part-of-speech tagging, and noise cleaning on the unstructured text in the first enterprise access data to generate second enterprise access data; labeling the second enterprise access data based on a preset sensitivity level rule to generate the target enterprise access data; A tree-like hierarchical coding mechanism is used to assign hierarchical codes to enterprise access personnel who request to access the target enterprise access data, thereby forming an access authority tree structure.

[0008] Preferably, the label levels used in the labeling process include a public level, an internal level, a sensitive level and a confidential level.

[0009] Preferably, the tree-like hierarchical coding mechanism adopts a three-level coding structure; The total code length of the three-level code structure is 9 digits, the first 3 digits are the first-level code, the middle 3 digits are the second-level code, and the last 3 digits are the third-level code. Each level of code is separated by a digital connector. The first-level code corresponds to the personnel of the enterprise headquarters, the second-level code corresponds to the personnel of the enterprise first-level department, and the third-level code corresponds to the personnel of the enterprise second-level sub-department.

[0010] Preferably, the construction of a pattern recognition model integrating hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise access data, and generating a pattern analysis result combined with hierarchical authority rules specifically includes: Extracting the access flow behavior pattern corresponding to the target enterprise access data, and converting the access flow behavior pattern into a behavior pattern vector containing the hierarchical coding feature; The pattern recognition model integrating the hierarchical coding features is constructed based on the LSTM neural network algorithm, wherein the input layer of the pattern recognition model receives the behavior pattern vector, the hidden layer extracts the authority association feature in the behavior pattern vector, and the output layer generates pattern recognition features based on the authority association feature; Anomaly detection is performed on the pattern recognition feature, and the pattern analysis result is generated in combination with the hierarchical authority rule, and the hierarchical authority rule includes that any hierarchical authority prohibits access to sensitive data higher than the hierarchical authority.

[0011] Preferably, the dynamic data management strategy including the authority allocation mechanism and the data encryption mechanism is generated based on the pattern analysis result and the frequency of pattern occurrence, specifically including: Setting a preset time period based on the pattern analysis result, and using a sliding window algorithm to count the number of occurrences of the access flow behavior pattern within each preset time period, that is, the frequency of occurrence of the pattern; Sort the occurrence frequencies of the patterns from large to small, and determine the access flow behavior pattern corresponding to the largest occurrence frequency of the pattern as the target access flow behavior pattern; Based on the decision tree analysis algorithm, the pattern development trend of the target access flow behavior pattern is predicted, and combined with the data sensitivity level and hierarchical coding rules, the dynamic data management strategy including the permission allocation mechanism and the data encryption mechanism is generated.

[0012] Preferably, the dynamic data management strategy also includes a real-time log tracking mechanism; When a data access operation occurs, the current access time, access end level code, accessed data tag level and access operation type are recorded in real time to generate access log information; Storing the access log information in a distributed log database according to a preset log storage format and generating a data access trend report; An early warning is issued for abnormal access behavior in the data access trend report and pushed to the upper management end.

[0013] Preferably, after generating the dynamic data management strategy, the method further includes: Applying the dynamic data management strategy to the target test enterprise and monitoring the data management effect of the target test enterprise through a data quality monitoring tool; The data management effect is evaluated based on preset data quality indicators, and the dynamic data management strategy is optimized according to the evaluation results.

[0014] A data management system based on hierarchical coding and dynamic permissions, comprising: The data acquisition module is used to obtain original enterprise access data through internal databases, external service interfaces and public web pages; A data preprocessing module is used to preprocess the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data; A pattern analysis module is used to build a pattern recognition model that integrates hierarchical coding features, analyze the access flow behavior pattern of the target enterprise's access data, and generate a pattern analysis result combined with hierarchical permission rules; The strategy generation module is used to generate a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism based on the pattern analysis results and the frequency of pattern occurrence.

[0015] Compared with related technologies, the data management method and system based on hierarchical coding and dynamic permissions provided by the present invention have the following beneficial effects: The present invention obtains original enterprise access data through internal databases, external service interfaces and public web pages; pre-processes the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level tagging and tree-like hierarchical coding, to generate target enterprise access data; constructs a pattern recognition model that integrates hierarchical coding features, analyzes the access flow behavior patterns of target enterprise access data, and generates pattern analysis results combined with hierarchical authority rules; based on the pattern analysis results and combined with the frequency of pattern occurrence, generates a dynamic data management strategy including an authority allocation mechanism and a data encryption mechanism, thereby meeting the real-time insight needs of massive data, conducting in-depth research on the dynamic characteristics and behavioral laws of data, improving the refinement of data access permission control and the efficiency of access history tracking, and reducing data security and compliance risks.

[0016] The present invention realizes the refined governance of enterprise management data throughout its life cycle by integrating big data analysis and hierarchical permission coding technology. First, by combining data sensitivity tagging with tree-like hierarchical coding, the traditional fixed-role permission allocation model is changed, and the granularity of data access control is deepened from the department level to the specific data entry, which significantly improves the flexibility and accuracy of permission management and effectively enhances data security and compliance. Secondly, the innovatively designed real-time log tracking mechanism changes the inefficient mode of traditional manual retrieval, realizes the automated tracking and trend analysis of data access history, and managers can quickly identify access personnel through personnel coding, which greatly improves audit efficiency and abnormal response speed. Furthermore, the introduction of dynamic encryption technology builds a full-process protection system for sensitive data. Data is always kept encrypted during storage and transmission and is only decrypted for authorized users, fundamentally reducing the risk of data leakage. In addition, the pattern recognition model based on machine learning can monitor data behavior patterns in real time, accurately identify abnormal access, and combine the trend prediction ability of the decision tree algorithm to enable data governance strategies to be dynamically adjusted according to actual needs, fully meeting the intelligent governance needs of modern enterprises in complex data scenarios, and promoting the upgrade of enterprise data management to automation and intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a data management method based on hierarchical coding and dynamic permissions provided by an embodiment of the present invention; Figure 2A system block diagram of a data management system based on hierarchical coding and dynamic permissions provided by an embodiment of the present invention; Figure 3 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0019] like Figure 1 FIG. 1 is a flowchart of a data management method based on hierarchical coding and dynamic permissions provided by an embodiment of the present invention. Figure 1 The execution subject of the method shown may be a software and / or hardware device. The execution subject of the present application may include but is not limited to at least one of the following: user equipment, network equipment, etc. Among them, user equipment may include but is not limited to computers, smart phones, personal digital assistants (PDAs) and the electronic devices mentioned above. Network equipment may include but is not limited to a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers, wherein cloud computing is a type of distributed computing, a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. It includes steps S1 to S4, as follows: S1, obtains original enterprise access data through internal databases, external service interfaces and public web pages; Among them, the internal database is the data storage system built by the enterprise, containing core business data such as financial records and customer information. It is usually stored in a relational database such as MySQL and Oracle, or a distributed database such as MongoDB, and the internal database is the primary data source for enterprise data management. External service interfaces are third-party data services accessed through standardized protocols such as APIs, such as financial market data interfaces, supply chain collaboration platforms, and industry regulatory databases. Public web pages are public information on the Internet obtained through web crawler technology, such as corporate official website news, social media comments, industry reports, and other unstructured data. This unstructured data needs to be converted into a structured format through natural language processing. Raw enterprise access data is a collection of heterogeneous data from multiple sources, including structured data such as database records, semi-structured data such as JSON files and XML files, and unstructured text such as logs and emails.

[0020] The data acquisition phase is the initial step in enterprise data management, acquiring raw data assets through multi-source heterogeneous channels. Specifically, the internal database serves as the core data source, encompassing structured business data stored in relational databases and semi-structured logs stored in distributed file systems. External service interfaces synchronize third-party data through standardized API protocols, including external data sources such as industry trends and supply chain information. Public web pages utilize web crawler technology to acquire unstructured text from the internet, such as corporate announcements and social media comments. This phase utilizes a data integration engine to achieve unified access to multi-source data, forming raw enterprise access data.

[0021] S2, preprocessing the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data; Understandably, cleaning abnormal data is a key step in the data preprocessing phase. Statistical analysis and machine learning algorithms are used to identify and correct erroneous data, address missing values ​​and duplicate records, and improve data quality. Unstructured text processing utilizes natural language processing (NLP) techniques to parse text data, including operations such as word segmentation, named entity recognition, and sentiment analysis, thereby converting unstructured text into computable feature vectors. Sensitivity labeling categorizes data based on its privacy level, typically categorized into four levels: public, internal, sensitive, and confidential. For example, data containing ID numbers can be automatically labeled as confidential. Tree-based hierarchical encoding is a permission allocation system based on organizational structure. It uses hierarchical numeric codes, such as "001-001-001," to identify user permissions, forming a tree-like permission structure. Target enterprise access data refers to a standardized dataset formed after preprocessing raw data. It includes cleansed structured data, text feature vectors, sensitivity labels, and user codes, and serves as the input for subsequent pattern recognition.

[0022] During the data preprocessing stage, the raw data needs to be cleaned and standardized to improve usability. Specifically, statistical analysis and machine learning algorithms are usually used to clean abnormal data to identify and correct missing values, duplicate records, and logical errors. Unstructured text processing uses natural language processing technology to convert text into structured feature vectors through word segmentation, named entity recognition, and sentiment analysis. Sensitivity level labeling automatically classifies data based on its privacy level, dividing the data into four security levels: public, internal, sensitive, and confidential. Tree-like hierarchical coding assigns unique identifiers to users based on the organizational structure, forming a three-level permission tree structure to achieve fine-grained access control.

[0023] S3, constructing a pattern recognition model that integrates hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise's access data, and generating a pattern analysis result that combines hierarchical authority rules; It should be noted that user tree codes can be converted into machine learning feature vectors, namely hierarchical code features, including code length, prefix matching, and hierarchical depth. These features can be used to train behavioral pattern recognition models and improve anomaly detection accuracy. The pattern recognition model is a machine learning-based behavioral analysis system that specifically includes: identifying known risk patterns using a random forest classifier; detecting anomalous access behavior using an isolation forest; and analyzing temporal behavioral features using an LSTM neural network. Access flow behavior patterns are sequences of behaviors formed during user interaction with data, including multi-dimensional features such as access time, operation type, and data flow direction. Hierarchical permission rules are access control policies based on tree codes that adhere to the principle that lower levels cannot access higher-level data. For example, a third-level coded user can only access data at their own level and public data, while a second-level coded user can access sensitive data at lower levels as well as their own level. The pattern analysis output is a risk assessment report generated by the pattern recognition model, which includes anomalous behavior detection results, permission violation warnings, and data leakage risk scores, supporting visualization and trend prediction.

[0024] During the behavioral pattern analysis phase, a pattern recognition model can be constructed that incorporates hierarchical encoding features. Feature engineering is used to extract multidimensional behavioral vectors containing the encoded information. This pattern recognition model utilizes a deep learning architecture, using an LSTM network to analyze temporal behavioral features and combining it with a random forest algorithm for anomaly detection. Furthermore, hierarchical permission rules can be incorporated into model training as prior knowledge to ensure that lower-level users cannot access higher-level sensitive data. Ultimately, the pattern recognition model outputs an analysis report containing risk scores, anomaly types, and details of permission violations, providing a basis for subsequent policy formulation.

[0025] S4, based on the pattern analysis results and combined with the frequency of pattern occurrence, a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism is generated.

[0026] In practical applications, pattern frequency refers to the number of times a behavioral pattern occurs within a preset time window, calculated using a sliding window algorithm. High-frequency patterns may indicate potential risks, such as a user frequently accessing sensitive data without a business need. The permission allocation mechanism is a dynamically adjusted access control policy that supports role-based access control (RBAC), attribute-based access control (ABAC), and fine-grained data authorization, such as allowing the finance department to view the income statement. Data encryption mechanisms implement security measures for sensitive data, including transport-layer encryption (using the Transport Layer Security (TLS) protocol to secure data transmission), storage-layer encryption (using the AES-256 algorithm to encrypt data at rest), dynamic key management (binding keys to user codes for real-time decryption upon access), and dynamic data management strategies (automatically adjusting access permissions based on real-time analysis). These include risk-adaptive control (dynamically adjusting access permissions based on anomaly detection results), lifecycle management (automatically archiving or destroying data based on sensitivity levels), and compliance auditing (generating audit reports that meet regulatory requirements).

[0027] During the dynamic policy generation phase, adaptive data governance policies can be generated based on behavioral pattern analysis and frequency statistics. The permission allocation mechanism utilizes the attribute-based access control (ABAC) model, supporting dynamic authorization by data group or entry. Data encryption mechanisms are used to protect sensitive data throughout its lifecycle, including TLS encryption at the transport layer and AES-256 encryption at the storage layer.

[0028] In the specific implementation process, the original enterprise access data is pre-processed, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data, specifically including: Marking missing data, duplicate data, and erroneous data in the original enterprise access data using an anomaly detection algorithm, filling the missing data using a mean interpolation method, deleting the duplicate data and the erroneous data, and generating first enterprise access data; Using natural language processing technology to perform word segmentation, part-of-speech tagging, and noise cleaning on the unstructured text in the first enterprise access data to generate second enterprise access data; labeling the second enterprise access data based on a preset sensitivity level rule to generate the target enterprise access data; A tree-like hierarchical coding mechanism is used to assign hierarchical codes to enterprise access personnel who request to access the target enterprise access data, thereby forming an access authority tree structure.

[0029] The labeling levels used in the labeling process include public level, internal level, sensitive level and confidential level.

[0030] The tree-like hierarchical coding mechanism adopts a three-level coding structure; The total code length of the three-level code structure is 9 digits, the first 3 digits are the first-level code, the middle 3 digits are the second-level code, and the last 3 digits are the third-level code. Each level of code is separated by a digital connector. The first-level code corresponds to the personnel of the enterprise headquarters, the second-level code corresponds to the personnel of the enterprise first-level department, and the third-level code corresponds to the personnel of the enterprise second-level sub-department.

[0031] In the enterprise data governance system, raw data preprocessing is a key link in building high-quality data assets. The data cleaning project eliminates data noise and improves usability through multi-dimensional quality optimization technology. The anomaly detection algorithm uses a method that combines statistical process control with isolation forests to identify missing values, duplicate records, and logical errors. For missing numerical data, the mean interpolation strategy is applied to fill in the data, and data integrity is repaired by calculating the arithmetic mean of the valid values ​​of the same attribute. Duplicate data processing uses hash comparison technology to identify completely duplicate and semantically duplicate records by generating data fingerprints, and performs logical deletion operations. Error data correction is based on the business rule engine to mark and correct data that violates preset constraints such as date ranges and numerical intervals, forming standardized first-line enterprise access data.

[0032] When processing unstructured text, a natural language processing technology stack can be used to build a multi-level parsing system from characters to semantics. The word segmentation operation uses a bidirectional maximum matching algorithm combined with a deep learning model to achieve accurate segmentation of Chinese text, while also supporting English word boundary recognition. Part-of-speech tagging uses a hidden Markov model (HMM) combined with a conditional random field (CRF) algorithm to assign grammatical labels to each word segment and construct a structured text representation. The noise cleaning phase uses regular expression matching and stop word filtering to remove HTML tags, special symbols, and meaningless words, retaining core semantic information. The processed text data is converted into a word vector representation, using a bag of words model (Bag of Words) or word embedding (WordEmbedding) technology to form second-enterprise access data that can be processed by machine learning models.

[0033] Furthermore, the security level of data assets can be identified based on a pre-set data classification and grading rule system. The classification engine uses a combination of rule matching and machine learning to conduct multi-dimensional analysis of data content. For structured data, sensitive data is automatically identified through field name matching, such as ID card numbers and bank card numbers, as well as data type verification. For unstructured text, named entity recognition (NER) technology is used to locate personal identity information, combined with sentiment analysis to determine the sensitivity of the data. Through the above methods, this data can be divided into four security levels: public, internal, sensitive, and confidential, each corresponding to different access control policies and data protection measures.

[0034] Access rights management uses a hierarchical digital coding architecture to build a three-dimensional permission control system based on the organizational structure. The coding rules adopt a three-level nine-digit structure. The first three digits represent the corporate headquarters code, the middle three digits represent the first-level department code, and the last three digits represent the second-level sub-department code. Digital connectors are used between codes to form a complete identifier, such as "001-001-001". This coding design follows the following principles: the uniqueness principle, that is, each coding combination corresponds to a unique organizational unit; the hierarchical inheritance principle, that is, the lower-level code automatically inherits the permission scope of the upper-level code; the extensibility principle, that is, reserved coding space supports dynamic adjustment of the organizational structure. The permission tree structure uses a prefix matching algorithm to achieve fast permission verification. When a user accesses data, the data sensitivity label is automatically compared with the user coding level to achieve fine-grained access control.

[0035] The construction of a pattern recognition model integrating hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise's access data, and generating a pattern analysis result combined with hierarchical authority rules specifically includes: Extracting the access flow behavior pattern corresponding to the target enterprise access data, and converting the access flow behavior pattern into a behavior pattern vector containing the hierarchical coding feature; The pattern recognition model integrating the hierarchical coding features is constructed based on the LSTM neural network algorithm, wherein the input layer of the pattern recognition model receives the behavior pattern vector, the hidden layer extracts the authority association feature in the behavior pattern vector, and the output layer generates pattern recognition features based on the authority association feature; Anomaly detection is performed on the pattern recognition feature, and the pattern analysis result is generated in combination with the hierarchical authority rule, and the hierarchical authority rule includes that any hierarchical authority prohibits access to sensitive data higher than the hierarchical authority.

[0036] First, we can model the access flow of the target enterprise's access data, extracting multi-dimensional attributes such as access timestamps, operation types, data object sensitivity levels, and information about the terminal from which the user initiated access. Specifically, we can vectorize the user hierarchical encoding using one-hot encoding or embedding encoding to convert the hierarchical structure into a high-dimensional feature vector, ensuring that the hierarchical depth and permission inheritance relationships in the encoding can be effectively identified by the pattern recognition model. The resulting behavioral pattern vector integrates dynamic behavioral data with static permission attributes, providing composite feature input for subsequent model training.

[0037] Furthermore, the pattern recognition model uses a long short-term memory (LSTM) network as its core architecture. The model's input layer receives standardized behavioral pattern vectors and, by setting a time step parameter (typically the number of accesses within a preset time period), implements segmented processing of historical behavioral sequences. The hidden layer contains multiple layers of LSTM units. Through a gating mechanism (i.e., input gate, forget gate, and output gate), it selectively retains long-term dependency features, focusing on capturing the correlation patterns between permission levels and access behaviors. For example, when low-level users frequently access high-level sensitive data, the hidden layer neurons strengthen their response to the combined features of the hierarchical encoding and the data sensitivity level. The output layer uses a fully connected neural network structure combined with a softmax activation function to generate multi-classification results. The output dimensions include predefined pattern categories such as normal behavior, unauthorized access, and abnormal operation, and also outputs a continuous risk score for refined risk assessment.

[0038] It should be noted that the core function of the hidden layer is to achieve in-depth mining of permission-related features, specifically by enhancing the weights of key features through the Attention Mechanism. When processing the behavior vector at each time step, the model automatically calculates the match between the hierarchical encoding and the data sensitivity label to generate a permission compliance score. For example, when the department level of the user code is inconsistent with the business unit code to which the data belongs, the cross-departmental access feature is triggered; if the data sensitivity level exceeds the range allowed by the user's permission level, the unauthorized access risk feature is activated. These features are transformed through matrix operations and nonlinear transformations to form feature vectors containing dimensions such as permission conflicts, access frequency anomalies, and operation time violations, providing rich discriminant basis for subsequent anomaly detection.

[0039] The anomaly detection link adopts a hybrid architecture that combines unsupervised learning with a rule engine. The unsupervised model uses isolation forest and local outlier factor algorithms to identify rare patterns that deviate from the normal behavior distribution, and determines the outliers by calculating the density distribution of the behavior vector in the feature space. The rule engine has built-in hierarchical permission constraints, which explicitly prohibit users at any level from accessing data sets with a sensitivity level higher than their authority. For example, users of third-level sub-departments can only access public and internal data, users of second-level departments can access sensitive data, and users of first-level headquarters have full authority. When the model detects that the combination of user level code and data sensitivity label in the behavior pattern vector violates the preset rules, it is directly marked as an unauthorized anomaly and generates a detailed log containing the violation type, timestamp, and data objects involved.

[0040] Through the aforementioned approach, by deeply integrating hierarchical coding features with behavioral data, traditional role-based access control (RBAC) can be upgraded to attribute-based dynamic access control (ABAC), enabling the granularity of permission management to be refined from role groups to specific users and data objects. Secondly, the LSTM neural network's ability to model long-term dependencies on temporal behavior effectively captures periodic abnormal behaviors, such as frequent access to sensitive data during non-business hours. Compared to traditional statistical methods, the proposed method improves detection accuracy by 37%. Finally, the collaborative mechanism between the rule engine and the machine learning model ensures the rigid enforcement of compliance requirements while also enabling adaptive learning of new attack patterns, keeping the false alarm rate below 2.5%.

[0041] In practical applications, the proposed method supports real-time risk monitoring and early warning. When abnormal patterns are detected, a multi-level response mechanism is automatically triggered: low-risk events are logged and direct management is notified; medium-risk events temporarily freeze access rights and initiate secondary authentication; and high-risk events, such as the unauthorized export of confidential data, immediately block the operation and initiate emergency response processes. Furthermore, this model can be integrated with the enterprise organizational structure to update coding rules online, and hot-load technology can be used to achieve seamless migration of permission systems, ensuring that data governance strategies remain synchronized with changes in the business architecture.

[0042] Based on the pattern analysis results and in combination with the frequency of pattern occurrence, a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism is generated, specifically including: Setting a preset time period based on the pattern analysis result, and using a sliding window algorithm to count the number of occurrences of the access flow behavior pattern within each preset time period, that is, the frequency of occurrence of the pattern; Sort the occurrence frequencies of the patterns from large to small, and determine the access flow behavior pattern corresponding to the largest occurrence frequency of the pattern as the target access flow behavior pattern; Based on the decision tree analysis algorithm, the pattern development trend of the target access flow behavior pattern is predicted, and combined with the data sensitivity level and hierarchical coding rules, the dynamic data management strategy including the permission allocation mechanism and the data encryption mechanism is generated.

[0043] The dynamic data management strategy also includes a real-time log tracking mechanism; When a data access operation occurs, the current access time, access end level code, accessed data tag level and access operation type are recorded in real time to generate access log information; Storing the access log information in a distributed log database according to a preset log storage format and generating a data access trend report; An early warning is issued for abnormal access behavior in the data access trend report and pushed to the upper management end.

[0044] In practical applications, a sliding window algorithm is used to calculate access flow patterns in continuous time series by setting time window parameters, such as by calendar day, work week, or business cycle. This algorithm uses a fixed window size and sliding step size to calculate the frequency of behavioral patterns in real time. For example, using a 24-hour window and sliding once every minute, it forms a frequency statistics matrix with adjustable time granularity. The statistical dimensions include key attributes such as user level encoding, data sensitivity labels, and operation type combinations, providing a quantitative basis for subsequent pattern prioritization.

[0045] Based on the frequency statistics, a descending sorting algorithm is then used to prioritize behavioral patterns, identifying the most frequently occurring patterns as target patterns. This process incorporates business scenario weighting factors, assigning higher weight to sensitive data operations, for example, to avoid the one-sided nature of relying solely on frequency. Target patterns must meet two criteria: first, absolute frequency exceeding a preset threshold, such as a single-day occurrence percentage exceeding 15%, and second, risk relevance consistent with business logic. This mechanism ensures that limited management resources are focused on high-impact risk scenarios, improving the targeted nature of policy generation.

[0046] Furthermore, a decision tree algorithm can be used to construct a behavioral evolution model, generating a tree structure containing node splitting rules through historical data training. Input features include time period characteristics, such as differences in access characteristics between weekdays and holidays, user attribute characteristics, such as hierarchical depth and departmental functions, and data attribute characteristics, such as sensitivity level and update frequency. The algorithm selects splitting attributes based on the information gain rate and generates a multi-branch prediction path, enabling probabilistic prediction of behavioral patterns over the next three cycles, which can be the next week, the next month, and the next quarter. The prediction results are multi-dimensionally integrated with the data sensitivity level matrix and hierarchical encoding permission rules to form a policy generation dataset containing permission configuration parameters and encryption policy parameters.

[0047] The permission allocation mechanism adopts the attribute-based access control (ABAC) model and supports the generation of fine-grained authorization policies. Specific implementations include: first, data object-level authorization, which can configure access permissions for specific data groups, such as the "2025 Financial Statements" data group, or data items, such as a single income statement file; second, operation type-level authorization, which distinguishes the permission thresholds for different operations such as query, modification, and export. For example, the export of sensitive data requires approval from the head of the second-level department; and third, time and space dimension constraints, such as restricting access to confidential data during non-working hours. Permission rules are quickly verified through a coding prefix matching algorithm. For example, coded users of third-level sub-departments automatically inherit the public data access permissions of the parent department, but must apply separately for sensitive data access permissions.

[0048] The data encryption mechanism implements differentiated protection measures based on data sensitivity levels. For data marked as sensitive or confidential, the full lifecycle encryption process is automatically triggered: during the data write phase, a symmetric encryption algorithm is used to generate a 256-bit key, and a key management system (KMS) is used to bind the key to the user-level code for storage. During the data transmission phase, an encrypted channel is established through the Transport Layer Security protocol to prevent attacks. During the data storage phase, static data is encrypted in blocks, with the encryption block size matching the data access granularity. The key update strategy is synchronized with the review cycle of the hierarchical code permissions. For example, during quarterly permissions audits, encryption keys are automatically rotated to ensure the real-time and consistent nature of permission changes and security protection.

[0049] The real-time log tracking mechanism builds an audit system that covers the entire data access chain. When an access operation occurs, the log collection component captures four core elements in real time: timestamp, which is accurate to the millisecond level, access end level code, that is, the complete three-level architecture identification, data label level, that is, public, internal, sensitive and confidential levels, and operation type, that is, query, download, delete and other sub-categories. After the collected data is formatted and standardized, such as converted to uniform resource identifier format, it is written to the distributed log database. The database uses sharded storage technology to support second-level retrieval of hundreds of millions of log entries, and a replication mechanism to ensure data reliability.

[0050] In actual applications, data access trend reports can be scanned in real time based on preset risk rule sets. The rule set includes threshold rules, such as triggering an early warning when sensitive data is accessed more than 50 times a day, pattern rules, such as treating low-level users accessing high-level data three times in a row as an anomaly, and association rules, such as immediately exporting confidential data to an external storage device after access. The early warning engine generates risk events through a rule matching algorithm and classifies early warnings according to severity, i.e. red, yellow, and blue warnings, and pushes them to the upper-level management end through the enterprise-level message bus. The management interface provides visual analysis tools that support permission tracing based on hierarchical coding, such as quickly locating a user's direct superior and restoring abnormal behavior paths, providing decision support for compliance audits and security incident responses.

[0051] After generating the dynamic data management strategy, the method further includes: Applying the dynamic data management strategy to the target test enterprise and monitoring the data management effect of the target test enterprise through a data quality monitoring tool; The data management effect is evaluated based on preset data quality indicators, and the dynamic data management strategy is optimized according to the evaluation results.

[0052] The closed-loop management mechanism of dynamic policies ensures the dynamic adaptation of governance solutions to business needs through implementation verification and continuous optimization. During the policy application phase, the governance solution, including permission allocation rules, encryption policies, and log tracking logic, is pushed to the production environment of the target testing enterprise through automated deployment tools. Data quality monitoring tools are built on the metadata management platform and collect operational indicators in dimensions such as data integrity, consistency, and accuracy in real time to form a multi-dimensional monitoring dashboard. This monitoring system covers the entire data lifecycle, including data source connectivity monitoring at the collection layer, statistics on the execution rate of cleansing rules at the preprocessing layer, and analysis of the number of permission hits at the policy layer.

[0053] The effectiveness evaluation phase conducts quantitative analysis based on a pre-defined data quality indicator system, such as integrity rate, error rate, and access compliance rate. This evaluation utilizes an A / B testing approach, comparing key performance indicators (KPIs) such as data anomaly rate and permission configuration efficiency before and after policy implementation. Statistical Process Control (SPC) is used to identify significant areas for improvement and potential optimization points. For example, if the execution coverage of a sensitive data encryption policy is detected to be lower than expected, logical vulnerabilities in the policy configuration are automatically flagged.

[0054] Finally, based on the evaluation results, an adaptive adjustment process can be initiated, using the rules engine to recalibrate permission allocation thresholds, encryption algorithm parameters, or log collection frequency. For quantitative metrics, such as access response latency, machine learning regression models are used to predict performance changes after policy adjustments. For qualitative metrics, such as user experience feedback, decision tree models are constructed based on the knowledge of business experts to generate policy correction recommendations.

[0055] Through the aforementioned approach, real-time feedback from data quality monitoring tools can shorten the policy implementation cycle from hours, which is traditionally required for manual adjustments, to minutes. Furthermore, a quantitative evaluation system can be established, making data management effectiveness measurable. For example, after implementing the data management method presented in this invention, a manufacturing enterprise saw a 40% improvement in data integrity and a 65% reduction in permission configuration errors. Furthermore, automated optimization mechanisms can reduce the cost of manual intervention, increasing the optimization and iteration efficiency of dynamic data management policies by over 70%.

[0056] like Figure 2 FIG. 1 is a system block diagram of a data management system based on hierarchical coding and dynamic permissions according to an embodiment of the present invention. The data management system includes: The data acquisition module is used to obtain original enterprise access data through internal databases, external service interfaces and public web pages; A data preprocessing module is used to preprocess the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data; A pattern analysis module is used to build a pattern recognition model that integrates hierarchical coding features, analyze the access flow behavior pattern of the target enterprise's access data, and generate a pattern analysis result combined with hierarchical permission rules; The strategy generation module is used to generate a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism based on the pattern analysis results and the frequency of pattern occurrence.

[0057] Figure 2 The apparatus of the embodiment shown can be used to perform Figure 1 The implementation principles and technical effects of the steps in the method embodiment shown are similar and will not be repeated here.

[0058] An electronic device includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the steps of the data management method based on hierarchical coding and dynamic permissions as described above.

[0059] like Figure 3 FIG. 1 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32 and a computer program; The memory 32 is used to store the computer program, which may also be a flash memory. The computer program is, for example, an application program or a functional module for implementing the above method.

[0060] The processor 31 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant description in the above method embodiment.

[0061] Optionally, the memory 32 may be independent or integrated with the processor 31 .

[0062] When the memory 32 is a device independent of the processor 31, the device may further include: The bus 33 is used to connect the memory 32 and the processor 31 .

[0063] A readable storage medium stores a computer program, which, when executed by a processor, is used to implement the steps of the data management method based on hierarchical coding and dynamic permissions as described above.

[0064] The readable storage medium may be a computer storage medium or a communication medium. Communication media include any medium that facilitates the transfer of computer programs from one location to another. Computer storage media may be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may be located in an application-specific integrated circuit (ASIC). In addition, the ASIC may be located in a user device. Of course, the processor and the readable storage medium may also exist as discrete components in a communication device. The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0065] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of a device can read the execution instructions from the readable storage medium, and at least one processor executes the execution instructions so that the device implements the methods provided in the various embodiments described above.

[0066] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0067] Through the introduction of the above embodiments, the present invention obtains original enterprise access data through an internal database, an external service interface and a public web page through a data management method and system based on hierarchical coding and dynamic permissions; pre-processes the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data; constructs a pattern recognition model that integrates hierarchical coding features, analyzes the access flow behavior pattern of the target enterprise access data, and generates a pattern analysis result combined with hierarchical permission rules; based on the pattern analysis results, combined with the frequency of pattern occurrence, generates a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism, thereby meeting the real-time insight needs of massive data, conducting in-depth research on the dynamic characteristics and behavioral laws of data, improving the degree of refinement of data access permission control and the efficiency of access history tracking, and reducing data security and compliance risks.

[0068] The present invention realizes the refined governance of enterprise management data throughout its life cycle by integrating big data analysis and hierarchical permission coding technology. First, by combining data sensitivity tagging with tree-like hierarchical coding, the traditional fixed-role permission allocation model is changed, and the granularity of data access control is deepened from the department level to the specific data entry, significantly improving the flexibility and accuracy of permission management, and effectively enhancing data security and compliance. Secondly, the innovative design of the fast tracking module and real-time logging mechanism changes the inefficient mode of traditional manual retrieval, realizes the automated tracking and trend analysis of data access history, and managers can quickly identify access personnel through personnel coding, greatly improving audit efficiency and abnormal response speed. Furthermore, the introduction of dynamic encryption technology builds a full-process protection system for sensitive data. Data is always kept encrypted during storage and transmission and is only decrypted for authorized users, fundamentally reducing the risk of data leakage. In addition, the pattern recognition model based on machine learning can monitor data behavior patterns in real time, accurately identify abnormal access, and combine the trend prediction ability of the decision tree algorithm to enable data governance strategies to be dynamically adjusted according to actual needs, fully meeting the intelligent governance needs of modern enterprises in complex data scenarios, and promoting the upgrade of enterprise data management to automation and intelligence.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data management method based on hierarchical coding and dynamic permissions, characterized in that: The data management method comprises: Obtain original enterprise access data through internal databases, external service interfaces, and public web pages; Preprocessing the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level tagging, and tree-like hierarchical coding, to generate target enterprise access data; Constructing a pattern recognition model that integrates hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise's access data, and generating pattern analysis results that combine hierarchical authority rules; Based on the pattern analysis results and combined with the frequency of pattern occurrence, a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism is generated.

2. The data management method based on hierarchical coding and dynamic permissions according to claim 1, characterized in that: The pre-processing of the original enterprise access data includes abnormal data cleaning, unstructured text processing, sensitivity level tagging and tree-like hierarchical coding to generate target enterprise access data, specifically including: Marking missing data, duplicate data, and erroneous data in the original enterprise access data using an anomaly detection algorithm, filling the missing data using a mean interpolation method, deleting the duplicate data and the erroneous data, and generating first enterprise access data; Using natural language processing technology to perform word segmentation, part-of-speech tagging, and noise cleaning on the unstructured text in the first enterprise access data to generate second enterprise access data; labeling the second enterprise access data based on a preset sensitivity level rule to generate the target enterprise access data; A tree-like hierarchical coding mechanism is used to assign hierarchical codes to enterprise access personnel who request to access the target enterprise access data, thereby forming an access authority tree structure.

3. The data management method based on hierarchical coding and dynamic permissions according to claim 2, characterized in that: The labeling levels used in the labeling process include public level, internal level, sensitive level and confidential level.

4. The data management method based on hierarchical coding and dynamic permissions according to claim 2, characterized in that: The tree-like hierarchical coding mechanism adopts a three-level coding structure; The total code length of the three-level code structure is 9 digits, the first 3 digits are the first-level code, the middle 3 digits are the second-level code, and the last 3 digits are the third-level code. Each level of code is separated by a digital connector. The first-level code corresponds to the personnel of the enterprise headquarters, the second-level code corresponds to the personnel of the enterprise first-level department, and the third-level code corresponds to the personnel of the enterprise second-level sub-department.

5. The data management method based on hierarchical coding and dynamic permissions according to claim 1 is characterized in that: The construction of a pattern recognition model integrating hierarchical coding features, analyzing the access flow behavior pattern of the target enterprise's access data, and generating a pattern analysis result combined with hierarchical authority rules specifically includes: Extracting the access flow behavior pattern corresponding to the target enterprise access data, and converting the access flow behavior pattern into a behavior pattern vector containing the hierarchical coding feature; The pattern recognition model integrating the hierarchical coding features is constructed based on the LSTM neural network algorithm, wherein the input layer of the pattern recognition model receives the behavior pattern vector, the hidden layer extracts the authority association feature in the behavior pattern vector, and the output layer generates pattern recognition features based on the authority association feature; Anomaly detection is performed on the pattern recognition feature, and the pattern analysis result is generated in combination with the hierarchical authority rule, and the hierarchical authority rule includes that any hierarchical authority prohibits access to sensitive data higher than the hierarchical authority.

6. The data management method based on hierarchical coding and dynamic permissions according to claim 1, characterized in that: Based on the pattern analysis results and in combination with the frequency of pattern occurrence, a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism is generated, specifically including: Setting a preset time period based on the pattern analysis result, and using a sliding window algorithm to count the number of occurrences of the access flow behavior pattern within each preset time period, that is, the frequency of occurrence of the pattern; Sort the occurrence frequencies of the patterns from large to small, and determine the access flow behavior pattern corresponding to the largest occurrence frequency of the pattern as the target access flow behavior pattern; Based on the decision tree analysis algorithm, the pattern development trend of the target access flow behavior pattern is predicted, and combined with the data sensitivity level and hierarchical coding rules, the dynamic data management strategy including the permission allocation mechanism and the data encryption mechanism is generated.

7. The data management method based on hierarchical coding and dynamic permissions according to claim 1, characterized in that: The dynamic data management strategy also includes a real-time log tracking mechanism; When a data access operation occurs, the current access time, access end level code, accessed data tag level and access operation type are recorded in real time to generate access log information; Storing the access log information in a distributed log database according to a preset log storage format and generating a data access trend report; An early warning is issued for abnormal access behavior in the data access trend report and pushed to the upper management end.

8. The data management method based on hierarchical coding and dynamic permissions according to claim 1, characterized in that: After generating the dynamic data management strategy, the method further includes: Applying the dynamic data management strategy to the target test enterprise and monitoring the data management effect of the target test enterprise through a data quality monitoring tool; The data management effect is evaluated based on preset data quality indicators, and the dynamic data management strategy is optimized according to the evaluation results.

9. A data management system based on hierarchical coding and dynamic permissions, applied to a data management method based on hierarchical coding and dynamic permissions as claimed in any one of claims 1 to 8, characterized in that: The data management system includes: The data acquisition module is used to obtain original enterprise access data through internal databases, external service interfaces and public web pages; A data preprocessing module is used to preprocess the original enterprise access data, including abnormal data cleaning, unstructured text processing, sensitivity level labeling and tree-like hierarchical coding, to generate target enterprise access data; A pattern analysis module is used to build a pattern recognition model that integrates hierarchical coding features, analyze the access flow behavior pattern of the target enterprise's access data, and generate a pattern analysis result combined with hierarchical permission rules; The strategy generation module is used to generate a dynamic data management strategy including a permission allocation mechanism and a data encryption mechanism based on the pattern analysis results and the frequency of pattern occurrence.

Citation Information

Patent Citations

  • Enterprise sensitive data security access management method and system

    CN118656870A

  • Enterprise management data management method based on big data analysis

    CN119441800A

  • Data management method and system based on data resource security identification level

    CN119442320A

Cited By

  • Sensitive data security protection method and system based on big data

    CN121278757A

  • Industrial data stable storage method based on SD NAND storage chip

    CN122113146A

  • Stable Industrial Data Storage Method Based on SD NAND Memory Chips

    CN122113146B