A personalized data governance optimization method based on behavioral time series analysis
By building a user-personalized data governance method through behavioral time series analysis, we solve the problems of inaccurate data classification and insufficient adaptability to multi-user environments in traditional data governance, implement personalized and dynamic data governance solutions, and improve the accuracy and security of data management.
Patent Information
- Application Number
- CN202510936085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Traditional data governance methods rely on static information and cannot capture the dynamic characteristics of user behavior, resulting in inaccurate data classification and difficulty in meeting personalized needs. They also lack support for multi-user environments and cannot achieve partition optimization and governance strategy sharing based on group behavior trends.
Through behavioral time series analysis, historical behavioral data of multiple user individuals are collected, and personalized feature vectors of operation and location behavior are constructed. Combined with the data hierarchical structure, individual and fusion data level distribution is generated, a data governance architecture is built, and personalized governance solutions are matched to dynamically adjust governance strategies.
It achieves accurate and dynamic data classification, adapts to personalized governance in multi-user environments, improves the intelligence and adaptability of data governance, and ensures efficient, secure and personalized data management.
Smart Images

Figure CN120429295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a personalized data governance optimization method based on behavior timing analysis. Background Art
[0002] In the digital age, data governance has become a critical link in ensuring data quality, security, and efficient utilization. Traditional data governance methods rely primarily on static information such as file type, directory location, and manual user annotations to manage data hierarchically. While these methods can achieve data classification to a certain extent, they lack a dynamic understanding and application of user behavior, making them difficult to adapt to complex and changing data environments and user needs.
[0003] Traditional data governance methods rely solely on static information and fail to capture the dynamic characteristics of user behavior, such as operation frequency, location migration, and usage scenarios. This results in inaccurate data grading and makes it difficult to meet the needs of personalized data governance. Furthermore, existing solutions often focus on single-user data governance, lacking effective support for multi-user shared environments and unable to achieve partition optimization and governance policy sharing based on group behavior trends. Existing methods often rely on manual configuration or fixed policy template invocation, lack intelligent governance solution generation mechanisms, and struggle to adapt to changes in user behavior. Summary of the Invention
[0004] The present invention provides a personalized data governance optimization method based on behavioral time series analysis to solve the technical problems in the existing technology of inaccurate data classification, inability to adapt to multi-user environments, and lack of personalization and dynamic adjustment capabilities of governance solutions, and achieves the technical effects of accurate and dynamic data classification, integrated governance of multi-user environments, and personalized and dynamically adjustable governance solutions.
[0005] The present invention provides a personalized data governance optimization method based on behavioral time series analysis, comprising:
[0006] In the target data environment, multiple user individuals are used as behavior collection objects, and historical behavior data is obtained according to the preset time window.
[0007] A time series traversal analysis is performed on the historical behavior data to establish the behavior personality of each user, wherein the behavior personality includes operation behavior personality and location behavior personality.
[0008] According to a preset data hierarchical structure, the behavioral personality is decomposed into components to obtain the individual data level distribution of each user individual, and multiple individual data level distributions are fused to obtain the fused data level distribution of the target data environment.
[0009] Combining the individual data level distribution and the fusion data level distribution, a data governance architecture for the target data environment is constructed.
[0010] According to the distribution of multiple individual data levels, corresponding data governance templates are matched from the preconfigured policy template library, and multiple personalized data governance solutions are constructed accordingly. Personalized data governance is performed in combination with the multiple personalized data governance solutions and the data governance architecture.
[0011] In a feasible implementation, the historical behavior data is subjected to a time-series traversal analysis to establish the behavior personality of each individual user, wherein the behavior personality includes an operation behavior personality and a location behavior personality, including:
[0012] Based on the historical behavior data, the user's individual operation behavior data is extracted to perform time series traversal analysis to obtain the file operation frequency distribution, common file type combinations, high-frequency operation sequence patterns and data creation-modification ratios, and corresponding operation behavior feature vectors are generated.
[0013] Based on the historical behavior data, the historical location behavior data of individual users is extracted to perform time series traversal analysis to obtain the main active locations, location switching frequency, location stay duration and location-operation mapping relationship, and corresponding location behavior feature vectors are generated.
[0014] The position behavior feature vector is used as the position behavior personality, and the operation behavior feature vector is used as the operation behavior personality, and they are associated and stored as the behavior personality.
[0015] In one feasible implementation, the data hierarchy is constructed based on data sensitivity, access frequency, data importance, and compliance rules, and includes:
[0016] When an individual user frequently operates on a data object during a high-frequency access period, the data object is mapped to a hot data level.
[0017] When a data object has not been accessed within a set silence threshold period, the data object is mapped to the cold data level.
[0018] When a data object satisfies any one of the conditions of being associated with a highly sensitive application behavior or being located in a preset sensitive geographical location, the data object is mapped to a high sensitivity level.
[0019] In a feasible implementation, the behavioral personality is decomposed into components according to a preset data hierarchical structure to obtain the individual data level distribution of each user, and multiple individual data level distributions are fused to obtain the fused data level distribution of the target data environment, including:
[0020] Taking data sensitivity, access frequency, importance and compliance as data classification features, the data classification structure is traversed to match the behavioral personality corresponding to each individual user with the corresponding data level, and a level matching result is constructed.
[0021] The frequency of each data level in the level matching result is counted, and a cluster analysis is performed with individual users as clustering dimensions to obtain the individual data level distribution of each individual user.
[0022] The individual data level distributions of multiple individual users are traversed, and secondary clustering analysis is performed with data level as a clustering dimension to determine the fused data level distribution of the target data environment.
[0023] In a feasible implementation, the individual data level distribution and the fusion data level distribution are combined to construct a data governance architecture for the target data environment, including:
[0024] A reference logical storage layer structure is constructed based on the fused data level distribution, and storage capacity and performance configuration are allocated, wherein the reference logical storage layer structure includes multiple logical storage layers, and each logical storage layer corresponds to a data level.
[0025] Based on the plurality of individual data level distributions, a user data distribution corresponding to each data level is determined.
[0026] According to the user data distribution, each of the logical storage layers is divided into a plurality of user individual storage blocks to form the data governance architecture.
[0027] In a feasible implementation, according to the multiple individual data level distributions, corresponding data governance templates are matched from a preconfigured policy template library, multiple personalized data governance solutions are constructed accordingly, and personalized data governance is performed in combination with the multiple personalized data governance solutions and the data governance architecture, including:
[0028] For any of the individual data level distributions, the included data levels are extracted to form a data level list.
[0029] Traverse the data level list, match the corresponding data governance template in the policy template library with the data level as the matching index, and form a data governance template cluster, wherein the data governance template includes data backup frequency, data backup rules and data synchronization time.
[0030] According to the data governance architecture, a mapping relationship between the data governance template cluster and the user's individual storage block is established to form the personalized data governance solution.
[0031] Repeatedly traverse multiple individual data level distributions to form multiple personalized data governance solutions, and send them to the target data environment for personalized data governance.
[0032] In a feasible implementation, personalized data governance is performed by combining multiple personalized data governance solutions with the data governance architecture, and then further comprising:
[0033] Continuously monitor the data access behavior of individual users after the implementation of personalized data governance.
[0034] According to the preset quantitative state migration rules, the data state change events corresponding to the data access behavior are identified, wherein the migration rules include: no access behavior within a set number of consecutive days triggers data state degradation, and the access frequency exceeding the preset threshold triggers a temporary state upgrade.
[0035] Governance policy adjustments are made based on the data status change events.
[0036] In a feasible implementation, adjusting the governance policy according to the data status change event further includes:
[0037] According to the adjustment frequency, the trigger category of the governance policy adjustment is determined, wherein the trigger category includes sudden adjustment and high-frequency adjustment.
[0038] If it is a sudden adjustment, the governance policy will be automatically rolled back after the preset retention window.
[0039] If it is a high-frequency adjustment, a retrospective analysis of the periodic behavior pattern is performed, and a corresponding periodic governance strategy is established, and the periodic governance strategy is merged into the corresponding personalized data governance solution.
[0040] The beneficial effects of the present invention are as follows: by taking multiple user individuals as behavioral observation objects in a target data environment, their historical behavior data is collected according to a set time window; performing time series traversal analysis on the collected historical behavior data to construct a behavior profile of the corresponding user individual, wherein the behavior profile includes operation behavior characteristics and location behavior characteristics; according to a predefined data hierarchical structure, the behavior profile is component-deconstructed, the data level distribution of each user individual is extracted, and the data level distribution of all individuals is integrated to generate an overall integrated level distribution of the target data environment; based on the individual data level distribution and the integrated data level distribution, a data governance system architecture adapted to the target data environment is designed and constructed; further, according to the data level distribution of each individual, a matching data governance template is screened from a preset policy template library to generate a corresponding personalized data governance plan, and combined with the constructed data governance system architecture, a targeted personalized data governance process is collaboratively executed. The personalized data governance optimization method based on behavioral time series analysis disclosed by the present invention solves the technical problems of inaccurate data classification, inability to adapt to multi-user environments, and lack of personalization and dynamic adjustment capabilities of governance plans, and achieves the technical effects of accurate and dynamic data classification, integrated governance of multi-user environments, and personalized and dynamically adjustable governance plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 The figure is a flow chart of a personalized data governance optimization method based on behavior time series analysis of the present invention.
[0042] Figure 2 This is a flow chart of constructing a data governance architecture for a target data environment in a personalized data governance optimization method based on behavioral timing analysis according to the present invention. DETAILED DESCRIPTION
[0043] The above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods of the specification to better understand the above technical solution. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments used only to explain the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, it should be noted that, for the convenience of description, only the parts related to the present invention, rather than all, are shown in the drawings.
[0044] Example, Figure 1 The figure is a flow chart of a personalized data management optimization method based on behavior time series analysis of the present invention, wherein the personalized data management optimization method based on behavior time series analysis includes:
[0045] S100: In a target data environment, multiple user individuals are used as behavior collection objects, and historical behavior data is obtained according to a preset time window.
[0046] Specifically, the behavior collection targets refer to multiple individual users in the target data environment, that is, behavioral entities with independent identities in the environment. The behavioral data of each individual user is independent and traceable. The preset time window refers to a time span pre-set based on data governance requirements. Historical behavioral data refers to various behavioral records generated in the target data environment and falling within the preset time window. Target data environments include but are not limited to home environments, studio environments, and enterprise environments.
[0047] Specifically, first, identify the target data environment, such as a home NAS server or a multi-user shared data storage environment like an enterprise or studio server, and then identify multiple individual users within it as behavior collection targets. Then, set a time window based on data governance requirements, such as one or three months, and filter historical behavior data that meets the time requirements within this time window. For example, this includes: file operation behaviors such as creation, modification, deletion, and access; and location behaviors such as user login locations and data access locations.
[0048] For example, in a home NAS environment, the operation behavior of family members on different devices for photos, documents and other files stored in the NAS and the access location information can be collected over the past three months.
[0049] Through the above process, comprehensive collection of multi-user behavior data can be achieved, thereby providing basic data support for subsequent personalized data governance based on behavioral time series analysis, ensuring that the collected data covers the behavioral characteristics of multiple user individuals within a specific time range, and can accurately reflect the behavioral habits and data usage patterns of each user.
[0050] S200: Performing a time-series traversal analysis on the historical behavior data to establish a behavior personality of each user, wherein the behavior personality includes an operation behavior personality and a location behavior personality.
[0051] Specifically, historical behavioral data is processed and analyzed in chronological order to construct the behavioral personalities of multiple individual users, namely, their preferences and habits regarding data governance (i.e., behavioral personalities). The operational behavioral personality is a set of features that reflects the habits and patterns of individual users in file manipulation, while the location behavioral personality is the behavioral patterns of users when accessing data in different locations, such as active locations and location switching frequency.
[0052] Through the above process, the unique behavioral characteristics of each individual user in the target data environment can be accurately portrayed, providing a key basis for the subsequent data classification and personalized data governance plan based on behavioral personality, realizing the transformation from traditional static data classification to dynamic behavior-driven classification, and thus improving the accuracy and personalization of data governance.
[0053] In some embodiments, a temporal traversal analysis is performed on the historical behavior data to establish a behavior personality of each individual user, wherein the behavior personality includes an operation behavior personality and a location behavior personality, including:
[0054] Based on the historical behavior data, the user's individual operation behavior data is extracted to perform a time-series traversal analysis to obtain the file operation frequency distribution, common file type combinations, high-frequency operation sequence patterns and data creation-modification ratios, and an operation behavior feature vector is generated accordingly; based on the historical behavior data, the user's individual historical location behavior data is extracted to perform a time-series traversal analysis to obtain the main active locations, location switching frequency, location stay duration and location-operation mapping relationship, and a location behavior feature vector is generated accordingly; the location behavior feature vector is used as the location behavior personality, and the operation behavior feature vector is used as the operation behavior personality, and they are associated and stored as the behavior personality.
[0055] Specifically, after acquiring historical behavior data, we first conduct a time-series traversal analysis of each user's individual operation behavior data. This analysis counts their file operation behaviors within a set time window, including daily or weekly operation frequency distribution, commonly used file type combinations (e.g., the combination of PDF and Word documents), typical high-frequency operation sequences (e.g., open, edit, save), and the ratio of data creation to modification. This generates an operation behavior feature vector. For example, if a user primarily operates on image files over a three-month period, modifies files an average of five times per day, and has a file creation to modification ratio of 1:3, their operation behavior personality may be characterized as a high-frequency editor.
[0056] Specifically, the same time-series traversal analysis is then performed on the location behavior data of individual users to extract their main active locations, location switching frequency, duration of stay at each location, and types of operation behaviors at different locations, such as mainly browsing files at home and mainly editing files at work, and a location behavior feature vector is generated based on this.
[0057] Furthermore, the operation behavior feature vector and the position behavior feature vector are associated and stored as the operation behavior personality and the position behavior personality respectively, so as to form a complete behavior personality file.
[0058] For example, taking the enterprise studio server environment as an example, it is assumed that the preset time window is the past 3 months. For employee A, the operational behavior data is extracted from the historical behavioral data, including the time, frequency, file type and other information of file creation, modification, deletion, access and other operations. After performing time sequence traversal analysis, it is obtained that the file operation frequency is concentrated from 9 am to 12 pm on weekdays, the commonly used file types are a combination of .docx and .xlsx, the high-frequency operation sequence is create file-modify file-save file, and the data creation-modification ratio is 1:3, and then the operational behavior feature vector is generated based on this. At the same time, the location behavior data of employee A is extracted, including the location where he logs into the server, the frequency of location switching, the length of time he stays at the location, and the location-operation mapping relationship. For example, file creation and modification operations are mainly performed at the office workstation, and file access operations are mainly performed in the conference room, etc., to generate a location behavior feature vector. Finally, these two feature vectors are stored in association to form the behavioral personality of employee A.
[0059] Through the above process, we can deeply explore the user's operation and location behavior characteristics based on the time dimension, and accurately characterize the user's individual behavior habits. This not only improves the accuracy of behavior modeling, but also provides high-quality personality feature input for subsequent personalized data governance strategies, thereby effectively enhancing the intelligence level and user adaptability of data governance.
[0060] S300: Decomposing the behavioral personality into components according to a preset data hierarchical structure, obtaining the individual data level distribution of each user, and fusing multiple individual data level distributions to obtain a fused data level distribution of the target data environment.
[0061] Specifically, the data hierarchy is a pre-defined data classification system used to assign different management strategies and resources to data. Based on this data hierarchy, the characteristics of operational and location behavior characteristics can be broken down and matched to corresponding data levels, generating a corresponding individual data level distribution.
[0062] Specifically, individual data level distribution refers to the distribution of each user across various data levels, reflecting the proportion of correlation between user behavior patterns and each data level. Fusion data level distribution combines the individual data level distributions of multiple users to obtain the comprehensive data level distribution of the entire target data environment, reflecting the overall distribution of multi-user group behavior characteristics across various data levels.
[0063] Through this process, individual user behaviors can be accurately mapped to a data hierarchy, achieving the transformation from behavioral characteristics to data levels. Simultaneously, by integrating the distribution of multiple individual users, the global data level distribution of the target data environment can be derived, providing a basis for the subsequent construction of a data governance architecture, ensuring that data governance strategies are both consistent with the behavioral characteristics of individual users and adaptable to the overall trends of a multi-user environment.
[0064] In some embodiments, the data hierarchy is constructed based on data sensitivity, access frequency, data importance, and compliance rules, and includes:
[0065] When an individual user frequently operates a data object during a high-frequency access period, the data object will be mapped to the hot data level; when a data object has not been accessed within a set silence threshold period, the data object will be mapped to the cold data level; when a data object meets any of the conditions of being associated with a highly sensitive application behavior or being located in a preset sensitive geographical location, the data object will be mapped to a highly sensitive level.
[0066] Specifically, the data classification structure is classified according to dimensions such as frequency of use, sensitivity, importance level, and compliance requirements of the data. For example, it can be divided into hot / cold data, general information and privacy, confidential information, and critical business data.
[0067] Specifically, the "hot data" level corresponds to data that is frequently accessed or frequently operated, typically with high usage activity; the "cold data" level corresponds to data that has not been accessed or used for a long time; and the "high sensitivity" level corresponds to data objects associated with sensitive behaviors or sensitive geographic locations, which may involve privacy, business secrets, or compliance risks. The "silence threshold period" refers to a set period of time. If data is not accessed during this period, it is considered silent and classified as cold data.
[0068] Specifically, in terms of hot data identification, we monitor users' operations on various data objects during high-frequency access periods. If a data object is frequently created, modified, accessed, etc. during this period, it is determined to be hot data. For example, if a user edited an Excel report file every day for the past week, the file would be mapped to the hot data level.
[0069] Specifically, in terms of cold data determination, a silent threshold period can be set. If a data object has not been accessed or operated during this period, it will be classified as cold data. For example, if a photo has not been accessed for 60 days since it was uploaded, it can be marked as cold data.
[0070] Specifically, in terms of highly sensitive data identification, it is achieved by analyzing the association between operational behavior and application context. For example, if the data object is used for highly sensitive applications such as medical and financial or its access behavior occurs in a preset sensitive geographical location, such as a company's confidential area, overseas access, etc., it will be mapped to a high sensitivity level.
[0071] Through the above process, data objects can be accurately classified according to the actual use and security requirements of the data, ensuring that the data governance strategy matches the actual status and importance of the data, improving the efficiency and security of data management, and better adapting to the personalized data governance needs in a multi-user environment, providing a scientific basis for the subsequent construction of data governance architecture and formulation of governance plans.
[0072] In some embodiments, the behavioral personality is decomposed into components according to a preset data hierarchy structure to obtain the individual data level distribution of each user, and multiple individual data level distributions are fused to obtain a fused data level distribution of the target data environment, including:
[0073] Taking data sensitivity, access frequency, importance and compliance as data classification features, traverse the data classification structure to match the behavioral personality corresponding to each user individual with the corresponding data level, and construct a level matching result; count the frequency of each data level in the level matching result, and perform a cluster analysis with the user individual as the clustering dimension to obtain the individual data level distribution of each user individual; traverse the individual data level distribution of multiple user individuals, perform a secondary cluster analysis with data level as the clustering dimension, and determine the fused data level distribution of the target data environment.
[0074] Specifically, component decomposition is the process of disassembling and mapping the behavioral personality of individual users according to established data classification characteristics, which is used to identify the distribution of data objects involved in the behavior of different individual users at different levels. For example, data classification characteristics may involve data sensitivity, access frequency, importance, compliance, etc.
[0075] Specifically, individual data level distribution refers to the proportion of various data levels, such as hot data, cold data, and highly sensitive data, in a single user's behavioral data. Fusion data level distribution is a global view formed by aggregating and analyzing the data level distributions of multiple users, reflecting the overall proportion and structure of data of different levels in the entire data environment.
[0076] Specifically, in a home NAS environment, family member A's behavioral personality indicates a high frequency of file operations, with a common file type combination of .jpg and .png. Their primary locations of activity are the bedroom computer and the living room tablet, with low switching frequency. Correspondingly, component decomposition based on the data hierarchy reveals that A's behavioral patterns are closely correlated with the hot data level. Therefore, their individual data level distribution has a high proportion of hot data (assuming 70%), a low proportion of cold data (assuming 20%), and a high sensitivity (assuming 10%). Similarly, component decomposition is performed on the behavioral personalities of the other family members to obtain their respective individual data level distributions.
[0077] For example, it is assumed that the behavioral personality data sample is as shown in Table 1 below:
[0078] Table 1 Behavioral personality data sample
[0079]
[0080] The corresponding individual data level distribution, that is, the result of a cluster analysis is as follows Table 2:
[0081] Table 2 Sample distribution of individual data levels
[0082]
[0083] Correspondingly, the user's individual data level distribution can be expressed as:
[0084] U1={L1, 1 / 3, L2, 1 / 3, L3, 1 / 3, L4, 0}; U2={L1, 0, L2, 1 / 2, L3, 0, L4, 1 / 2}; U3={L1, 1 / 3, L2, 1 / 3, L3, 1 / 3, L4, 0}. Specifically, the individual data level distributions of all family members are traversed. Using data level as the clustering dimension, secondary clustering analysis is performed using methods such as weighted average to calculate the integrated data level distribution of the entire family NAS environment. This is the global distribution pattern of all types of data levels in the entire data environment. This integrated distribution can be used to guide the construction of the subsequent data governance architecture.
[0085] For example, the rank distribution of all users in the above example is merged by user, and clustered based on rank. The obtained secondary clustering analysis result, i.e., the fused data rank distribution, can be expressed as shown in Table 3 below:
[0086] Table 3. Example of fusion data level distribution
[0087]
[0088] Correspondingly, the fusion data level distribution can be expressed as:
[0089] Denv (Fused distribution) = {L1, 2 / 9, L2, 3 / 9, L3, 2 / 9, L4, 1 / 9}.
[0090] Through the above process, not only can data level modeling be achieved at the individual level, but also a global data level portrait can be integrated at the group level, which helps to improve the level of refinement of data governance and enables the system to dynamically adjust resource allocation strategies according to the actual distribution of user behavior, thereby achieving efficient, secure and intelligent data management.
[0091] S400: Build a data governance architecture for the target data environment by combining the individual data level distribution and the fusion data level distribution.
[0092] Specifically, the data governance architecture is a logical framework built based on data hierarchical distribution to guide data storage, management, and access. It includes the storage layer structure, storage capacity allocation, and performance configuration corresponding to different data levels. This architecture combines individual behavioral characteristics with global data distribution, embodying the principles of on-demand allocation and hierarchical governance.
[0093] In some embodiments, as Figure 2 As shown, combining the individual data level distribution and the fusion data level distribution, a data governance architecture for the target data environment is constructed, including:
[0094] A baseline logical storage layer structure is constructed based on the fused data level distribution, and storage capacity and performance configuration are allocated, wherein the baseline logical storage layer structure includes multiple logical storage layers, each logical storage layer corresponds to a data level; based on multiple individual data level distributions, the user data distribution corresponding to each data level is determined; according to the user data distribution, each logical storage layer is divided into multiple user individual storage blocks to form the data governance architecture.
[0095] Specifically, the logical storage layer structure refers to an abstract storage layer divided above the physical storage resources. Each layer corresponds to a specific data level, such as the hot data layer, cold data layer, and highly sensitive data layer, enabling hierarchical data storage and management. The user individual storage block refers to a dedicated storage area within the logical storage layer designated for each individual user, used to store data of the corresponding level generated during their activities.
[0096] For example, in an enterprise studio server environment, component decomposition and cluster analysis are performed to determine the individual data level distribution for each employee. A baseline logical storage layer structure is then constructed based on the integrated data level distribution. For example, high-speed storage resources (SSD storage or high-speed areas in hybrid storage) are allocated to the hot data level; low-speed storage resources (HDD storage) are allocated to the cold data level; and encrypted storage resources are allocated to the highly sensitive level. The distribution of each employee's data across the logical storage layers is then determined based on the individual data level distribution. Specifically, the individual data level distributions of all users are traversed, and each user's contribution to the data at each level is calculated to form a user data distribution map (i.e., the user data distribution corresponding to the data level). Each logical layer is then further divided into multiple user-specific storage blocks to ensure that each user's data is properly isolated and managed within the logical layer of their corresponding level. For example, if employee A's hot data accounts for 30% of all hot data, then 30% of the high-speed storage resources (logical storage layer) allocated to the hot data level will be allocated to employee A's user-specific storage blocks.
[0097] Furthermore, based on the above method steps, each logical storage layer is divided into multiple user individual storage blocks according to the user data distribution to form a data governance architecture, ensuring that each user's data is reasonably stored according to their individual data level distribution, and at the same time the overall architecture meets the requirements of the integrated data level distribution.
[0098] Through this process, we can fully leverage the complementary advantages of individual data hierarchical distribution and integrated data hierarchical distribution to build a data governance architecture that not only conforms to the behavioral characteristics of individual users but also adapts to the overall trends of multi-user environments. This in turn enables the rational allocation of data storage resources, improves data access efficiency, and enhances data security.
[0099] S500: According to the distribution of multiple individual data levels, match corresponding data governance templates from the preconfigured policy template library, construct multiple personalized data governance solutions accordingly, and perform personalized data governance in combination with the multiple personalized data governance solutions and the data governance architecture.
[0100] Specifically, the policy template library is a preconfigured resource collection containing multiple data governance policy templates, each corresponding to a specific data level distribution pattern and governance requirements. A personalized data governance plan is generated based on multiple data governance templates matched to the user's individual data level distribution, specifically tailored to that user's data governance needs. A data governance template is a data management policy selected from the policy template library that applies to a specific data level and is used to guide data governance for the corresponding user's individual storage blocks, such as storage, updates, migration, or deletion.
[0101] The above process accurately matches and generates personalized data governance solutions based on the data level distribution characteristics of each individual user, and combines the data governance architecture to ensure the effective implementation of the solution. It can achieve refined management of different user data, improve the pertinence and effectiveness of data governance, and meet the diverse data governance needs in a multi-user environment.
[0102] In some embodiments, according to the multiple individual data level distributions, corresponding data governance templates are matched from a preconfigured policy template library, multiple personalized data governance solutions are constructed accordingly, and personalized data governance is performed in combination with the multiple personalized data governance solutions and the data governance architecture, including:
[0103] For any of the individual data level distributions, extract the data levels contained therein to form a data level list; traverse the data level list, and use the data level as a matching index to match the corresponding data governance template in the policy template library to form a data governance template cluster, wherein the data governance template includes data backup frequency, data backup rules and data synchronization time; according to the data governance architecture, establish a mapping relationship between the data governance template cluster and the user's individual storage block to form the personalized data governance plan; repeatedly traverse multiple individual data level distributions to form multiple personalized data governance plans, and send them to the target data environment for personalized data governance.
[0104] Specifically, a data governance template cluster is a collection of governance templates that match a user's data levels. A personalized data governance plan is a customized data governance implementation plan for a user, created by binding the template cluster to the user's corresponding storage block in the data governance architecture.
[0105] Specifically, first, the individual data level distribution of each user is traversed, and all the data levels involved are extracted to form a data level list. For example, the behavior of user A involves two levels: hot data and highly sensitive data. Then, using the level in the data level list as an index, the corresponding governance template is searched in the policy template library. For example, the template corresponding to hot data is daily incremental backup + real-time synchronization, and the template corresponding to highly sensitive data is hourly full backup + scheduled encryption synchronization. The above two templates constitute the governance template cluster of user A. Then, according to the aforementioned constructed data governance architecture, the corresponding storage block of user A in the logical storage layer is found, such as Block_A1 in the hot data layer and Block_A2 in the highly sensitive data layer, and each template in the governance template cluster is bound to its corresponding storage block to form a personalized data governance plan for user A.
[0106] For example, the personalized data management solution for user A is: Block_A1 performs daily incremental backup + real-time synchronization, and Block_A2 performs hourly full backup + scheduled encryption synchronization.
[0107] Furthermore, the above process is repeated to process the individual data level distribution of all users, forming multiple personalized data governance plans, and these plans are sent to the corresponding modules in the data environment for execution.
[0108] Through the above process, accurate matching from policy templates to individual behaviors is achieved, ensuring that each user's data can obtain the most appropriate governance strategy at its specific level, thereby realizing dynamic personalized data governance centered on user behavior and driven by data levels, and improving data security, compliance and system operation efficiency.
[0109] In some embodiments, personalized data governance is performed by combining multiple personalized data governance solutions with the data governance architecture, and then further comprising:
[0110] Continuously monitor the data access behavior of individual users after the implementation of personalized data governance; identify data state change events corresponding to the data access behavior based on preset quantitative state migration rules, wherein the migration rules include: no access behavior for a set number of consecutive days triggers data state degradation, and access frequency exceeding a preset threshold triggers a temporary state upgrade; adjust the governance strategy based on the data state change event.
[0111] Specifically, a data state change event refers to an event during data usage that requires a data level adjustment due to changes in access behavior, such as downgrading data from hot to cold, or temporarily upgrading it from cold to hot. Quantitative state transition rules are triggers set based on access behavior to determine whether data should undergo a state transition. For example, downgrading data after seven consecutive days of no access or upgrading it after more than ten daily accesses.
[0112] Specifically, governance strategy adjustment refers to dynamically modifying the corresponding data backup frequency, synchronization strategy, storage level and other parameters in the data governance plan according to changes in data status, so as to maintain consistency between the governance strategy and the actual usage status of the data.
[0113] Specifically, after implementing a personalized data governance solution, users' access behavior data, such as access frequency, access time, and access patterns, is collected and analyzed in real time or periodically. For example, user A may not have accessed a cold data block in the past seven days, or user B may have accessed sensitive data more than 200 times in 24 hours.
[0114] Specifically, then, based on the preset quantitative state migration rules, data state change events triggered by access behaviors are identified. For example, user A's data block triggers the conditions for cold data to be further downgraded to archived data, and user B's data block triggers the conditions for temporary upgrade to hot data.
[0115] Furthermore, based on the determined data status change events, the governance policy applicable to the corresponding data blocks is automatically adjusted. For example, the backup frequency of user A's data blocks is adjusted from once a week to once a month, and the data synchronization policy is adjusted from scheduled synchronization to no longer synchronization after archiving; for user B's data blocks, they are temporarily migrated to a high-performance storage layer, and the backup frequency is increased to once a day, and the synchronization policy is adjusted to real-time synchronization.
[0116] This process enables a closed-loop feedback loop for data governance solutions, encompassing a complete process from policy formulation and execution to behavioral feedback and policy adjustments. This ensures that data governance solutions can respond to changes in user behavior in real time, preventing governance policy failures and resource waste caused by unrecognized data state changes. Furthermore, the state migration approach based on quantitative rules enhances the objectivity and controllability of judgments, improving the accuracy and automation of data governance.
[0117] In some implementations, adjusting the governance policy based on the data state change event further includes:
[0118] According to the adjustment frequency, the trigger category of the governance policy adjustment is determined, where the trigger category includes sudden adjustment and high-frequency adjustment; if it is a sudden adjustment, the governance policy is automatically rolled back after the preset retention window; if it is a high-frequency adjustment, a periodic behavior pattern retrospective analysis is performed, and a corresponding periodic governance policy is established, and the periodic governance policy is merged into the corresponding personalized data governance solution.
[0119] Specifically, the trigger category is used to classify the causes and characteristics of policy adjustments, including sudden adjustments, i.e., policy adjustments caused by sudden changes in access behavior within a short period of time, and high-frequency adjustments, i.e., policy adjustments that occur frequently within a period of time, reflecting the periodicity or volatility of user behavior.
[0120] Specifically, rollback governance policies are used for sudden adjustments to prevent misjudgments that could impact system performance or waste resources. For example, they can automatically revert to the original governance policy after a set time window. Retrospective analysis of periodic behavior patterns identifies patterns in governance policy adjustments triggered by user access behavior over a period of time, such as weekday-weekday patterns or sleep patterns, thereby building more stable and predictable governance policies.
[0121] Specifically, after identifying data state change events based on data access behavior and triggering policy adjustments, further classification is performed based on the frequency of the adjustments. If a particular adjustment (or a category of adjustments on a user's individual storage block) is sudden, such as a user accessing a piece of data three times in a single day after not accessing it for 30 consecutive days, this can be identified as a sudden adjustment. A retention window, such as 48 hours, is then set after the policy adjustment is executed. If the access behavior does not persist after this 48-hour window, the policy is automatically rolled back to the original policy to avoid wasting resources. Conversely, if a piece of data undergoes frequent policy adjustments within a week, for example, if the daily access frequency fluctuates dramatically, resulting in repeated policy changes, this can be identified as a high-frequency adjustment. A retrospective analysis of periodic behavior patterns is then performed. For example, methods such as sliding window analysis and spectrum analysis are used to identify the periodicity of access behavior and generate corresponding periodic governance policies. For example, data levels can be upgraded in advance on peak days, backup frequency can be increased, and this policy can be incorporated into the user's personalized data governance plan.
[0122] Through this process, we can extract and leverage the cyclical patterns of user behavior while addressing sudden behavioral changes, dynamically optimizing data governance strategies. On the one hand, the rollback mechanism for sudden adjustments avoids resource waste and policy volatility caused by short-term behavioral fluctuations. On the other hand, pattern recognition and cyclical strategy construction for high-frequency adjustments enhance the foresight and stability of governance strategies. This helps enhance the adaptability and intelligence of personalized data governance solutions, improving resource utilization efficiency and user experience.
[0123] In summary, the personalized data governance optimization method based on behavior time series analysis provided by the present invention has the following technical effects:
[0124] By taking multiple user individuals as behavioral observation objects in the target data environment, their historical behavior data are collected according to the set time window; time-series traversal analysis is performed on the collected historical behavior data to construct a behavior portrait of the corresponding user individual, where the behavior portrait includes operation behavior characteristics and location behavior characteristics; based on the predefined data hierarchical structure, the behavior portrait is decomposed into components, the data level distribution of each user individual is extracted, and the data level distribution of all individuals is integrated to generate the overall fusion level distribution of the target data environment; based on the individual data level distribution and the fusion data level distribution, a data governance system architecture adapted to the target data environment is designed and constructed; further, based on the data level distribution of each individual, matching data governance templates are screened from the preset policy template library to generate corresponding personalized data governance solutions, and combined with the constructed data governance system architecture, targeted personalized data governance processes are collaboratively executed to achieve the technical effects of accurate and dynamic data classification, integrated governance of multi-user environments, and personalized and dynamically adjustable governance solutions.
[0125] It should be understood that the embodiments disclosed in the present invention and the above description can enable those skilled in the art to use the present invention to implement the present invention. At the same time, the present invention is not limited to the embodiments mentioned above. It should be understood that those skilled in the art can still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention and are all included in the scope of protection of the present invention.
Claims
1. A personalized data governance optimization method based on behavioral time series analysis, characterized in that: The method comprises: In the target data environment, multiple user individuals are used as behavior collection objects, and historical behavior data is obtained according to the preset time window; Performing a time series traversal analysis on the historical behavior data to establish a behavior personality of each individual user, wherein the behavior personality includes an operation behavior personality and a location behavior personality; Decomposing the behavioral personality into components according to a preset data hierarchical structure, obtaining the individual data level distribution of each user, and fusing multiple individual data level distributions to obtain a fused data level distribution of the target data environment; Combining the individual data level distribution and the fusion data level distribution, a data governance architecture for the target data environment is constructed; According to the multiple individual data level distributions, corresponding data governance templates are matched from a pre-configured policy template library, multiple personalized data governance solutions are constructed accordingly, and personalized data governance is performed in combination with the multiple personalized data governance solutions and the data governance architecture; The fusion data level distribution of the target data environment is obtained, including: Taking data sensitivity, access frequency, importance and compliance as data classification features, traversing the data classification structure to match the behavioral personality corresponding to each individual user with the corresponding data level, and constructing a level matching result; Counting the frequency of each data level in the level matching results, and performing a cluster analysis with individual users as clustering dimensions to obtain the individual data level distribution of each individual user; Traversing the individual data level distribution of multiple individual users, performing secondary cluster analysis with data level as a clustering dimension, and determining the fused data level distribution of the target data environment; The data governance architecture for the target data environment is constructed, including: Constructing a baseline logical storage layer structure based on the fused data level distribution and allocating storage capacity and performance configuration, wherein the baseline logical storage layer structure includes a plurality of logical storage layers, each logical storage layer corresponding to a data level; determining, based on the plurality of individual data level distributions, a user data distribution corresponding to each data level; According to the user data distribution, each of the logical storage layers is divided into a plurality of user individual storage blocks to form the data governance architecture.
2. The personalized data management optimization method based on behavior time series analysis according to claim 1 is characterized in that: Performing a time series traversal analysis on the historical behavior data to establish the behavior personality of each individual user, wherein the behavior personality includes operation behavior personality and location behavior personality, including: Based on the historical behavior data, extract the user's individual operation behavior data to perform time series traversal analysis to obtain file operation frequency distribution, common file type combinations, high-frequency operation sequence patterns and data creation-modification ratios, and generate corresponding operation behavior feature vectors; Based on the historical behavior data, extract the historical location behavior data of individual users to perform time series traversal analysis to obtain the main active locations, location switching frequency, location stay duration, and location-operation mapping relationship, and generate a corresponding location behavior feature vector; The position behavior feature vector is used as the position behavior personality, and the operation behavior feature vector is used as the operation behavior personality, and they are associated and stored as the behavior personality.
3. The personalized data management optimization method based on behavior time series analysis according to claim 2 is characterized in that: The data hierarchy is constructed based on data sensitivity, access frequency, data importance and compliance rules, and includes: When an individual user frequently operates on a data object during a high-frequency access period, the data object is mapped to the hot data level; When a data object has not been accessed within the set silence threshold period, the data object is mapped to the cold data level; When a data object satisfies any one of the conditions of being associated with a highly sensitive application behavior or being located in a preset sensitive geographical location, the data object is mapped to a high sensitivity level.
4. The personalized data management optimization method based on behavior time series analysis according to claim 3 is characterized in that: According to the multiple individual data level distributions, corresponding data governance templates are matched from a pre-configured policy template library, multiple personalized data governance solutions are constructed accordingly, and personalized data governance is performed in combination with the multiple personalized data governance solutions and the data governance architecture, including: For any of the individual data level distributions, extract the included data levels to form a data level list; Traversing the data level list, matching the corresponding data governance template in the policy template library with the data level as a matching index to form a data governance template cluster, wherein the data governance template includes data backup frequency, data backup rules and data synchronization time; According to the data governance architecture, a mapping relationship between the data governance template cluster and the user's individual storage block is established to form the personalized data governance solution; Repeatedly traverse multiple individual data level distributions to form multiple personalized data governance solutions, and send them to the target data environment for personalized data governance.
5. The personalized data management optimization method based on behavior time series analysis according to claim 1 is characterized in that: Combining multiple personalized data governance solutions with the data governance architecture to perform personalized data governance, and then further comprising: Continuously monitor the data access behavior of individual users after the implementation of personalized data governance; Identify data state change events corresponding to the data access behavior based on preset quantitative state migration rules, where the migration rules include: no access behavior for a set number of consecutive days triggers data state degradation, and access frequency exceeding a preset threshold triggers a temporary state upgrade; Governance policy adjustments are made based on the data status change events.
6. The personalized data management optimization method based on behavior time series analysis according to claim 5 is characterized in that: Adjusting the governance policy based on the data status change event also includes: According to the adjustment frequency, the trigger category of the governance policy adjustment is determined, wherein the trigger category includes sudden adjustment and high-frequency adjustment; If it is a sudden adjustment, the governance policy will be automatically rolled back after the preset retention window; If it is a high-frequency adjustment, a retrospective analysis of the periodic behavior pattern is performed, and a corresponding periodic governance strategy is established, and the periodic governance strategy is merged into the corresponding personalized data governance solution.
Citation Information
Patent Citations
Cloud storage data grading method based on user behaviors
CN109918448A
Cloud data caching method and device, equipment and storage medium
CN114840140A