Data asset management method
By introducing a multi-level feature extraction and dynamic classification mechanism, combined with the timing optimal data scheduling algorithm, the shortcomings of traditional data asset management methods in real-time evaluation, dynamic classification, storage resource optimization, etc. are solved, and efficient, refined and dynamic management of data assets are achieved, data response efficiency and resource utilization are improved, while reducing storage and operation costs.
Patent Information
- Application Number
- CN202510251515.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional data asset management methods have significant shortcomings in real-time data evaluation, dynamic classification, storage resource optimization, response time management, energy consumption control, and life cycle management, which is difficult to meet the needs of efficient, refined and dynamic data asset management, which restricts the effective use of the value of data assets.
By collecting data and introducing a multi-level feature extraction mechanism, the initial multi-dimensional attribute vector is generated and projected into the high-dimensional embedding space to generate time-dependent mapping vectors. Then, based on the energy difference between data, a dynamic threshold adjustment mechanism is introduced for dynamic classification, and the storage location and access priority of data are dynamically optimized through the timing optimal data scheduling algorithm.
It realizes efficient, refined and dynamic management of data assets, automatically adjusts the allocation of data between efficient storage devices and low-cost storage devices, ensures the optimization of response speed and storage costs, improves data response efficiency and resource utilization, and reduces storage and operation costs.
Smart Images

Figure CN120179168A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and particularly to a method for data asset management. Background Art
[0002] The importance of data asset management has become increasingly prominent in today's digital age. However, the management of data assets is not just about simple storage or backup of data, but rather through scientific management methods to achieve refined utilization, life cycle management, and efficient allocation of data assets, enabling data assets to maximize their value under the premise of security and compliance.
[0003] The management and utilization of data assets involve multiple aspects, including classification, storage, access control, life cycle management, and security management of data, etc. First of all, due to different sources, types, and uses of data, the value and importance of each piece of data vary greatly. Therefore, how to scientifically classify and evaluate data assets and identify data of great significance has become an important task in data management; secondly, the usage frequency, storage location, and security requirements of data will also change continuously with the change of business needs. For example, core data often needs to be preferentially stored in a high-performance storage system to ensure immediate access and secure access, while other low-frequency access or historical data is suitable for storage in low-cost storage media to reduce costs.
[0004] However, traditional data asset management methods have significant deficiencies in aspects such as real-time evaluation, dynamic classification, storage resource optimization, response time management, energy consumption control, and life cycle management of data, making it difficult to meet the requirements of efficient, refined, and dynamic data asset management, and restricting the effective play of the value of data assets. Summary of the Invention
[0005] The present invention provides a method for data asset management to solve the problem that traditional data asset management has significant deficiencies in aspects such as real-time evaluation, dynamic classification, storage resource optimization, response time management, energy consumption control, and life cycle management of data, making it difficult to meet the requirements of efficient, refined, and dynamic data asset management, and restricting the effective play of the value of data assets.
[0006] A method for data asset management according to the present invention specifically includes the following technical solutions:
[0007] A method for data asset management includes the following steps:
[0008] S1. Collect data and introduce a multi-level feature extraction mechanism to generate an initial multi-dimensional attribute vector; project the initial multi-dimensional attribute vector into a high-dimensional embedding space to generate a time-related mapping vector; perform adaptive dynamic evaluation on the mapping vector and calculate the energy difference between data;
[0009] S2. Introduce a dynamic threshold adjustment mechanism based on the energy difference between data for dynamic classification to obtain the classification result of the data; based on the classification result of the data, dynamically optimize the storage location and access priority of the data through the time-series optimal data scheduling algorithm.
[0010] Preferably, the S1 specifically includes:
[0011] The multi-level feature extraction mechanism quantifies the metrics of the data to generate an initial multi-dimensional attribute vector.
[0012] Preferably, the S1 specifically includes:
[0013] Project the initial multi-dimensional attribute vector into a high-dimensional embedding space through a non-linear mapping formula to generate a time-related mapping vector. The non-linear mapping formula is:
[0014]
[0015] Among them, Z i (t) represents the mapping vector of the i-th data mapped to the embedding space at time t, indicating the time-related mapping vector; w ik (t) is the time-related weight factor of the i-th data in the k-th dimension at time t; α k is the frequency factor in the sine function; x ik is the original attribute value of the i-th data in the k-th dimension; β k is the phase shift in the sine function; γ ik (t) is the time-related exponential mapping coefficient of the i-th data in the k-th dimension at time t; λ k is the growth rate parameter of the exponential function in the k-th dimension; n is the total number of dimensions.
[0016] Preferably, the S1 specifically includes:
[0017] Introduce a dynamic non-linear energy equation to adaptively dynamically evaluate the mapping vector and calculate the energy difference between data.
[0018] Preferably, the S2 specifically includes:
[0019] In the implementation process of the dynamic threshold adjustment mechanism, based on the energy difference between data and combined with the classification threshold, dynamically classify the data to obtain the classification result of the data.
[0020] Preferably, the S2 specifically includes:
[0021] In the implementation process of the dynamic threshold adjustment mechanism, combined with the data usage frequency, introduce a sine function and an integral term to calculate the classification threshold.
[0022] Preferably, the S2 specifically includes:
[0023] In the implementation process of the optimal timing data scheduling algorithm, based on the classification result of the data, a scheduling optimization function is constructed and minimized to adjust the storage location and access priority of the data. The formula of the scheduling optimization function is:
[0024]
[0025] Where is the scheduling optimization function; N is the total number of data in the dataset; C i (t) represents the classification result of the i-th data at time t; S i (t) is the storage cost of the i-th data; τ i (t) is the response time of the i-th data; R i (t) is the access frequency of the i-th data at time t; μ is the energy cost control coefficient; P i (s) represents the energy consumption of the i-th data at time s; s is the time index.
[0026] Preferably, the S2 specifically includes:
[0027] Based on the result calculated by the scheduling optimization function, data storage optimization, access priority management, resource configuration and cost control, and asset life cycle management are carried out.
[0028] The beneficial effects of the technical solution of the present invention are:
[0029] 1. The present invention introduces a scheduling optimization function to automatically adjust the allocation of data between high-efficiency storage devices and low-cost storage devices; data with high access frequency or high importance is preferentially stored on high-performance storage to ensure the response speed; while data that is not frequently accessed or used for archiving is stored in a lower-cost storage medium, effectively reducing the overall storage cost and achieving refined configuration of storage resources.
[0030] 2. The present invention combines the response time and access frequency parameters to automatically dynamically adjust the access priority of the data, making the call of critical data more efficient; for high-priority data access requests, the system pre-loads the relevant data, significantly shortening the access latency, ensuring the timely response of business-critical data, avoiding resource waste, and effectively improving the utilization rate of resources and data response efficiency.
[0031] 3. By introducing an energy cost control coefficient, the present invention can select appropriate storage devices according to energy consumption costs. On the premise of meeting data access requirements, it preferentially uses low-energy-consuming storage devices to further reduce resource costs. It periodically monitors the energy consumption of each storage device and realizes higher energy efficiency management through automated resource adjustment, significantly reducing the operating costs of the devices while ensuring the availability of data assets.
[0032] 4. The scheduling optimization function of the present invention can automatically identify the life cycle stages of data. For example, it can identify data assets that have been in a low priority, low usage rate, and high storage cost for a long time, and automatically archive or destroy them according to the situation, thereby releasing storage resources, reducing management costs, ensuring the reasonable management of data assets throughout the life cycle, and achieving efficient life cycle control through dynamic monitoring and policy adjustment. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of a data asset management method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0036] The following specifically describes the specific solution of a data asset management method provided by the present invention in conjunction with the accompanying drawings.
[0037] Referring to the attached Figure 1 , which shows a flowchart of a data asset management method provided by an embodiment of the present invention. The method includes the following steps:
[0038] S1. Collect data and introduce a multi-level feature extraction mechanism to generate an initial multi-dimensional attribute vector; project the initial multi-dimensional attribute vector into a high-dimensional embedding space to generate a time-related mapping vector; perform adaptive dynamic evaluation on the mapping vector and calculate the energy difference between data;
[0039] Data is obtained from various data sources (including but not limited to databases, file systems, real-time stream processing platforms, etc.) through a standardized interface, covering important information in the asset management process. During the acquisition process, a multi-level feature extraction mechanism is introduced to automatically quantify key indicators such as data sensitivity, relevance, and usage frequency, generating an initial multi-dimensional attribute vector X i =(x i1 , x i2 ,..., x in ), where i is the i-th data sample in the dataset, each multi-dimensional attribute vector is a sample, and n is the total number of dimensions, that is, the number of features contained in the i-th data sample X i . The acquisition process not only obtains the information of the data itself but also captures the operation behaviors of the data (such as access records, permission information), providing a reliable basis for subsequent dynamic evaluation.
[0040] The initial multi-dimensional attribute vector is precisely described by a set of quantization criteria. For example, sensitivity can be quantified according to the leakage risk score, usage frequency represents the invocation frequency of the data within a period of time, and relevance represents the degree of association between the data and other data.
[0041] Through a non-linear mapping formula, the initial multi-dimensional attribute vector is projected into a high-dimensional embedding space to generate a time-related mapping vector, and its formula is:
[0042]
[0043] where Z i (t) represents the mapping vector of the i-th data mapped to the embedding space at time t, containing the high-dimensional mapping features of the data, representing the time-related mapping vector; w ik (t) is the time-related weight factor of the i-th data in the k-th dimension at time t, used to control the projection influence in the k-th dimension, obtained through experiments; α k is the frequency factor in the sine function, corresponding to the k-th dimension, used to control the amplitude of the non-linear mapping in the k-th dimension, obtained through experiments; x ik is the original attribute value of the i-th data in the k-th dimension, such as the leakage risk score or invocation frequency, etc.; β k is the phase shift in the sine function, corresponding to the k-th dimension, used to control the initial phase of the non-linear mapping in the k-th dimension, obtained through experiments; γ ik (t) is the time-related exponential mapping coefficient of the i-th data in the k-th dimension at time t, determining the mapping weight of the k-th dimension at time t, obtained through experiments; λ kIt is the growth rate parameter in the exponential function of the k-th dimension, corresponding to the k-th dimension, determining the non-linear growth rate of the mapping in the k-th dimension, and obtained through experiments. Through non-linear mapping, the characteristics of each data in different dimensions can be dynamically adjusted, so that the data can dynamically exhibit time correlation in the high-dimensional embedding space and capture the complexity of features.
[0044] To perform an adaptive dynamic evaluation of the mapping vector, a dynamic non-linear energy equation is introduced. By calculating the "energy difference" between each pair of data, the dynamic characteristics and classification differences of the data are evaluated, providing a basis for subsequent classification. The formula for the dynamic non-linear energy equation is as follows:
[0045]
[0046] Among them, E ij (t) represents the energy difference between the i-th data and the j-th data at time t; p is the total number of dimensions of the embedding space; Z im (t) is the value of the i-th data mapped to the m-th dimension at time t; Z jm (t) is the value of the j-th data mapped to the m-th dimension at time t; the dynamic non-linear energy equation analyzes the change rate of the i-th data in the m-th dimension to reflect the dynamic characteristics of the data; δ m is the energy weight of the m-th dimension, determining the influence degree of the m-th dimension on energy calculation, and obtained through experiments; κ is a parameter used to control the non-linear growth of the energy difference, making the "energy difference" between data in the high-dimensional embedding space show complex non-linear changes, further ensuring that the data changes with the flow of time in the embedding space.
[0047] S2. Based on the energy difference between data, a dynamic threshold adjustment mechanism is introduced for dynamic classification to obtain the classification result of the data; based on the classification result of the data, through the time-series optimal data scheduling algorithm, the storage location and access priority of the data are dynamically optimized.
[0048] A dynamic threshold adjustment mechanism is introduced. Based on the output result of the dynamic non-linear energy equation (i.e., the energy difference between data), combined with the classification threshold of the data, and dynamic classification is performed to obtain the classification result of the data; the data is divided by the classification threshold so that the classification standard of the data can be automatically adjusted with time. Specifically, the core formula of the classification algorithm is:
[0049]
[0050] Among them, C i(t) represents the classification result of the i-th data at time t, that is, to determine whether the i-th data belongs to a certain specific category; N is the total number of data in the dataset; the indicator function Ⅱ is used to determine whether the energy difference value between data is less than the classification threshold θ(t), and returns 1 when E ij (t) < θ(t) is true, otherwise returns 0. The classification threshold θ(t) is adjusted in real time according to the following formula:
[0051] θ(t) = θ0 + η·sin(ωt) + ζ·∫0 t f(s)ds
[0052] where, θ0 is the initial classification threshold; η is the oscillation amplitude coefficient of the classification threshold, used to control the fluctuation amplitude of the classification threshold, obtained through experiments; ω is a parameter used to control the sine wave frequency, determining the speed of the dynamic threshold fluctuation, obtained through experiments; f(s) is the usage frequency of the data changing with time; ζ is the regulation coefficient of the integral term, used to control the cumulative impact of the frequency change on the classification threshold, obtained through experiments. Through the dynamic threshold adjustment mechanism, the data is adaptively divided into different categories according to the energy difference between data, and the classification result is updated in real time, thus completing the adaptive classification of the data in the embedding space.
[0053] To achieve the optimal utilization of resources and the efficient scheduling of dynamic data, a timing-optimal data scheduling algorithm is designed. Based on the classification result of the data, the storage location and access priority of the data are scheduled and optimized to ensure the data access efficiency while minimizing the storage cost and response time. The goal of the scheduling optimization is to minimize the following scheduling optimization function:
[0054]
[0055] where, is the scheduling optimization function, used to evaluate the optimization effect of data storage and access scheduling; S i (t) is the storage cost of the i-th data, used to measure the storage cost of this data on different storage media; τ i (t) is the response time of the i-th data, indicating the delay time from the request to the access of this data; R i (t) is the access frequency of the i-th data at time t; μ is the energy cost control coefficient, used to control the influence of storage consumption in the scheduling optimization; P i (s) represents the energy consumption of the i-th data at time s, used to measure the energy consumption of its storage location, and s is the time index. By minimizing the scheduling optimization function, the storage strategy and access priority are dynamically optimized to minimize the storage energy consumption and improve the access efficiency. The output of the scheduling optimization function will regulate the migration of data between different storage media in real time to further ensure the optimization of data access efficiency.
[0056] The results calculated by the scheduling optimization function directly drive a series of data asset management operations:
[0057] 1. Data storage optimization: Based on the output results of the scheduling optimization function, the storage location of each data is adjusted in real time. Data with higher value and more frequent access will be preferentially arranged on storage media with higher performance (such as cache or high-performance storage devices) to ensure response speed; while data with low-frequency access or archived data will be migrated to storage devices with lower costs (such as cold storage or disk backup), thereby reducing the overall storage cost.
[0058] 2. Access priority management: The response time and access frequency parameters included in the scheduling optimization function can automatically adjust the access priority of each data asset. In actual operation, this means that for high-priority data access requests, relevant data will be pre-loaded in advance to shorten the access latency and improve the data response efficiency; it can not only ensure the efficient utilization of key data, but also dynamically manage assets in resource scheduling to avoid resource waste.
[0059] 3. Resource configuration and cost control: According to the energy consumption part in the scheduling optimization function, identify and preferentially schedule storage devices with lower energy consumption to control the overall resource cost. Periodically monitor the power consumption of each storage device, and perform automated resource adjustments for data or storage locations with higher energy consumption to reduce the equipment operation cost while maintaining data availability, so that the storage and use costs of data assets during their life cycle can be effectively controlled, truly realizing the refined and economical management of data assets.
[0060] 4. Asset life cycle management: By continuously analyzing and updating the output of the scheduling optimization function, identify the life cycle stage of the data and decide when to archive or destroy a certain data asset. For example, when the results of the scheduling optimization function show that a certain type of data has been in a low-priority, low-usage and high-storage-cost state for a long time, automatically archive this type of data and set up a destruction process to ensure that data assets are reasonably managed in different life cycle stages, release storage resources, and reduce management costs.
[0061] Under the guidance of the scheduling optimization function, not only the dynamic adjustment of data storage, access and resource costs is realized, but also the state of data in different life cycle stages can be effectively managed, ensuring the all-round optimization of data assets in terms of use, storage and cost control, and fully achieving the theme goal of "asset management".
[0062] In summary, a data asset management method is completed.
[0063] The sequential order of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0064] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0065] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A data asset management method, characterized in that: The following steps are involved: S1. Collect data and introduce a multi-level feature extraction mechanism to generate an initial multi-dimensional attribute vector; project the initial multi-dimensional attribute vector into a high-dimensional embedding space to generate a time-dependent mapping vector; perform adaptive dynamic evaluation on the mapping vector and calculate the energy difference between the data; S2. Based on the energy difference between data, a dynamic threshold adjustment mechanism is introduced to perform dynamic classification and obtain the classification results of the data; based on the classification results of the data, the storage location and access priority of the data are dynamically optimized through the time-optimal data scheduling algorithm.
2. A data asset management method according to claim 1, characterized in that: The S1 specifically includes: The multi-level feature extraction mechanism quantifies the indicators of the data and generates an initial multi-dimensional attribute vector.
3. A data asset management method according to claim 2, characterized in that: The S1 specifically includes: Through the nonlinear mapping formula, the initial multi-dimensional attribute vector is projected into the high-dimensional embedding space to generate a time-related mapping vector. The nonlinear mapping formula is: Among them, Z i (t) represents the mapping vector of the i-th data at time t to the embedding space, and represents the time-related mapping vector; w ik (t) is the time-dependent weight factor of the i-th data in the k-th dimension at time t; α k is the frequency factor in the sine function; x ik is the original attribute value of the i-th data in the k-th dimension; β k is the phase shift in the sine function; γ ik (t) is the time-related index mapping coefficient of the i-th data in the k-th dimension at time t; λ k is the growth rate parameter of the exponential function in the kth dimension; n is the total number of dimensions.
4. A data asset management method according to claim 3, characterized in that: The S1 specifically includes: The dynamic nonlinear energy equation is introduced to perform adaptive dynamic evaluation on the mapping vector and calculate the energy difference between the data.
5. A data asset management method according to claim 1, characterized in that: The S2 specifically includes: In the process of implementing the dynamic threshold adjustment mechanism, the data is dynamically classified based on the energy difference between the data and combined with the classification threshold to obtain the classification result of the data.
6. A data asset management method according to claim 5, characterized in that: The S2 specifically includes: In the process of implementing the dynamic threshold adjustment mechanism, the sine function and integral term are introduced in combination with the data usage frequency to calculate the classification threshold.
7. A data asset management method according to claim 6, characterized in that: The S2 specifically includes: In the process of implementing the time-series optimal data scheduling algorithm, based on the classification results of the data, the scheduling optimization function is constructed and minimized, and the storage location and access priority of the data are adjusted. The scheduling optimization function formula is: in, is the scheduling optimization function; N is the total number of data in the data set; C i (t) represents the classification result of the i-th data at time t; S i (t) is the storage cost of the i-th data; τ i (t) is the response time of the i-th data; R i (t) is the access frequency of the i-th data at time t; μ is the energy cost control coefficient; P i (s) represents the energy consumption of the i-th data at time s; s is the time index.
8. A data asset management method according to claim 7, characterized in that: The S2 specifically includes: Based on the calculation results of the scheduling optimization function, data storage optimization, access priority management, resource allocation and cost control, and asset lifecycle management are performed.