Data asset classification method and system based on big data
Through the big data-based data asset classification method, through standardized processing, correlation analysis and knowledge graph construction, the problem of enterprise classification and management of data assets in digital transformation is solved, efficient utilization and security management of data assets are achieved, and data-driven growth of enterprises is promoted.
Patent Information
- Application Number
- CN202510071947.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of digital transformation, enterprises lack detailed classification and management of data assets, resulting in inefficient data circulation and difficult to mine data value.
The data asset classification method based on big data is adopted, and through standardized processing, correlation analysis, knowledge graph construction and other steps, efficient classification of multi-source data assets is achieved, and the classification of data assets is dynamically updated to match business needs.
It improves the utilization rate of data assets, strengthens data security management, promotes data compliance, supports data analysis and decision-making, thereby promoting enterprise data-driven business growth and innovation.
Smart Images

Figure CN119989129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and in particular to a data asset classification method and system based on big data. Background Art
[0002] With the rapid development of information technology and Internet technology, big data technology has become an important means of digital transformation for enterprises. However, in the process of digital transformation, managers' understanding of company assets usually remains based on the traditional definition of assets, and they lack the understanding of the value of the company's intangible data assets. Problems in the management and utilization of enterprise data assets still exist.
[0003] Enterprises lack detailed classification and management of data assets, resulting in low data circulation efficiency and difficulty in mining data value. Traditional data management methods can no longer meet the growing data management needs of enterprises, so a data asset classification management method based on big data technology is needed to help enterprises better manage and utilize their data assets.
[0004] The present invention provides a method and system for enterprise data asset classification management based on big data. The method realizes efficient classification of multi-source data assets through standardized processing, association analysis, knowledge graph construction and other steps, and can dynamically update the classification of data assets to ensure the matching degree between classification results and business needs. This method can not only improve the utilization rate of data assets and strengthen data security management, but also promote data compliance, support data analysis and decision-making, and thus promote data-driven business growth and innovation of enterprises. Summary of the invention
[0005] The purpose of the present invention is to provide a data asset classification method and system based on big data.
[0006] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0007] A first aspect of the present invention provides a data asset classification method based on big data, comprising:
[0008] S1 obtains multi-source data assets of the business system and performs standardized processing;
[0009] S2 performs association analysis on the standardized multi-source data assets to construct a knowledge graph; obtains the flow of data assets in the business system and dynamically updates the knowledge graph;
[0010] S3 determines key data assets based on the usage frequency of data assets, obtains the association path and association strength of key data assets from the knowledge graph, and predicts the prospects of key data assets;
[0011] S4 clusters the prospect prediction results according to data complexity and calculates the average value of the data asset evaluation of each cluster as the cluster center indicator;
[0012] S5 classifies all data assets according to the cluster center index to form a data asset portrait;
[0013] S6 isolates all the data asset portraits from each other, and there is only a unique connection between them. It also follows the principle of minimizing storage to obtain the life cycle of the data asset portraits and performs distributed storage.
[0014] As a further method, the method of performing association analysis on the standardized multi-source data assets to construct a knowledge graph includes:
[0015] Perform association rule mining on standardized multi-source data assets to extract association relationships;
[0016] Obtain the operations corresponding to multi-source data assets in the business process, treat each operation as a state in the Markov model, and use the Markov model to trace the path of the business results;
[0017] According to the path obtained by backtracking, a contribution is assigned to each operation. The closer the operation is to the final business result, the higher the contribution. The expression is:
[0018]
[0019] Among them, C(o i ) is represented by o i The contribution of i is the ith operation in the business process, e is a constant, ψ is the attenuation factor, which is 0.01, o f For the final business result operation, d(o i ,o f ) is the operation o i With o f The distance between the two in the business process, Cost(o i ) and T(o i ) are respectively operation o i cost and execution time;
[0020] The knowledge graph model is constructed by taking multi-source data assets as nodes of the knowledge graph, taking association relationships as edges, and taking the contribution of operations as attributes of corresponding nodes.
[0021] As a further method, the method of obtaining the data asset flow in the business system and dynamically updating the knowledge graph includes:
[0022] Real-time monitoring of the flow of data assets in business systems, including the creation, modification, access, and deletion of data assets;
[0023] According to the monitored data asset flow information, the entities and relationships in the knowledge graph are updated, and the model parameters in the knowledge graph are updated using the incremental learning algorithm. The expression is:
[0024]
[0025] Among them, θ t+1 and θ t are the model parameters at time step t+1 and time step t respectively, α is the learning rate, which is used to control the step size of each parameter update, It represents the loss function L with respect to the model parameter θ at the t+1th update t The gradient of X t+1 is the feature vector set at the t+1th update, Y t+1 is the set of true label values corresponding to the t+1th update, and λ is the regularization coefficient;
[0026] Perform consistency checks on the updated knowledge graph to check whether the edge connections comply with business process rules, and delete edges between data assets that do not comply with business process rules.
[0027] As a further method, the method of determining key data assets based on the usage frequency of data assets, obtaining the association path and association strength of key data assets from the knowledge graph, and predicting the prospects of key data assets includes:
[0028] The contribution of each node in the knowledge graph is normalized from 0 to 1, and the difference in contribution between nodes is used as the association strength between nodes, and the association direction is from high-contribution nodes to low-contribution nodes;
[0029] Data assets whose usage frequency is greater than the average value within the preset time window are regarded as key data assets;
[0030] Starting from the key data asset node, the heuristic search algorithm is used to search in the knowledge graph to obtain the association path. Specifically, if the association strength between two segments of nodes is less than 0.4 and the association direction is consistent, then the three nodes are judged to be on the same association path;
[0031] The association path and association strength are used as the characteristics of key data assets to build a spatiotemporal graph convolutional network model, and the historical time series feature information is obtained to train the model. The trained model is used to predict the characteristics of key data assets. The expression is:
[0032] h t =σ(Wio x t +W ho h t-1 +b o )⊙tanh(σ(W if x t +W hf h t-1 +b f )⊙c t-1
[0033] +σ(W ii x t +W hi h t-1 +b i )⊙tanh(W ig x t +W hg h t-1 +b g ))
[0034] Among them, h t is the hidden state representation of the key data asset feature at time step t, σ is the sigmoid function, and W io , W if , W ii and W ig The input layer weight matrices of the output gate, forget gate, input gate, and memory unit are d×d in dimension, where d is the hidden layer dimension and x is the input layer weight matrix of the output gate, forget gate, input gate, and memory unit. t is the input at time step t, W ho , W hf , W hi and W hg They are the hidden layer weight matrices of the output gate, forget gate, input gate, and memory unit, with dimensions d and h. t-1 is the hidden layer state at time step t-1, b o 、b f 、b i and b g are the bias vectors of the output gate, forget gate, input gate and memory unit respectively, and tanh is the hyperbolic tangent function.
[0035] As a further method, the method of clustering the prospect prediction results according to data complexity includes:
[0036] Obtain forecast results for all key data assets;
[0037] The data complexity of calculating the prospect prediction results is expressed as:
[0038]
[0039] Where D is the feature set in the prospect prediction result, ξ, β, γ, δ and ε are the weight coefficients of the number of features, feature difference, feature information entropy, feature distribution and feature time series in the prediction result, n is the number of D, ζ is the weight coefficient of i and j are the i-th and j-th eigenvalues in D, respectively, ω ij For i and j The association weight between them, H(D) is the information entropy of D, μ and τ are the mean and standard deviation of feature set D, T is the length of the time series, ζ t and t-1 are the eigenvalues in D at time steps t and t-1, respectively, and Δt is the interval of time steps;
[0040] The data complexity is normalized from 0 to 1, and key data assets with data complexity difference percentage less than 0.1 are clustered into one category.
[0041] As a further method, the method of calculating the average value of the data asset evaluation of each cluster as a cluster center indicator includes:
[0042] Obtain the business usage scenarios of all key data assets in each cluster;
[0043] Calculate the data asset evaluation value of key data assets in each business usage scenario. The expression is:
[0044]
[0045] Among them, S ij is the data asset evaluation value of key data asset i in business usage scenario j, T is the observation time range, and f ijt is the usage frequency of key data asset i in business usage scenario j at time step t, β is the periodic change parameter of usage frequency, L is the number of data assets associated with key data asset i in the knowledge graph, ω il is the strength of association between key data asset i and the lth associated data asset, M il is the number of key data assets i and the lth associated data asset used in the same business usage scenario, A il is the sum of the number of business usage scenarios of key data asset i and the lth associated data asset currently used, D i is the data redundancy of key data asset i, P i Score the data quality of key data asset i;
[0046] For each cluster, the average value of the data asset evaluation of key data assets in all business usage scenarios is obtained as the central indicator of the cluster to characterize the overall value level of the key data assets in the cluster.
[0047] As a further method, the method of classifying all data assets according to the cluster center index to form a data asset portrait includes:
[0048] The cosine similarity method is used to calculate the similarity between the data asset evaluation values of all data assets and each cluster center indicator in each business usage scenario;
[0049] Using the calculated similarity, each data asset is assigned to the category represented by the most similar cluster center indicator through the K-nearest neighbor algorithm;
[0050] Assign corresponding attributes to the classified data assets to form a data asset portrait, which includes the frequency of use, correlation strength and future prediction information of the data assets.
[0051] As a further method, the method of isolating all the data asset portraits from each other and having only a unique connection between them includes:
[0052] Store each data asset profile in a different storage partition, and use access control lists to set different access permissions for it, allowing only specific data asset classification-related roles and servers to access the corresponding storage partitions;
[0053] Based on the classification characteristics of data asset portraits, a hash algorithm is used to generate a unique contact identifier to uniquely map data asset portraits.
[0054] The connection identifiers between data asset portraits are only called when cross-classification data asset portrait association analysis is required and when it is used to verify the security and integrity of different data asset classification information.
[0055] As a further method, the method of obtaining the life cycle of the data asset portrait in accordance with the principle of minimizing storage and performing distributed storage includes:
[0056] The objective function is constructed with the goal of minimizing storage, and the expression is:
[0057]
[0058] Among them, s i is the storage size of the i-th data asset portrait, l i is the life cycle of the i-th data asset portrait, r ijis the access frequency of the ith data asset portrait on the jth day, p is the cost per unit storage capacity, q is the number of data asset portraits, e is the storage overhead coefficient for establishing a unique connection between data asset portraits, δ ik is the strength of association established between the i-th data asset portrait and the k-th data asset portrait, and c is the cost coefficient of data storage;
[0059] Under the constraints of the business process, the life cycle that optimizes the objective function is obtained, and data asset portraits that exceed the life cycle are compressed;
[0060] Set a coefficient that is proportional to the access frequency of the data asset portrait to control the life cycle of the data asset portrait. When the access frequency changes, the life cycle of the data asset portrait changes proportionally.
[0061] Distribute various types of data asset portraits to different storage nodes for distributed storage.
[0062] A second aspect of the present invention provides a data asset classification system based on big data, comprising:
[0063] The data processing module is used to obtain multi-source data assets of the business system and perform standardized processing;
[0064] A graph construction and update module is used to perform association analysis on the standardized multi-source data assets to construct a knowledge graph; obtain the flow of data assets in the business system and dynamically update the knowledge graph;
[0065] A key data prediction module is used to determine key data assets based on the usage frequency of data assets, obtain the association path and association strength of key data assets from the knowledge graph, and perform prospect prediction on key data assets;
[0066] The cluster center module is used to cluster the prospect prediction results according to the data complexity and calculate the average value of the data asset evaluation of each cluster as the cluster center indicator;
[0067] A data asset classification module is used to classify all data assets according to the cluster center index to form a data asset portrait;
[0068] The storage isolation management module is used to perform mutual isolation for all the data asset portraits, with only a unique connection between them, and follow the principle of minimizing storage to obtain the life cycle of the data asset portraits for distributed storage.
[0069] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0070] (1) By constructing a knowledge graph and dynamically updating it, the present invention can reflect the flow and changes of data assets in real time, ensure the timeliness and accuracy of data assets, help enterprises better understand and utilize data assets, and improve the efficiency and accuracy of data-driven decision-making;
[0071] (2) The present invention determines key data assets based on the frequency of data asset usage, obtains the association paths and association strengths of key data assets from the knowledge graph, and performs prospect prediction, which helps enterprises to identify and respond to potential data needs and risks in advance and optimize the management and utilization of data assets;
[0072] (3) The present invention classifies all data assets according to cluster centers to form data asset portraits, and ensures that the portraits are isolated from each other and connected only by unique connection functions, which helps to better manage and protect data assets and prevent data leakage and abuse;
[0073] (4) The present invention follows the principle of minimizing storage, obtains the life cycle of data assets, and performs distributed storage, which not only optimizes the use of storage resources and reduces storage costs, but also improves the scalability and reliability of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 The present invention is a flowchart of a method for classifying data assets based on big data in an embodiment of the present invention. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0076] Reference Figure 1 As shown, the present invention provides a data asset classification method based on big data, comprising:
[0077] S1 obtains multi-source data assets of the business system and performs standardized processing;
[0078] In the actual evaluation, in an implementation case applied to e-commerce enterprises, data acquisition was carried out, specifically: nearly 20,000 product information was extracted from the product management system, covering detailed fields such as product name, ID, category, brand, price, inventory, description, and image link; nearly 90 million order data in the past year were obtained from the order processing system, including order number, user ID, product ID, order time, order amount, payment method, delivery address, and order status; 5.4 million user data were extracted from the user management system, including user ID, name, gender, age, registration time, login frequency, historical purchase category preferences, and consumption amount range information; nearly 70 million logistics data were collected from the logistics distribution system, including logistics order number, order ID, shipping time, distribution node, estimated delivery time, actual delivery time, and distribution status;
[0079] In the actual evaluation, the product names were cleaned to remove special characters and garbled characters, and the character encoding was unified to UTF-8. The order time in the order data was converted to a timestamp format, the age in the user data was divided into age groups, and the distribution node names in the logistics data were standardized.
[0080] S2 performs association analysis on the standardized multi-source data assets to construct a knowledge graph; obtains the flow of data assets in the business system and dynamically updates the knowledge graph;
[0081] It is necessary to explain that, for standardized multi-source data assets, association analysis algorithms are used to deeply explore the intrinsic connections between data. For example, in medical data, the connections between symptoms, diagnosis results, treatment plans and basic patient information can be found. These data entities are used as nodes and the association relationships are used as edges to construct a knowledge graph, which can intuitively display the complex relationships between data, provide a structured basis for subsequent analysis, and help understand and utilize data assets from a holistic perspective.
[0082] In the actual evaluation, the Apriori algorithm was used to mine association rules on the standardized data, including: for users who buy "mobile phones", there is a 30% probability that they will buy "mobile phone cases" and "earphones" at the same time, and the association rule is: mobile phone → mobile phone cases, earphones (support: 0.05, confidence: 0.3); it was found that during the big promotion, the average customer unit price of the goods purchased by users will increase by 40%, and the association rule is: big promotion → customer unit price increase (support: 0.1, confidence: 0.4); in the order processing process, the operations of order creation, inventory check, payment confirmation, order delivery, logistics distribution, and order receipt are defined as the state of the Markov model. Take a successfully completed order as an example, its path is: order creation → inventory check (in stock) → payment confirmation (success) → order delivery → logistics distribution (normal) → order receipt. The cost of the inventory check operation is 0.2 yuan, and the execution time is 1 minute; the cost of the payment confirmation operation is 3 yuan, and the execution time is 2 minutes. Calculate the contribution of each operation, among which the contribution of inventory check is 1.535; take the goods, orders, users, and logistics data as nodes of the knowledge graph, take the mined associations, including the purchase relationship between goods and users, the distribution relationship between orders and logistics, etc. as edges, and take the contribution of each operation as the attribute of the corresponding node to build a knowledge graph model;
[0083] In the actual evaluation, the business system is monitored in real time. When data changes are detected, the entities and relationships in the knowledge graph are updated in time. The model parameters in the knowledge graph are updated using the incremental learning algorithm, where the set of true label values is the corresponding business results (whether the order is successful). The regularization coefficient is 0.05. In one day's monitoring, 43,000 product information updates, 300,000 new order creations, 2,000 user information modifications, and 150,000 logistics status updates were captured. The knowledge graph was updated and a consistency check was performed on the updated knowledge graph. Through predefined business rules (such as canceled orders can no longer have a shipping status), 500 edges that did not comply with the rules were found and deleted.
[0084] S3 determines key data assets based on the usage frequency of data assets, obtains the association path and association strength of key data assets from the knowledge graph, and predicts the prospects of key data assets;
[0085] It is necessary to explain that the association paths and association strengths of key data assets are obtained from the knowledge graph. By analyzing these association paths and association strengths, we can better understand the role and impact of data assets in business processes, thereby providing a basis for the management and optimization of data assets.
[0086] It should be understood that the future forecast of key data assets is aimed at evaluating the potential value and risks of key data assets in future business development. Through predictive analysis, enterprises can plan the utilization and protection strategies of data assets in advance to support the continuous development and innovation of the business.
[0087] In the actual evaluation, the usage frequency of each data asset is obtained, and the data assets with a usage frequency higher than the average are defined as key data assets. The contribution of each node in the knowledge graph is normalized to make its value between 0 and 1. For the key data assets of the popular mobile phone product node, the heuristic search algorithm is used to search in the knowledge graph, and it is found that the correlation strength between the mobile phone product node and the mobile phone accessories product node is 0.3, and the correlation direction is from mobile phone products to accessories products. At the same time, the correlation strength between the mobile phone product node and the high-spending user node is 0.25, and the correlation direction is consistent. It is judged that these three nodes are on the same association path, that is, popular mobile phone products → mobile phone accessories products → high-spending users;
[0088] In the actual evaluation, the association path and association strength are used as the characteristics of key data assets, and a spatiotemporal graph convolutional network model is constructed. The historical time series feature information of the past two years (monthly mobile phone sales, user search popularity, and competitor product dynamics) is collected to train the model. The trained model is used to predict the key data assets and predict the sales of a popular mobile phone in the next month. It is assumed that the market share of the current mobile phone, the intensity of recent promotional activities and other characteristic values are input at the time step. The hidden state representation of the key data asset characteristics at the time step is obtained through model calculation, and it is predicted that the sales of the mobile phone in the next month are expected to increase by 15%.
[0089] S4 clusters the prospect prediction results according to data complexity and calculates the average value of the data asset evaluation of each cluster as the cluster center indicator;
[0090] It needs to be explained that according to the differences in these complexities, clustering algorithms are used to group the prospect prediction results, and data with similar complexity are grouped into one category, so that subsequent targeted management and analysis of prediction results of different categories can be carried out, which can more clearly divide the data into layers and improve the efficiency and pertinence of data asset processing;
[0091] In the actual evaluation, the prospect prediction results of all key data assets are collected and the data complexity of the prospect prediction results is calculated. For popular clothing products, the feature set of the prospect prediction results includes sales volume, price fluctuations, seasonal factors, fashion trends and other features. The data complexity is calculated and normalized. The key data assets with a data complexity difference percentage of less than 0.1 are clustered into one category. Specifically, the data complexity of several popular fashion clothing items of the season is similar after normalization and is clustered into one category.
[0092] In the actual evaluation, the business usage scenarios of the key data assets in each cluster are determined, and the data asset valuation of the key data assets in each business usage scenario is calculated. For a certain popular clothing in the new product promotion scenario, the observation time range is one month after the new product is launched. The usage frequency of the clothing (daily views, collections, and orders), the periodic change parameters of the usage frequency (fitted according to historical new product promotion data), and the number of data assets associated with the clothing (including model wear displays, other products of the same brand, etc.) are 10, of which the data redundancy is 0.12, and the data asset valuation of the clothing in the new product promotion scenario is 28.1; then for each cluster, the average data asset valuation of the key data assets in all business usage scenarios is calculated as the cluster center indicator.
[0093] S5 classifies all data assets according to the cluster center index to form a data asset portrait;
[0094] It needs to be explained that after the cluster center index is obtained, all data assets are classified, each data asset is compared with each cluster center index, and the data assets are classified into corresponding categories according to the similarity. Attributes are given to the data assets in each category, such as frequency of use, association strength, etc. These attributes are combined to form a data asset portrait, which is used to clearly display the characteristics of data assets;
[0095] It should be understood that data asset portraits include information such as the attributes, value, usage, and relationship with other data assets of data assets, which helps to more intuitively understand the characteristics and potential value of each data asset and provide support for data-driven decision-making;
[0096] In the actual evaluation, the cosine similarity is used to calculate the similarity of the data asset evaluation values of all data assets and each cluster center indicator in each business usage scenario. The K-nearest neighbor algorithm is used to assign each data asset to the category represented by the most similar cluster center indicator. Among them, the new clothing data mentioned above is assigned to the popular clothing category because it has the highest similarity with the popular clothing cluster center indicator. The classified data assets are assigned corresponding attributes to form a data asset portrait. For a high-consumption user, his data asset portrait includes usage frequency (10 logins and 3 orders per month), association strength (the association strength with high-end products is 0.6, and the association strength with popular brands is 0.5) and future forecast information (the consumption amount is expected to increase by 20% in the next three months).
[0097] S6 isolates all the data asset portraits from each other, and there is only a unique connection between them. It also follows the principle of minimizing storage to obtain the life cycle of the data asset portraits and performs distributed storage.
[0098] It needs to be explained that mutual isolation ensures that each data asset profile is stored independently, preventing data leakage and unauthorized access, enhancing data security, and protecting the integrity and privacy of data assets;
[0099] It should be understood that a special association is established between data asset portraits. This association is unique and directional. This connection helps to track and understand the interaction between data assets when necessary, while maintaining their isolation.
[0100] In the actual evaluation, each data asset portrait is stored in a different storage partition, where high-value user data asset portraits are stored in a high-performance storage partition using an all-flash array to ensure fast access; historical order data asset portraits are stored in a storage partition composed of large-capacity mechanical hard disks to reduce costs; different access rights are set using access control lists, and only senior managers and data analysis experts in the marketing department can access the high-value user data asset portrait storage partition; financial department personnel can access the historical order data asset portrait storage partition for financial accounting and auditing. Based on the classification characteristics of the data asset portrait, the SHA-256 hash algorithm is used to generate a unique contact identifier, where the contact identifier between the high-value user data asset portrait and the historical order data asset portrait is "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08";
[0101] In the actual evaluation, the objective function is constructed with the goal of minimizing storage. The optimal life cycle of high-value user data asset portraits is 15 months. If the life cycle is exceeded, a compression algorithm is used. At the same time, a coefficient is set that is proportional to the frequency of access to the data asset portrait. When the access frequency of a historical order data asset portrait increases from an average of 20 times per day to 40 times per day, its life cycle is adjusted to twice the previous one, and each type of data asset portrait is distributed to different storage nodes for distributed storage.
[0102] Specifically, the effect comparison before and after the data asset classification method is applied is shown in Table 1:
[0103] Table 1 Comparison of experimental results
[0104]
[0105] In this embodiment, the method of performing association analysis on the standardized multi-source data assets to construct a knowledge graph includes:
[0106] Perform association rule mining on standardized multi-source data assets to extract association relationships;
[0107] Obtain the operations corresponding to multi-source data assets in the business process, treat each operation as a state in the Markov model, and use the Markov model to trace the path of the business results;
[0108] According to the path obtained by backtracking, a contribution is assigned to each operation. The closer the operation is to the final business result, the higher the contribution. The expression is:
[0109]
[0110] Among them, C(o i ) is represented by o i The contribution of i is the ith operation in the business process, e is a constant, ψ is the attenuation factor, which is 0.01, o f For the final business result operation, d(o i ,o f ) is the operation o i With o f The distance between the two in the business process, Cost(o i ) and T(o i ) are respectively operation o i cost and execution time;
[0111] The knowledge graph model is constructed by taking multi-source data assets as nodes of the knowledge graph, taking association relationships as edges, and taking the contribution of operations as attributes of corresponding nodes.
[0112] In this embodiment, the method of obtaining the data asset flow in the business system and dynamically updating the knowledge graph includes:
[0113] Real-time monitoring of the flow of data assets in business systems, including the creation, modification, access, and deletion of data assets;
[0114] According to the monitored data asset flow information, the entities and relationships in the knowledge graph are updated, and the model parameters in the knowledge graph are updated using the incremental learning algorithm. The expression is:
[0115]
[0116] Among them, θ t+1 and θ t are the model parameters at time step t+1 and time step t respectively, α is the learning rate, which is used to control the step size of each parameter update, It represents the loss function L with respect to the model parameter θ at the t+1th update t The gradient of X t+1 is the feature vector set at the t+1th update, Yt+1 is the set of true label values corresponding to the t+1th update, and λ is the regularization coefficient;
[0117] Perform consistency checks on the updated knowledge graph to check whether the edge connections comply with business process rules, and delete edges between data assets that do not comply with business process rules.
[0118] In this embodiment, the method of determining key data assets based on the usage frequency of data assets, obtaining the association path and association strength of key data assets from the knowledge graph, and predicting the prospects of key data assets includes:
[0119] The contribution of each node in the knowledge graph is normalized from 0 to 1, and the difference in contribution between nodes is used as the association strength between nodes, and the association direction is from high-contribution nodes to low-contribution nodes;
[0120] Data assets whose usage frequency is greater than the average value within the preset time window are regarded as key data assets;
[0121] Starting from the key data asset node, the heuristic search algorithm is used to search in the knowledge graph to obtain the association path. Specifically, if the association strength between two segments of nodes is less than 0.4 and the association direction is consistent, then the three nodes are judged to be on the same association path;
[0122] The association path and association strength are used as the characteristics of key data assets to build a spatiotemporal graph convolutional network model, and the historical time series feature information is obtained to train the model. The trained model is used to predict the characteristics of key data assets. The expression is:
[0123] h t =σ(W io x t +W ho h t-1 +b o )⊙tanh(σ(W if x t +W hf h t-1 +b f )⊙c t-1
[0124] +σ(W ii x t +W hi h t-1 +b i )⊙tanh(W ig x t +W hg h t-1 +b g ))
[0125] Among them, h t is the hidden state representation of the key data asset feature at time step t, σ is the sigmoid function, and W io , W if , W ii and W ig The input layer weight matrices of the output gate, forget gate, input gate, and memory unit are d×d in dimension, where d is the hidden layer dimension and x is the input layer weight matrix of the output gate, forget gate, input gate, and memory unit. t is the input at time step t, W ho , W hf , W hi and W hg They are the hidden layer weight matrices of the output gate, forget gate, input gate, and memory unit, with dimensions d and h. t-1 is the hidden layer state at time step t-1, b o 、b f 、b i and b g are the bias vectors of the output gate, forget gate, input gate and memory unit respectively, and tanh is the hyperbolic tangent function.
[0126] In this embodiment, the method for clustering the prospect prediction results according to data complexity includes:
[0127] Obtain forecast results for all key data assets;
[0128] The data complexity of calculating the prospect prediction results is expressed as:
[0129]
[0130] Where D is the feature set in the prospect prediction result, ξ, β, γ, δ and ε are the weight coefficients of the number of features, feature difference, feature information entropy, feature distribution and feature time series in the prediction result, n is the number of D, ζ is the weight coefficient of i and j are the i-th and j-th eigenvalues in D, respectively, ω ij For i and j The association weight between them, H(D) is the information entropy of D, μ and τ are the mean and standard deviation of feature set D, T is the length of the time series, ζ t and t-1 are the eigenvalues in D at time steps t and t-1, respectively, and Δt is the interval of time steps;
[0131] The data complexity is normalized from 0 to 1, and key data assets with data complexity difference percentage less than 0.1 are clustered into one category.
[0132] In this embodiment, the method of calculating the average value of the data asset evaluation of each cluster as the cluster center indicator includes:
[0133] Obtain the business usage scenarios of all key data assets in each cluster;
[0134] Calculate the data asset evaluation value of key data assets in each business usage scenario. The expression is:
[0135]
[0136] Among them, S ij is the data asset evaluation value of key data asset i in business usage scenario j, T is the observation time range, and f ijt is the usage frequency of key data asset i in business usage scenario j at time step t, β is the periodic change parameter of usage frequency, L is the number of data assets associated with key data asset i in the knowledge graph, ω il is the strength of association between key data asset i and the lth associated data asset, M il is the number of key data assets i and the lth associated data asset used in the same business usage scenario, A il is the sum of the number of business usage scenarios of key data asset i and the lth associated data asset currently used, D i is the data redundancy of key data asset i, P i Score the data quality of key data asset i;
[0137] For each cluster, the average value of the data asset evaluation of key data assets in all business usage scenarios is obtained as the central indicator of the cluster to characterize the overall value level of the key data assets in the cluster.
[0138] In this embodiment, the method of classifying all data assets according to the cluster center index to form a data asset portrait includes:
[0139] The cosine similarity method is used to calculate the similarity between the data asset evaluation values of all data assets and each cluster center indicator in each business usage scenario;
[0140] Using the calculated similarity, each data asset is assigned to the category represented by the most similar cluster center indicator through the K-nearest neighbor algorithm;
[0141] Assign corresponding attributes to the classified data assets to form a data asset portrait, which includes the frequency of use, correlation strength and future prediction information of the data assets.
[0142] In this embodiment, the method of isolating all the data asset portraits from each other and having only one connection between them includes:
[0143] Store each data asset profile in a different storage partition, and use access control lists to set different access permissions for it, allowing only specific data asset classification-related roles and servers to access the corresponding storage partitions;
[0144] Based on the classification characteristics of data asset portraits, a hash algorithm is used to generate a unique contact identifier to uniquely map data asset portraits.
[0145] The connection identifiers between data asset portraits are only called when cross-classification data asset portrait association analysis is required and when it is used to verify the security and integrity of different data asset classification information.
[0146] In this embodiment, the method of obtaining the life cycle of the data asset portrait and performing distributed storage in accordance with the principle of minimizing storage includes:
[0147] The objective function is constructed with the goal of minimizing storage, and the expression is:
[0148]
[0149] Among them, s i is the storage size of the i-th data asset portrait, l i is the life cycle of the i-th data asset portrait, r ij is the access frequency of the ith data asset portrait on the jth day, p is the cost per unit storage capacity, q is the number of data asset portraits, e is the storage overhead coefficient for establishing a unique connection between data asset portraits, δ ik is the strength of association established between the i-th data asset portrait and the k-th data asset portrait, and c is the cost coefficient of data storage;
[0150] Under the constraints of the business process, the life cycle that optimizes the objective function is obtained, and data asset portraits that exceed the life cycle are compressed;
[0151] Set a coefficient that is proportional to the access frequency of the data asset portrait to control the life cycle of the data asset portrait. When the access frequency changes, the life cycle of the data asset portrait changes proportionally.
[0152] Distribute various types of data asset portraits to different storage nodes for distributed storage.
[0153] The second aspect of the present invention also provides a data asset classification system based on big data, comprising:
[0154] The data processing module is used to obtain multi-source data assets of the business system and perform standardized processing;
[0155] A graph construction and update module is used to perform association analysis on the standardized multi-source data assets to construct a knowledge graph; obtain the flow of data assets in the business system and dynamically update the knowledge graph;
[0156] A key data prediction module is used to determine key data assets based on the usage frequency of data assets, obtain the association path and association strength of key data assets from the knowledge graph, and perform prospect prediction on key data assets;
[0157] The cluster center module is used to cluster the prospect prediction results according to the data complexity and calculate the average value of the data asset evaluation of each cluster as the cluster center indicator;
[0158] A data asset classification module is used to classify all data assets according to the cluster center index to form a data asset portrait;
[0159] The storage isolation management module is used to perform mutual isolation for all the data asset portraits, with only a unique connection between them, and follow the principle of minimizing storage to obtain the life cycle of the data asset portraits for distributed storage.
[0160] The above contents are merely examples and explanations of the structure of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, they should all fall within the protection scope of the present invention.
Claims
1. A data asset classification method based on big data, characterized in that: The following steps are involved: Acquire multi-source data assets of business systems and perform standardized processing; Performing association analysis on the standardized multi-source data assets to construct a knowledge graph; obtaining the flow of data assets in the business system to dynamically update the knowledge graph; Determine key data assets based on the usage frequency of data assets, obtain the association path and association strength of key data assets from the knowledge graph, and make prospect predictions for key data assets; The prospect prediction results are clustered according to the data complexity, and the average value of the data asset evaluation of each cluster is calculated as the cluster center indicator; Classify all data assets according to the cluster center indicators to form a data asset portrait; All the data asset portraits are isolated from each other and have only a unique connection between them. The life cycle of the data asset portraits is obtained in accordance with the principle of minimizing storage and distributed storage is performed.
2. According to the data asset classification method based on big data in claim 1, it is characterized in that: The method of performing association analysis on the standardized multi-source data assets to construct a knowledge graph includes: Perform association rule mining on standardized multi-source data assets to extract association relationships; Obtain the operations corresponding to multi-source data assets in the business process, treat each operation as a state in the Markov model, and use the Markov model to trace the path of the business results; According to the path obtained by backtracking, a contribution is assigned to each operation. The closer the operation is to the final business result, the higher the contribution. The expression is: Among them, C(o i ) is represented by o i The contribution of i is the ith operation in the business process, e is a constant, ψ is the attenuation factor, which is 0.01, o f For the final business result operation, d(o i , o f ) is the operation o i With o f The distance between the two in the business process, Cost(o i ) and T(o i ) are respectively operation o i cost and execution time; The knowledge graph model is constructed by taking multi-source data assets as nodes of the knowledge graph, taking association relationships as edges, and taking the contribution of operations as attributes of corresponding nodes.
3. The data asset classification method based on big data according to claim 1 is characterized in that: The method of obtaining the data asset flow in the business system and dynamically updating the knowledge graph includes: Real-time monitoring of the flow of data assets in business systems, including the creation, modification, access, and deletion of data assets; According to the monitored data asset flow information, the entities and relationships in the knowledge graph are updated, and the model parameters in the knowledge graph are updated using the incremental learning algorithm. The expression is: Among them, θ t+1 and θ t are the model parameters at time step t+1 and time step t respectively, α is the learning rate, which is used to control the step size of each parameter update, It represents the loss function L with respect to the model parameter θ at the t+1th update t The gradient of X t+1 is the feature vector set at the t+1th update, Y t+1 is the set of true label values corresponding to the t+1th update, and λ is the regularization coefficient; Perform consistency checks on the updated knowledge graph to check whether the edge connections comply with business process rules, and delete edges between data assets that do not comply with business process rules.
4. The data asset classification method based on big data according to claim 1 is characterized in that: The method of determining key data assets based on the usage frequency of data assets, obtaining the association path and association strength of key data assets from the knowledge graph, and predicting the prospects of key data assets includes: The contribution of each node in the knowledge graph is normalized from 0 to 1, and the difference in contribution between nodes is used as the association strength between nodes, and the association direction is from high-contribution nodes to low-contribution nodes; Data assets whose usage frequency is greater than the average value within the preset time window are regarded as key data assets; Starting from the key data asset node, the heuristic search algorithm is used to search in the knowledge graph to obtain the association path. Specifically, if the association strength between two segments of nodes is less than 0.4 and the association direction is consistent, then the three nodes are judged to be on the same association path; The association path and association strength are used as the characteristics of key data assets to build a spatiotemporal graph convolutional network model, and the historical time series feature information is obtained to train the model. The trained model is used to predict the characteristics of key data assets. The expression is: Among them, h t is the hidden state representation of the key data asset feature at time step t, σ is the sigmoid function, and W io , W if , W ii and W ig The input layer weight matrices of the output gate, forget gate, input gate, and memory unit are d×d in dimension, where d is the hidden layer dimension and x is the input layer weight matrix of the output gate, forget gate, input gate, and memory unit. t is the input at time step t, W ho , W hf , W hi and W hg They are the hidden layer weight matrices of the output gate, forget gate, input gate, and memory unit, with dimensions d and h. t-1 is the hidden layer state at time step t-1, b o 、b f 、b i and b g are the bias vectors of the output gate, forget gate, input gate and memory unit respectively, and tanh is the hyperbolic tangent function.
5. The data asset classification method based on big data according to claim 1 is characterized in that: The method for clustering the prospect prediction results according to data complexity includes: Obtain forecast results for all key data assets; The data complexity of calculating the prospect prediction results is expressed as: Where D is the feature set in the prospect prediction result, ξ, β, γ, δ and ε are the weight coefficients of the number of features, feature difference, feature information entropy, feature distribution and feature time series in the prediction result, n is the number of D, ζ is the weight coefficient of i and j are the i-th and j-th eigenvalues in D, respectively, ω ij For i and j The association weight between them, H(D) is the information entropy of D, μ and τ are the mean and standard deviation of feature set D, T is the length of the time series, ζ t and t-1 are the eigenvalues in D at time steps t and t-1, respectively, and Δt is the interval of time steps; The data complexity is normalized from 0 to 1, and key data assets with data complexity difference percentage less than 0.1 are clustered into one category.
6. The data asset classification method based on big data according to claim 1 is characterized in that: The method of calculating the average value of the data asset evaluation of each cluster as a cluster center indicator includes: Obtain the business usage scenarios of all key data assets in each cluster; Calculate the data asset evaluation value of key data assets in each business usage scenario. The expression is: Among them, S ij is the data asset evaluation value of key data asset i in business usage scenario j, T is the observation time range, and f ijt is the usage frequency of key data asset i in business usage scenario j at time step t, β is the periodic change parameter of usage frequency, L is the number of data assets associated with key data asset i in the knowledge graph, ω il is the strength of association between key data asset i and the lth associated data asset, M il is the number of key data assets i and the lth associated data asset used in the same business usage scenario, A il is the sum of the number of business usage scenarios of key data asset i and the lth associated data asset currently used, D i is the data redundancy of key data asset i, P i Score the data quality of key data asset i; For each cluster, the average value of the data asset evaluation of key data assets in all business usage scenarios is obtained as the central indicator of the cluster to characterize the overall value level of the key data assets in the cluster.
7. The data asset classification method based on big data according to claim 1 is characterized in that: The method of classifying all data assets according to the cluster center index to form a data asset portrait includes: The cosine similarity method is used to calculate the similarity between the data asset evaluation values of all data assets and each cluster center indicator in each business usage scenario; Using the calculated similarity, each data asset is assigned to the category represented by the most similar cluster center indicator through the K-nearest neighbor algorithm; Assign corresponding attributes to the classified data assets to form a data asset portrait, which includes the frequency of use, correlation strength and future prediction information of the data assets.
8. The data asset classification method based on big data according to claim 1 is characterized in that: The method of isolating all the data asset portraits from each other and having only one connection between them includes: Store each data asset profile in a different storage partition, and use access control lists to set different access permissions for it, allowing only specific data asset classification-related roles and servers to access the corresponding storage partitions; Based on the classification characteristics of data asset portraits, a hash algorithm is used to generate a unique contact identifier to uniquely map data asset portraits. The connection identifiers between data asset portraits are only called when cross-classification data asset portrait association analysis is required and when it is used to verify the security and integrity of different data asset classification information.
9. The data asset classification method based on big data according to claim 1 is characterized in that: The method of obtaining the life cycle of the data asset portrait in accordance with the principle of minimizing storage and performing distributed storage includes: The objective function is constructed with the goal of minimizing storage, and the expression is: Among them, s i is the storage size of the i-th data asset portrait, l i is the life cycle of the i-th data asset portrait, r ij is the access frequency of the ith data asset portrait on the jth day, p is the cost per unit storage capacity, q is the number of data asset portraits, e is the storage overhead coefficient for establishing a unique connection between data asset portraits, δ ik is the strength of association established between the i-th data asset portrait and the k-th data asset portrait, and c is the cost coefficient of data storage; Under the constraints of the business process, the life cycle that optimizes the objective function is obtained, and data asset portraits that exceed the life cycle are compressed; Set a coefficient that is proportional to the access frequency of the data asset portrait to control the life cycle of the data asset portrait. When the access frequency changes, the life cycle of the data asset portrait changes proportionally. Distribute various types of data asset portraits to different storage nodes for distributed storage.
10. A data asset classification system based on big data, used to execute a data asset classification method based on big data according to any one of claims 1 to 9, characterized in that: The system comprises: The data processing module is used to obtain multi-source data assets of the business system and perform standardized processing; A graph construction and update module is used to perform association analysis on the standardized multi-source data assets to construct a knowledge graph; obtain the flow of data assets in the business system and dynamically update the knowledge graph; A key data prediction module is used to determine key data assets based on the usage frequency of data assets, obtain the association path and association strength of key data assets from the knowledge graph, and perform prospect prediction on key data assets; The cluster center module is used to cluster the prospect prediction results according to the data complexity and calculate the average value of the data asset evaluation of each cluster as the cluster center indicator; A data asset classification module is used to classify all data assets according to the cluster center index to form a data asset portrait; The storage isolation management module is used to perform mutual isolation for all the data asset portraits, with only a unique connection between them, and follow the principle of minimizing storage to obtain the life cycle of the data asset portraits for distributed storage.
Citation Information
Cited By
Communication engineering informatization system and method based on big data
CN120672548A
Full-life-cycle intelligent operation method and system for data assets
CN120811900A
Targeted partitioning method and system for water network system
CN120911918A
Data asset analysis method and device, storage medium and program product
CN121478870A
Data asset analysis method and apparatus, and program product
CN121478871A