A data catalog intelligent arrangement system based on data asset management
By generating a time-series association path map and call index for data assets, the problem of low efficiency in manual annotation in data asset management is solved, and efficient utilization and precise management of data assets are achieved.
Patent Information
- Application Number
- CN202411437569.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-15
AI Technical Summary
Existing data asset management relies too heavily on manual data labeling, resulting in low efficiency and poor accuracy. Manually labeled data tags cannot effectively link enterprise data assets, leading to low value in data asset cataloging.
Data asset information is acquired through the data acquisition module, an initial data asset pool is established, a time-series correlation path graph is generated, and a data asset call index is generated by combining project schedule requirements and call frequency using the random forest algorithm and data call event model, thereby optimizing the data asset catalog.
Reduce data redundancy, improve the efficiency of data asset utilization, and enhance the management and utilization efficiency of data assets.
Smart Images

Figure CN119597973B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, in particular to a data directory intelligent arrangement system based on data asset management. BACKGROUND
[0002] The data directory intelligent arrangement of data asset management refers to helping users quickly locate, understand and utilize data through automatic metadata collection, intelligent classification and grading, correlation mapping, etc., so as to maximize the value of data assets, and intelligently arrange index directory, realize efficient support for data governance, data analysis and data application, and promote enterprise digital transformation and intelligent upgrading.
[0003] The existing data asset management relies too much on manual data labeling, and the enterprise data volume is usually large, so that the efficiency of manual data labeling is poor and the accuracy is low, secondly, after the data labels labeled by manual labeling are classified, the correlation between enterprise data assets cannot provide actual reference for subsequent projects, resulting in low value of enterprise data asset directory arrangement. SUMMARY
[0004] To solve the above technical problems, a data directory intelligent arrangement system based on data asset management is provided, which solves the above problems that the existing data asset management relies too much on manual data labeling, and the enterprise data volume is usually large, so that the efficiency of manual data labeling is poor and the accuracy is low, secondly, after the data labels labeled by manual labeling are classified, the correlation between enterprise data assets cannot provide actual reference for subsequent projects, resulting in low value of enterprise data asset directory arrangement.
[0005] To achieve the above purposes, the technical scheme adopted by the present application is:
[0006] A data directory intelligent arrangement system based on data asset management, comprising:
[0007] A data acquisition module, the data acquisition module is used for querying the enterprise database based on the enterprise database, acquiring all data asset information attributes in the enterprise, and performing data preprocessing to establish an initialized enterprise data asset pool;
[0008] A data pool module, the data pool module is electrically connected with the data acquisition module, the data pool module is used for binding the initialization enterprise data asset information pool according to the enterprise data generation time stamp based on the initialized enterprise data asset pool, and obtaining an initialized enterprise data asset time sequence pool;
[0009] A relationship path module, the relationship path module is electrically connected with the data pool module, the relationship path module is used for analyzing the correlation between each group of data assets according to the enterprise data link path under the time sequence based on the initialized enterprise data asset time sequence pool, and generating an enterprise data asset time sequence correlation path map.
[0010] Project log module, the project log module is used to obtain the project progress log of each department in the enterprise, and determine the project progress demand of each department in the enterprise;
[0011] Data division module, the data division module is electrically connected with the project log module and the relationship path module, the data division module is used for data division according to the project progress demand of each department in the enterprise, and the data asset time sequence correlation path atlas of the enterprise is obtained.
[0012] Data state analysis module, the data state analysis module is electrically connected with the data division module, the data state analysis module is used for analyzing the state of each department project data asset based on the classified collection of each department project data asset, according to the calling frequency of each department project data, and generating data asset calling frequency index;
[0013] Catalog module, the catalog module is electrically connected with the data state analysis module, the catalog module is used for sorting and establishing data asset catalog based on the data asset calling frequency index in the classified collection of each department project data asset.
[0014] Preferably, the relationship path module comprises:
[0015] Scatter plot visualization unit, based on the initialized enterprise data asset time sequence pool, the enterprise data generation link path under unit time is marked, and the enterprise data generation link path scatter plot is established;
[0016] Correlation analysis unit, based on the enterprise data generation link path scatter plot, the correlation analysis of each adjacent scatter point of the enterprise data link path in the scatter plot is calculated, and the correlation index scatter plot of the adjacent link path scatter point of the enterprise data is obtained.
[0017] Significance verification unit, based on the correlation index scatter plot of the adjacent link path scatter point of the enterprise data, the significance index of each adjacent scatter point of the enterprise data link path in the correlation index scatter plot is verified, and the significance index scatter plot of the adjacent link path scatter point of the enterprise data is obtained.
[0018] Data division unit, the significance range threshold of each adjacent link path scatter point in the significance index scatter plot of the adjacent link path scatter point of the enterprise data is judged, and the enterprise data asset time sequence correlation path atlas is obtained.
[0019] Preferably, the data division module comprises:
[0020] Decision tree unit, based on Random Forest random forest, the decision tree corresponding to the enterprise department project demand is established, and the random forest model of the project demand of each department in the enterprise is established.
[0021] Preference feature unit, based on the project progress demand of each department of the enterprise, determine the project progress demand preference feature data of each department;
[0022] Preference label unit, according to the decision tree corresponding to the project demand of the enterprise department, take the project progress demand preference feature data of each department as the splitting condition of the leaf node in the random forest model of the project demand of each department of the enterprise, and substitute the root node with the time sequence correlation path atlas of the enterprise data asset, and perform leaf node splitting according to the maximum information gain of the enterprise data under each path, to obtain the project demand preference label data of each department;
[0023] Data packaging unit, data packaging is carried out on the project demand preference label data of each department, and a classified collection of department project data assets is generated.
[0024] Preferably, the data state analysis module comprises:
[0025] Call information collection unit, obtain the historical enterprise data call information log of the project progress of each department; the call information log includes: call time, call department, call category data;
[0026] Call event collection unit, based on the historical enterprise data call information log of the project progress of each department, statistics the call events corresponding to all call information logs in unit time; the call events include data addition, deletion, query and modification;
[0027] Frequency reference calculation unit, based on the historical enterprise data call information log of the project progress of each department and the call events corresponding to all call information logs in unit time, calculate the call event frequency reference value of the enterprise data in the classified collection of department project data assets;
[0028] Model construction unit, according to the call event frequency reference value of the enterprise data in the classified collection of department project data assets, establish the enterprise data call event probability prediction model of each department;
[0029] Event prediction unit, based on the classified collection of department project data assets, substitute the enterprise data call event probability prediction model of each department to generate the enterprise data call event prediction probability of each department in unit time;
[0030] Normalization unit, using the normalization formula, the enterprise data call event prediction probability of each department in unit time is normalized to obtain the relative value of the enterprise data call event prediction probability of each department;
[0031] Index generation unit, taking the relative value of the enterprise data call event prediction probability of each department in unit time as the call attribute of the data asset classified collection, generating the data asset call frequency index;
[0032] The enterprise data calling event probability prediction model of each department is specifically:
[0033] ;
[0034] In the formula, is the relative value of the calling event prediction probability of the i-th department in the j-th time period calling the k-th enterprise data, is the calling event prediction probability value of the i-th department in the j-th time period calling the k-th enterprise data, is the calling event reference value of the i-th department in the j-th time period calling the k-th enterprise data, is the number of calling events of the i-th department in the j-th time period calling the k-th enterprise data, is the duration of the j-th time period, is the total number of data assets of the i-th department, is an exponential function, is a linear regression coefficient.
[0035] Compared with the prior art, the beneficial effects of the present application are that:
[0036] The present application proposes a data directory intelligent arrangement scheme based on data asset management, generates an enterprise data asset time sequence correlation path atlas by analyzing the production cycle and reference link path of enterprise data, divides data assets according to the project progress requirements of each department of the enterprise, determines the data classification set of each department, and creates a data asset calling index and establishes a data asset directory according to the data asset calling frequency of each department. The present application has the advantages of reducing data redundancy and improving the utilization efficiency of data assets. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 It is a data directory intelligent arrangement system framework based on data asset management;
[0038] Figure 2 It is an internal framework diagram of the relationship path module;
[0039] Figure 3 It is an internal framework diagram of the data division module;
[0040] Figure 4 The internal framework diagram of the data state analysis module. DETAILED DESCRIPTION
[0041] The following description is provided to enable those skilled in the art to practice the invention. The preferred embodiments described below are only examples of the invention and are not intended to limit the scope of the invention.
[0042] Referring to Figure 1 As shown in the figure, a data catalog intelligent arrangement system based on data asset management includes:
[0043] A data acquisition module is used to obtain all data asset information attributes in the enterprise based on enterprise database queries, and to perform data preprocessing and establish an initial enterprise data asset pool.
[0044] A data pool module is electrically connected to the data acquisition module, and is used to obtain an initial enterprise data asset time sequence pool based on the initial enterprise data asset pool, and to associate and bind the enterprise data generation time stamp with the initial enterprise data asset information pool.
[0045] A relationship path module is electrically connected to the data pool module, and is used to analyze the correlation between each group of data assets based on the initial enterprise data asset time sequence pool and the enterprise data link path under the time sequence, and to generate an enterprise data asset time sequence correlation path map.
[0046] A project log module is used to obtain project progress logs of each department in the enterprise, and to determine the project progress requirements of each department in the enterprise.
[0047] A data division module is electrically connected to the project log module and the relationship path module, and is used to divide data based on the enterprise data asset time sequence correlation path map according to the project progress requirements of each department in the enterprise, and to obtain a classified collection of department project data assets.
[0048] A data state analysis module is electrically connected to the data division module, and is used to analyze the state of department project data assets based on the classified collection of department project data assets, and to generate a data asset call frequency index according to the call frequency of department project data.
[0049] A catalog module is electrically connected to the data state analysis module, and is used to sort and organize a data asset catalog based on the data asset call frequency index in the classified collection of department project data assets.
[0050] The scheme generates an enterprise data asset time sequence correlation path map by analyzing the production cycle and reference link path of enterprise data, divides the data assets according to the project progress requirements of each department of the enterprise, determines the data classification set of each department, and creates a data asset call index and a data asset directory according to the data asset call frequency of each department.
[0051] Referring to Figure 2 As shown in the figure, the relationship path module internally includes:
[0052] The scatter plot visualization unit marks the enterprise data generation link path under unit time based on the initialized enterprise data asset time sequence pool, and establishes an enterprise data generation link path scatter plot;
[0053] The correlation analysis unit performs correlation analysis on each adjacent scatter point in the enterprise data link path scatter plot based on the enterprise data generation link path scatter plot, and obtains a correlation index scatter plot of the adjacent link path scatter points of the enterprise data;
[0054] The significance verification unit verifies the significance index of each adjacent scatter point in the correlation index scatter plot of the adjacent link path scatter points of the enterprise data based on the correlation index scatter plot of the adjacent link path scatter points of the enterprise data, and obtains a significance index scatter plot of the adjacent link path scatter points of the enterprise data;
[0055] The data division unit divides each adjacent link path scatter point in the significance index scatter plot of the adjacent link path scatter points of the enterprise data according to the significance range threshold value where the adjacent link path scatter point is located, and obtains an enterprise data asset time sequence correlation path map.
[0056] It should be noted that the significance range threshold value is used to verify the correlation strength between adjacent link path scatter points, and the significance range threshold value is set as follows: |r|>0.95: significant correlation, |r|>=0.8: high correlation, 0.5<=|r|<0.8: moderate correlation, 0.3<=|r|<0.5: low correlation, |r|<0.3: weak correlation; the enterprise data corresponding to the adjacent link path scatter points is segmented through significance verification, and the accuracy of data division is improved;
[0057] It should be noted that the link path of the adjacent scatter points in the enterprise data generation link path scatter plot usually includes one or more, and therefore, in the present scheme, when the significance verification segmentation is performed, the cross-adjacent path is segmented into one or more link paths to ensure the integrity of the data.
[0058] Referring to Figure 3 As shown in the figure, the data division module internally includes:
[0059] The decision tree unit establishes a decision tree corresponding to the project demand of each department of the enterprise based on a Random Forest random forest, and forms a Random Forest model of the project demand of each department of the enterprise.
[0060] The preference feature unit determines the project progress demand preference feature data of each department based on the project progress demand of each department of the enterprise.
[0061] The preference label unit takes the project progress demand preference feature data of each department as the splitting condition of the leaf node in the Random Forest model of the project demand of each department of the enterprise, and takes the enterprise data time sequence correlation path atlas as the root node, and performs leaf node splitting on the maximum information gain of the enterprise data under each path to obtain the project progress demand preference label data of each department.
[0062] The data packaging unit packages the project progress demand preference label data of each department to generate a classified collection of project data assets of each department.
[0063] The scheme constructs a decision tree model of the project demand of each department of the enterprise through a Random Forest algorithm, and performs leaf node splitting using the preference feature data of the project progress demand to maximize the information gain and optimize the model, thereby accurately classifying and predicting the project demand. The beneficial effects are to improve the decision accuracy, optimize the resource allocation, enhance the project management capability, and promote the data-driven enterprise decision-making mode.
[0064] Referring to Figure 4 As shown in the figure, the data state analysis module internally includes:
[0065] The call information collection unit acquires historical enterprise data call information logs of the project progress of each department; the call information logs include call time, call department, and call category data.
[0066] The call event collection unit, based on the historical enterprise data call information logs of the project progress of each department, counts the call events corresponding to all call information logs in a unit time; the call events include data addition, deletion, query, and modification.
[0067] The frequency benchmark calculation unit, based on the historical enterprise data call information logs of the project progress of each department and the call events corresponding to all call information logs in a unit time, calculates the call event frequency benchmark value of the enterprise data in the classified collection of project data assets of each department.
[0068] The model construction unit establishes an enterprise data call event probability prediction model of each department according to the call event frequency benchmark value of the enterprise data in the classified collection of project data assets of each department.
[0069] The event prediction unit will input the project data asset classification set of each department into the enterprise data call event probability prediction model of each department to generate the enterprise data call event prediction probability of each department per unit time.
[0070] The normalization unit uses a normalization formula to normalize the predicted probability of enterprise data call events for each department within a unit of time, thereby obtaining the relative value of the predicted probability of enterprise data call events for each department.
[0071] The index generation unit uses the relative values of the predicted probabilities of enterprise data call events for each department within a unit of time as the call attributes of the data asset classification set to generate a data asset call frequency index.
[0072] Specifically, the enterprise data access event probability prediction model for each department is as follows:
[0073] ;
[0074] In the formula, For the first The department is in the first Call the first time period Relative probability of data retrieval events for individual enterprises For the first The department is in the first Call the first time period Predicted probability value of data retrieval events for an individual enterprise. For the first The department is in the first Call the first time period Baseline value for data retrieval events for individual enterprises. For the first The department is in the first Call the first time period Number of times enterprise data is accessed. For the first Duration of the time period For the first Total number of data assets of each department It is an exponential function. These are the linear regression coefficients.
[0075] The principle behind this solution is to collect and analyze historical enterprise data access logs for project progress across departments, statistically calculate baseline values for the frequency of data asset access events for each department, and establish an enterprise data access event probability prediction model based on this. This model can predict the probability of enterprise data access events for each department within a future unit of time, and through normalization processing, obtain relative values, ultimately generating a data asset access frequency index. This provides support for enterprises to optimize data asset management and data utilization efficiency.
[0076] The use process of the present application is:
[0077] Step 1: Based on the enterprise database query, all data asset information attributes in the enterprise are obtained, data preprocessing is carried out, and an initialized enterprise data asset pool is established;
[0078] Step 2: Based on the initialized enterprise data asset pool, the initialization enterprise data asset information pool is associated and bound according to the enterprise data generation time stamp, and an initialized enterprise data asset time sequence pool is obtained;
[0079] Step 3: Based on the initialized enterprise data asset time sequence pool, the enterprise data generation link path under the unit time is marked, and an enterprise data generation link path scatter plot is established;
[0080] Step 4: Based on the enterprise data generation link path scatter plot, the correlation analysis of the adjacent scatter points of each enterprise data link path in the scatter plot is carried out, and an enterprise data adjacent link path scatter point correlation index scatter plot is obtained;
[0081] Step 5: Based on the enterprise data adjacent link path scatter point correlation index scatter plot, the significance index of each enterprise data link path adjacent scatter point in the correlation index scatter plot is verified, and an enterprise data adjacent link path scatter point significance index scatter plot is obtained;
[0082] Step 6: The significance range threshold of each adjacent link path scatter point in the enterprise data adjacent link path scatter point significance index scatter plot is judged for data division, and an enterprise data asset time sequence correlation path atlas is obtained;
[0083] Step 7: Obtain the project progress log of each department in the enterprise, and determine the project progress demand of each department in the enterprise;
[0084] Step 8: Based on the Random Forest, an enterprise department project demand corresponding decision tree is established, and an enterprise department project demand random forest model is formed;
[0085] Step 9: Based on the project progress demand of each department, the project progress demand preference feature data of each department is determined;
[0086] Step 10: According to the enterprise department project demand corresponding decision tree, the project progress demand preference feature data of each department is taken as the splitting condition of the leaf node in the enterprise department project demand random forest model, the enterprise data asset time sequence correlation path atlas is taken as the root node, and the maximum information gain of the enterprise data under each path is taken for leaf node splitting, and the project progress demand preference label data of each department is obtained;
[0087] Step 11: data packaging is performed on the demand preference label data of each department project, and a classified collection of data assets of each department project is generated;
[0088] Step 12: historical enterprise data calling information logs of each department project progress are obtained;
[0089] Step 13: based on the historical enterprise data calling information logs of each department project progress, all calling information logs corresponding to calling events in a unit time are counted;
[0090] Step 14: based on the historical enterprise data calling information logs of each department project progress and the calling events corresponding to all calling information logs in a unit time, a calling event frequency benchmark value of enterprise data in the classified collection of data assets of each department project is calculated;
[0091] Step 15: according to the calling event frequency benchmark value of enterprise data in the classified collection of data assets of each department project, an enterprise data calling event probability prediction model of each department is established;
[0092] Step 16: the classified collection of data assets of each department is substituted into the enterprise data calling event probability prediction model of each department to generate enterprise data calling event prediction probability of each department in a unit time;
[0093] Step 17: the normalized formula is used to normalize the enterprise data calling event prediction probability of each department in a unit time, and the relative value of the enterprise data calling event prediction probability of each department is obtained;
[0094] Step 18: the relative value of the enterprise data calling event prediction probability of each department in a unit time is taken as a calling attribute of the classified collection of data assets, and a data asset calling frequency index is generated;
[0095] Step 19: the data asset calling frequency index in the classified collection of data assets of each department project is sorted to establish a data asset directory.
[0096] In summary, the advantages of the present application are that: by analyzing the production cycle and reference link path of enterprise data, an enterprise data asset time sequence correlation path atlas is generated, the data assets are divided according to the progress demand of each department of the enterprise, the data classification collection of each department is determined, and the data asset calling index is created according to the data asset calling frequency of each department, and the data asset directory is established.
[0097] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only the principles of the present application. Various changes and improvements can be made without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A data catalog intelligent arrangement system based on data asset management, characterized in that, include: The data acquisition module is used to query the enterprise database, obtain the information attributes of all data assets within the enterprise, perform data preprocessing, and establish an initial enterprise data asset pool. The data pool module is electrically connected to the data acquisition module. The data pool module is used to obtain the initial enterprise data asset time-series pool based on the initial enterprise data asset pool and to associate and bind the enterprise data with the initial enterprise data asset information pool according to the generation timestamp of the enterprise data. The relationship path module is electrically connected to the data pool module. The relationship path module is used to analyze the correlation between each group of data assets based on the initialized enterprise data asset time-series pool and the enterprise data link path under time sequence, and generate an enterprise data asset time-series correlation path map. The project log module is used to obtain project progress logs from various departments within the enterprise and determine the project progress requirements of each department. The data partitioning module is electrically connected to the project log module and the relationship path module. The data partitioning module is used to partition the data according to the project progress requirements of each department of the enterprise and the time-series association path graph of the enterprise's data assets to obtain the classification set of project data assets of each department. The data status analysis module is electrically connected to the data partitioning module. The data status analysis module is used to analyze the status of the data assets of each department based on the classification set of project data assets of each department and according to the calling frequency of project data of each department, and generate a data asset calling frequency index. The data status analysis module includes: The event prediction unit will input the project data asset classification set of each department into the enterprise data call event probability prediction model of each department to generate the enterprise data call event prediction probability of each department per unit time. The index generation unit uses the relative values of the predicted probabilities of enterprise data call events for each department within a unit of time as the call attributes of the data asset classification set to generate a data asset call frequency index. Specifically, the enterprise data access event probability prediction model for each department is as follows: ; In the formula, For the first The department is in the first Call the first time period Relative probability of data retrieval events for individual enterprises For the first The department is in the first Call the first time period Predicted probability value of data retrieval events for an individual enterprise. For the first The department is in the first Call the first time period Baseline value for data retrieval events for individual enterprises. For the first The department is in the first Call the first time period Number of times enterprise data is accessed. For the first Duration of the time period For the first Total number of data assets of each department It is an exponential function. These are the linear regression coefficients; The catalog module is electrically connected to the data status analysis module. The catalog module is used to sort and build a data asset catalog based on the data asset call frequency index in the data asset classification set of each department's project data assets.
2. The intelligent data cataloging system based on data asset management according to claim 1, characterized in that, The relational path module includes: The scatter visualization unit, based on the initialized enterprise data asset time series pool, marks the enterprise data generation link path under a unit of time and establishes a scatter plot of the enterprise data generation link path. The correlation analysis unit generates a scatter plot of link paths based on enterprise data, calculates the correlation analysis of adjacent scatter points of each enterprise data link path in the scatter plot, and obtains a scatter plot of correlation index of adjacent link path scatter points of enterprise data. The significance verification unit verifies the significance index of each scatter point adjacent to the enterprise data link path in the correlation index scatter plot based on the correlation index scatter plot of the enterprise data adjacent link path scatter plot, and obtains the significance index scatter plot of the enterprise data adjacent link path scatter plot. Data is divided into units, and the significance index of the scatter points of adjacent link paths of enterprise data is determined. The data is divided according to the significance range threshold of each scatter point of adjacent link path in the scatter plot, and the time-series correlation path map of enterprise data assets is obtained.
3. The intelligent data cataloging system based on data asset management according to claim 2, characterized in that, The data partitioning module includes: The decision tree unit, based on Random Forest, establishes decision trees corresponding to the project requirements of enterprise departments, and constructs a Random Forest model of project requirements for each department of the enterprise. The preference feature unit determines the preference feature data of project schedule requirements for each department based on the project schedule requirements of each department in the enterprise. The preference label unit, based on the decision tree corresponding to the project requirements of the enterprise departments, uses the project progress requirement preference feature data of each department as the splitting condition of the leaf nodes in the random forest model of the project requirements of each department. The enterprise data asset time series association path map is used as the root node, and the leaf nodes are split by the maximum information gain of the enterprise data under each path to obtain the requirement preference label data of each department's project. The data packaging unit packages the demand preference tag data of each department's projects into a data asset classification set for each department's projects.
4. The intelligent data cataloging system based on data asset management according to claim 3, characterized in that, The data status analysis module also includes: The information collection unit is invoked to obtain historical enterprise data on project progress from various departments, including the call information log. The call information log includes: call time, calling department, and call category data. The event collection unit calls the historical enterprise data call information logs based on the project progress of each department, and counts the call events corresponding to all call information logs within a unit of time; the call events include data addition, deletion, query, and modification; The frequency benchmark calculation unit calculates the frequency benchmark value of the enterprise data call events in the project data asset classification set of each department based on the historical enterprise data call information logs of each department's project progress and the call events corresponding to all call information logs under each unit of time. The model building unit establishes a probability prediction model for enterprise data call events for each department based on the benchmark value of the call event frequency in the enterprise data asset classification set of each department's project. The normalization unit uses a normalization formula to normalize the predicted probability of enterprise data retrieval events for each department within a unit of time, thus obtaining the relative value of the predicted probability of enterprise data retrieval events for each department.
Citation Information
Patent Citations
Intelligent identification system for hierarchical relationship of data assets
CN118410405A