A data management method and system based on data governance and data value extraction
By building a data governance system and deep learning methods, we have solved the problems of data silos and data quality, achieved centralized data management and deep value mining, improved data accuracy and sharing efficiency, and provided enterprises with more comprehensive data support.
Patent Information
- Application Number
- CN202510040760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-18
- Filing Date
- 2025-01-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing data governance and data value extraction technologies have serious problems such as data silos, inconsistent data formats, lack of unified standards, difficulty in ensuring data quality, and traditional methods are time-consuming and labor-intensive and difficult to meet real-time analysis and decision-making needs.
Build a data governance system, including a data collection module, a data quality screening and processing module, and a data resource pool construction module. Use deep learning methods to extract features from embedded expressions, combine data collection strategies, quality assessment and verification, and external data source supplementation to achieve centralized data management and deep value mining.
It realizes the centralized management and efficient use of data, improves the accuracy, integrity and consistency of data, enhances the efficiency of data sharing and circulation, improves the accuracy and efficiency of data value extraction, and provides more comprehensive data support for enterprises.
Smart Images

Figure CN119961496B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management technology, and in particular to a data management method and system based on data governance and data value extraction. Background Art
[0002] With the rapid development of information technology, enterprises have accumulated unprecedented amounts of data resources. This data, encompassing a variety of forms, including structured business data, unstructured documents, and social media content, provides a rich source of material for corporate decision-making, business optimization, and market insights. However, effectively managing this data and extracting valuable information from it has become a major challenge for enterprises.
[0003] Data governance, as a key means to ensure data quality and enhance data value, has received widespread attention in recent years. Traditional data governance methods usually include data collection, data cleaning, data conversion and data storage. However, these methods have exposed a series of problems in actual application: First, data sharing between departments is not smooth, resulting in serious data silos and difficulty in fully utilizing the value of data. Second, data sources are diverse and formats are different, and there is a lack of unified data standards and quality control mechanisms, which affects the accuracy and reliability of data. Finally, faced with massive amounts of data, traditional data processing methods are often unable to cope with the needs of real-time analysis and decision-making.
[0004] Data value extraction is the process of extracting valuable information from data. It is crucial for a company's strategic planning and business optimization. However, existing data value extraction technologies also face many challenges: how to extract useful features from massive amounts of data is a major problem in data value extraction. Traditional methods often rely on manually designed feature engineering, which is time-consuming and labor-intensive and difficult to guarantee results.
[0005] To sum up, although the existing data governance and data value extraction technologies have solved the problems of data management and value mining to a certain extent, they still have many limitations. Summary of the Invention
[0006] In view of this, the present invention proposes a data management method and system based on data governance and data value extraction, which can achieve comprehensive data governance and in-depth value mining.
[0007] The technical solution of the present invention is achieved as follows:
[0008] A data management method based on data governance and data value extraction, specifically including:
[0009] Constructing a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module;
[0010] Construct a data value extraction system that uses individual indicator data as the smallest unit of data to convert it into an embedded expression. This system then uses deep learning methods to extract features from the embedded expression to extract the value of the data.
[0011] Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy;
[0012] The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing;
[0013] The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification.
[0014] As a further optional solution to the data management method based on data governance and data value extraction, the data collection module collects structured business data and unstructured data according to the data collection strategy, specifically including:
[0015] Clarify the goals and requirements of data collection and select appropriate data sources;
[0016] Collect structured data from corresponding data sources based on API interface technology or database synchronization technology;
[0017] Unstructured data is collected from corresponding data sources based on text mining technology, natural language processing technology, or web crawler technology.
[0018] As a further optional solution to the data management method based on data governance and data value extraction, the data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing, specifically including:
[0019] After data collection, the data were checked for missing values, outliers, and duplicate values;
[0020] Use ETL tools to monitor data quality in real time;
[0021] Generate data quality reports regularly to assess trends and changes in data quality;
[0022] Introducing external data sources or third-party analytical models;
[0023] Supplement data that cannot be shared or is temporarily missing based on external data sources or third-party analysis models.
[0024] As a further optional solution to the data management method based on data governance and data value extraction, when supplementing the data that cannot be shared or is temporarily missing based on external data sources or third-party analysis models, data integration and conversion are required, specifically including:
[0025] Identify and finalize any external data sources or third-party analytical models that need to be incorporated;
[0026] Select the corresponding data access method to extract data based on the type and access rights of the external data source or third-party analysis model;
[0027] Preprocess the extracted data;
[0028] Establish field mapping relationships based on field information from different data sources or third-party analysis models;
[0029] Based on the field mapping relationship, the pre-processed data is merged into a unified storage system.
[0030] As a further optional solution to the data management method based on data governance and data value extraction, the data resource pool includes:
[0031] Data resource architecture, which determines how data is organized, stored, and accessed;
[0032] Data resource management, which ensures data integrity, consistency, and availability, supports data access and sharing, and meets business needs and data regulatory requirements;
[0033] Data resource optimization, used to optimize data storage structure, query statements and indexes;
[0034] Data resource security, which is used to ensure the confidentiality, integrity, and availability of data, and to prevent data leakage, tampering, and damage;
[0035] The data resource architecture includes:
[0036] Data model design, used to describe data entities, attributes and their relationships;
[0037] Data partitioning strategy is used to split a data set into multiple data blocks, each of which is called a partition;
[0038] Data indexing strategy is used to quickly access data in database tables.
[0039] As a further optional solution to the data management method based on data governance and data value extraction, the deep learning method includes using a supervised learning model and training the model with a manually labeled data set so that the model can automatically learn the characteristic expressions of various enterprise indicators based on the input, and use the input and output to update the weights in the deep learning network layer.
[0040] As a further optional solution of the data management method based on data governance and data value extraction, the method further includes:
[0041] The integrated iterative update system updates the data resource pool, indicator system and data value extraction model accordingly based on business changes and changes in data resources.
[0042] A data management system based on data governance and data value extraction, including:
[0043] The first construction module is used to build a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module;
[0044] The second building block is used to build a data value extraction system. The data value extraction system uses individual indicator data as the smallest unit of data to convert it into an embedded expression, which is used as the input of the deep learning model. The deep learning method is used to extract features from the embedded expression to achieve data value extraction;
[0045] Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy;
[0046] The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing;
[0047] The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification.
[0048] A computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned data management method based on data governance and data value extraction are implemented.
[0049] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned data management method based on data governance and data value extraction.
[0050] The beneficial effects of the present invention are: through a clear data collection strategy, structured and unstructured data from different sources can be efficiently collected, which helps to break down data silos and realize centralized management and utilization of data. It can process various types of data, including text, images, audio, etc., to meet diverse data needs. Through a complete data quality monitoring system, the collected data is strictly evaluated and verified for quality, which helps to ensure the accuracy, completeness and consistency of the data and improve the quality of the data. The data that cannot be shared or is temporarily missing is supplemented to ensure the completeness and availability of the data, which helps to reduce the impact of missing data on subsequent analysis. A data resource pool is constructed based on the data that has completed quality assessment and verification, providing reliable data support for subsequent data analysis and application. A complete indicator system is constructed, which helps to classify, summarize and organize the data, improve the readability and interpretability of the data, and uses individual indicator data as the smallest unit of data for embedded expression conversion, which helps to capture the characteristics of individual data and provide a basis for subsequent feature extraction. As the input of the deep learning model, the deep learning method is used for feature extraction, which can extract deeper feature information and improve the accuracy and efficiency of data value extraction. The deep learning algorithm is used to extract features from embedded expressions, which can automatically learn the inherent laws and characteristics of the data, which helps to discover potential information and patterns in the data and provide strong support for data analysis and application. Through feature extraction, feature information useful for business decisions can be extracted, and the value and utilization of data can be improved. This helps enterprises better understand market dynamics, customer needs, business processes, etc., and provide data support for optimizing business decisions. Through the construction of data governance and data value extraction systems, data integration and utilization are realized, which helps to break data silos, improve data sharing and circulation efficiency, and provide enterprises with more comprehensive and accurate data support. Through the role of data quality screening and processing modules and data resource pool construction modules, the quality and accuracy of data can be ensured, which helps to reduce the impact of data errors and redundancy on subsequent analysis and improve the reliability of data analysis and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 A flow chart of a data management method based on data governance and data value extraction according to the present invention;
[0053] Figure 2 A schematic diagram of the composition of a data management system based on data governance and data value extraction according to the present invention;
[0054] Figure 3 The figure is a schematic diagram of the composition of a computer device of the present invention. DETAILED DESCRIPTION
[0055] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0056] refer to Figures 1 to 3 , a data management method based on data governance and data value extraction, specifically including:
[0057] Constructing a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module;
[0058] Construct a data value extraction system that uses individual indicator data as the smallest unit of data to convert it into an embedded expression. This system then uses deep learning methods to extract features from the embedded expression to extract the value of the data.
[0059] Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy;
[0060] The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing;
[0061] The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification.
[0062] In this embodiment, through a clear data collection strategy, structured and unstructured data from different sources can be efficiently collected, which helps to break down data silos and realize centralized management and utilization of data. It can process various types of data, including text, images, audio, etc., to meet diverse data needs. Through a complete data quality monitoring system, the collected data is strictly evaluated and verified for quality, which helps to ensure the accuracy, completeness and consistency of the data and improve the quality of the data. The data that cannot be shared or is temporarily missing is supplemented to ensure the integrity and availability of the data, which helps to reduce the impact of missing data on subsequent analysis. A data resource pool is built based on the data that has completed quality evaluation and verification, providing reliable data support for subsequent data analysis and application. A complete indicator system is built, which helps to classify, summarize and organize the data and improve the readability and interpretability of the data. Individual indicator data is used as the smallest unit of data for embedded expression conversion, which helps to capture the characteristics of individual data and provide a basis for subsequent feature extraction. As the input of the deep learning model, deep learning methods are used for feature extraction, which can extract deeper feature information and improve the accuracy and efficiency of data value extraction. Deep learning algorithms are used to extract features from embedded expressions, which can automatically learn the inherent laws and characteristics of the data. This helps to discover potential information and patterns in the data and provide strong support for data analysis and application. Through feature extraction, feature information useful for business decisions can be extracted, and the value and utilization of data can be improved. This helps companies better understand market dynamics, customer needs, business processes, etc., and provide data support for optimizing business decisions. Through the construction of data governance and data value extraction systems, data integration and utilization are achieved, which helps to break data silos, improve data sharing and circulation efficiency, and provide enterprises with more comprehensive and accurate data support. Through the role of data quality screening and processing modules and data resource pool construction modules, the quality and accuracy of data can be ensured, which helps to reduce the impact of data errors and redundancy on subsequent analysis and improve the reliability of data analysis and application.
[0063] Preferably, the data collection module collects structured business data and unstructured data according to a data collection strategy, specifically including:
[0064] Clarify the goals and requirements of data collection and select appropriate data sources;
[0065] Collect structured data from corresponding data sources based on API interface technology or database synchronization technology;
[0066] Unstructured data is collected from corresponding data sources based on text mining technology, natural language processing technology, or web crawler technology.
[0067] In this embodiment, through diversified data collection technologies, both structured and unstructured types of data can be collected to meet diverse data needs. Clear goals and demand settings, as well as efficient and standardized data collection technologies, help improve data quality and accuracy. Through API interface technology and database synchronization technology, real-time or quasi-real-time data collection and updating can be achieved, improving the timeliness of data. The collected data can provide enterprises with more comprehensive and accurate information support, help enterprises better understand market dynamics, customer needs, business processes, etc., and provide data support for optimizing business decisions.
[0068] It should be noted that according to data requirements, appropriate data sources should be selected. For structured business data, data sources may include internal enterprise databases, third-party data interfaces, etc. For unstructured data, data sources may include social media, web page text, images, audio, etc.; use API interfaces to obtain data from internal enterprise databases or third-party data interfaces. This method can ensure the real-time and accuracy of data, but it is necessary to ensure the stability and security of API interfaces; use database synchronization technology to synchronize data in internal enterprise databases to data collection systems in real time or regularly. This method is suitable for situations where data needs to be updated in real time; use text mining technology to extract useful information from text data, which includes pre-processing steps such as word segmentation, stop word removal, stemming, and feature extraction, Mining steps such as model training and evaluation use natural language processing technology to extract information from unstructured text, such as named entity recognition, relationship extraction, and event extraction. For unstructured data that needs to be obtained from the Internet, web crawler technology can be used. Web crawlers can automatically capture information on the Internet according to established rules and convert it into a structured data format; according to the selected data collection strategy, configure the corresponding collection tools. For example, for API interface collection, it is necessary to configure the API interface call parameters; for database synchronization, it is necessary to configure the database connection and synchronization rules; for web crawlers, it is necessary to configure crawler rules and parsing logic; execute the collection task according to the predetermined collection strategy and schedule. During the collection process, it is necessary to monitor the execution status of the collection task and the quality of the collection results.
[0069] Preferably, the data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing, specifically including:
[0070] After data collection, the data were checked for missing values, outliers, and duplicate values;
[0071] Use ETL tools to monitor data quality in real time;
[0072] Generate data quality reports regularly to assess trends and changes in data quality;
[0073] Introducing external data sources or third-party analytical models;
[0074] Supplement data that cannot be shared or is temporarily missing based on external data sources or third-party analysis models.
[0075] In this embodiment, through a series of data quality assessment and verification measures, the quality of the data can be significantly improved, ensuring the accuracy, completeness and reliability of the data. High-quality data can provide enterprises with more accurate and comprehensive information support, helping enterprises to make more informed decisions. Automated data quality monitoring and processing tools can improve work efficiency, reduce manual intervention and errors, and through the role of the data quality screening and processing module, the value of the data can be fully explored and utilized, providing strong data support for the business development of the enterprise.
[0076] It should be noted that after data collection, the data quality needs to be assessed, including indicators such as accuracy, completeness, and consistency. Scripts can be written in programming languages such as Python to check for missing values, abnormal values, and duplicate values. A real-time monitoring mechanism should be established to continuously monitor the quality of the data to ensure that the data is always maintained at a high level. ETL tools such as Ta l end and Informat i can be used. ca and Pentaho, etc., to realize data extraction, conversion and loading, as well as quality monitoring; regularly generate data quality reports to evaluate data quality trends and changes, which helps to timely discover data quality problems and take corresponding measures to improve them; establish a data quality feedback mechanism to collect feedback from data users and timely improve data quality management measures; when the platform cannot share or temporarily lacks certain data, it should actively seek external data sources for supplementation, and can obtain the required data by cooperating with third-party institutions, purchasing data services, etc.; in order to improve the efficiency of data processing and analysis, third-party analysis models can be introduced for assistance. These models have more advanced algorithms and technologies, which can help enterprises better tap the value of data; when introducing external data sources or third-party analysis models, data integration and conversion are required, which includes integrating data from different data sources to form a complete data set, and converting the data format for unified analysis and processing.
[0077] Preferably, when supplementing the data that cannot be shared or is temporarily missing based on the external data source or third-party analysis model, the data needs to be integrated and converted, specifically including:
[0078] Identify and finalize any external data sources or third-party analytical models that need to be incorporated;
[0079] Select the corresponding data access method to extract data based on the type and access rights of the external data source or third-party analysis model;
[0080] Preprocess the extracted data;
[0081] Establish field mapping relationships based on field information from different data sources or third-party analysis models;
[0082] Based on the field mapping relationship, the pre-processed data is merged into a unified storage system.
[0083] In this embodiment, by identifying and determining all external data sources or third-party analysis models that need to be merged, data source integration can be achieved to ensure the comprehensiveness and integrity of the data, which helps to solve the problem of data silos and improve data availability; merging data from different sources can increase data diversity and provide more information and perspectives for data analysis; according to the type and access rights of the external data source or third-party analysis model, selecting the corresponding data access method to extract data can ensure data security and compliance, which helps to avoid the risk of data leakage and illegal access; selecting the appropriate data access method can improve the efficiency of data extraction and reduce the time cost of data acquisition; pre-processing the extracted data, including data cleaning, deduplication, format conversion, etc., can ensure the quality and consistency of the data, which helps to reduce data analysis and application. In order to eliminate errors and deviations in the use of data, the preprocessing process can also achieve data standardization, so that data from different sources can be compared and analyzed under a unified format and standard; establishing field mapping relationships based on the field information of different data sources or third-party analysis models can ensure the accuracy of data during the merging process, which helps to avoid confusion and mismatching of data fields; the establishment of field mapping relationships can also ensure the consistency of data after the merger, so that data from different sources can be integrated and analyzed under a unified framework, and the preprocessed data can be merged into a unified storage system based on the field mapping relationship to ensure data integrity, which helps to solve the problems of data missing and inconsistency and improve data availability; the merged data is stored in a unified system, which can improve data access efficiency and reduce the time cost of data query and analysis.
[0084] Preferably, the data resource pool includes:
[0085] Data resource architecture, which determines how data is organized, stored, and accessed;
[0086] Data resource management, which ensures data integrity, consistency, and availability, supports data access and sharing, and meets business needs and data regulatory requirements;
[0087] Data resource optimization, used to optimize data storage structure, query statements and indexes;
[0088] Data resource security, which is used to ensure the confidentiality, integrity, and availability of data, and to prevent data leakage, tampering, and damage;
[0089] The data resource architecture includes:
[0090] Data model design, used to describe data entities, attributes and their relationships;
[0091] Data partitioning strategy is used to split a data set into multiple data blocks, each of which is called a partition;
[0092] Data indexing strategy is used to quickly access data in database tables.
[0093] In this embodiment, by describing data entities, attributes and their relationships, the data model design provides a clear data structure view, which helps to understand the nature and relationships of the data. A reasonable data model design can reduce data redundancy, improve data storage efficiency and query performance, and a good data model design can support complex query requirements, making data analysis more flexible and efficient. Dividing a data set into multiple partitions can significantly improve query performance, because queries can only be performed on relevant partitions without scanning the entire data set. The partitioning strategy makes data easier to manage and expand, and partitions can be dynamically increased or decreased according to demand. Through a reasonable partitioning strategy, data load balancing can be achieved to avoid overloading of a single node. The data index strategy can significantly improve data access speed, making database table queries more efficient. A reasonable index design can support multiple query modes, including single-field queries, combined queries and range queries. Although the index will take up a certain amount of storage space, a reasonable index design can optimize the use of storage space and avoid unnecessary waste. Fee; through data resource management, the integrity of data, that is, the accuracy and consistency of data, can be ensured. Data resource management supports data access and sharing, ensuring that data can be obtained in a timely and accurate manner when needed. Data resource management can manage and configure data according to business needs, ensuring that data can support business decisions and operations. Data resource management can also ensure that data complies with relevant regulatory requirements, such as data privacy protection, data security, and cross-border data transmission; by optimizing the data storage structure, data storage overhead can be reduced and storage efficiency can be improved. Optimizing query statements and indexes can significantly improve query performance and reduce query time. Data resource optimization can reasonably allocate and utilize system resources and improve the overall performance and stability of the system; through data encryption, access control and other means, the confidentiality of data can be ensured not to be leaked. Data resource security policies can prevent data from being tampered with or damaged, ensuring the authenticity and accuracy of data. Data backup and recovery mechanisms can ensure that data can be quickly restored when it is damaged or lost, thereby ensuring data availability.
[0094] Preferably, the deep learning method includes using a supervised learning model and using a manually labeled data set to train the model so that the model can automatically learn the characteristic expressions of various indicators of the enterprise based on the input, and use the input and output to update the weights in the deep learning network layer.
[0095] In this embodiment, individual indicator data is collected from various data sources, which may include databases, files, API interfaces, etc. The collected data is cleaned, including removing duplicate data, processing missing values, correcting erroneous data, etc., to ensure the accuracy and completeness of the data; data from different sources and in different formats are standardized to have a unified format and unit for subsequent processing; according to the characteristics of the data and business needs, appropriate embedding methods are selected, such as word embedding, graph embedding, etc. ng), etc.; convert individual indicator data into embedded expressions. This step usually involves mapping the original data into a high-dimensional vector space so that similar data points are closer in the vector space; optimize the embedded expressions to improve their expressiveness and accuracy. This can be achieved by training the embedding model, such as using a neural network to fine-tune the embedding vector; select an appropriate deep learning model based on business needs and data characteristics, such as a convolutional neural network (CNN), recurrent neural network (RNN), deep neural network (DNN), etc.; use the embedded expression as input to the deep learning model for model training. During the training process, the model will learn how to extract useful features from the embedded expression; use the trained deep learning model to extract features from the embedded expression. This step will generate a set of feature vectors that can represent the value of the data.
[0096] It should be noted that the embedded expression conversion includes:
[0097] Build a vocabulary: For text data, you need to build a vocabulary and assign an integer ID to each unique word. In this way, the text is converted into a sequence of integers.
[0098] Word embedding: The integer sequence is further converted into a dense vector representation of fixed dimension, which is usually called word embedding, including Word2Vec, G loVe and FastText. These vectors can capture the semantic similarity between words;
[0099] Embedding of other data types: For non-text data (such as images, audio, sensor data, etc.), corresponding embedding methods are required. For example, image data can be feature extracted and converted into embedding vectors through convolutional neural networks (CNNs). Audio data can be converted into embedding vectors through Mel-frequency cepstral coefficients (MFCCs) or other feature representations. Sensor data can be converted into embedding vectors through methods such as time series analysis.
[0100] Preferably, the method further comprises:
[0101] The integrated iterative update system updates the data resource pool, indicator system and data value extraction model accordingly based on business changes and changes in data resources.
[0102] In this embodiment, the first step is to clarify the goal of the update, that is, to determine the specific needs brought about by business changes and data resource changes, which includes understanding new business needs, changes in data sources, data quality requirements, etc.; re-organize the business processes to ensure that the updated system can fit the actual business processes, which helps to ensure that the data resource pool, indicator system and data value extraction model are consistent with the business; according to business changes, connect new data sources or integrate existing data sources to ensure the stability of the data source and the accuracy of the data; clean the newly connected data, remove redundant, erroneous and invalid data, and convert the data so that the data format meets the system requirements; verify the cleaned and converted data to ensure that it meets the business rules and quality requirements, and store the verified data in the updated data resource pool.
[0103] Based on business changes, the indicator system is reorganized and redefined to ensure that the indicators can accurately reflect business changes and data characteristics; the calculation formulas and algorithms of the indicators are updated to ensure the accuracy and efficiency of the indicator calculations, and the indicators are optimized to improve the calculation efficiency and response speed of the indicators; the updated indicators are verified to ensure that they comply with business rules and quality requirements, and the verified indicators are published to the system for business use.
[0104] Redesign or select a suitable data value extraction model based on business changes and the characteristics of data resources to ensure that the model can accurately extract valuable information from the data; use the updated data resource pool to train the model, optimize the model, and improve the model's accuracy and generalization capabilities; verify the trained model to ensure that it complies with business rules and quality requirements, and deploy the verified model to the system for data value extraction and analysis.
[0105] It should be noted that regarding the update of data elements and data resource pool system:
[0106] During the data element aggregation stage, the system focuses on the timeliness and dynamic updating of data, regularly processes and updates system data according to project requirements to ensure the timeliness and accuracy of the data, and compares and verifies data during the data update process to ensure the accuracy and consistency of the new data. Through these measures, the comprehensiveness and quality of the data are guaranteed, providing a reliable basis for further analysis and decision-making; in terms of resource pool construction, the system continuously updates and expands the data resource pool according to business needs and changes in data sources to ensure the timeliness and richness of the data. While retaining the latest data elements, it also retains the indicators that have been derived to ensure the continuity and integrity of data analysis. This dynamic maintenance strategy not only enhances the practical value of the data, but also provides more accurate and timely support for business decisions.
[0107] Regarding the business iteration of data value indicators:
[0108] As the environment changes and business needs adjust, the business indicator system for data value extraction needs to be iterated accordingly to maintain its effectiveness. The iteration of the indicator system is mainly aimed at two application scenarios: First, the evaluation of data value is dynamically updated over time. Therefore, in order to cope with the routine update of individual evaluation data, the indicator system needs to replace data regularly; second, the table items and fields in the system may change. For example, new table items may be added according to the business needs of the regulated individuals, and the newly added data structure will make the system face the need to expand the indicator system. In this case, the iteration of the indicator system is necessary to improve the accuracy of data value evaluation. Therefore, when constructing and iteratively updating the indicator system, this project adopted an integrated data processing and deep learning method to calculate the data value, and designed the indicator system iteration for the two major application scenarios of data update of the evaluation object and system table item change; in terms of indicator system iteration and model update, in the data update scenario, for the latest data, data verification is first performed to ensure the quality of the new data can be used for data value evaluation. The verified data is used to update the resource pool according to the processing logic. The updated resource pool data is used for model training to adapt the model to the latest data value distribution. In the scenario of system table changes, this project has designed a complete set of indicator system expansion methods to adapt to the ever-changing business needs. First, according to the adjustment of the table structure, new indicators are designed and the original indicator dimensions are expanded to ensure that all data dimensions in the new business scenario can be fully expressed and analyzed. At the same time, in order to ensure the validity and reliability of the new indicators, when formulating new indicators, special attention is paid to the verification of data quality and the design of processing logic. Next, according to the expanded indicator dimensions, The network structure of the deep learning model is adjusted accordingly. These adjustments may involve increasing or decreasing the number of network layers, adjusting the number of nodes, and updating activation functions. The purpose is to enable the model to more accurately capture the complex patterns in the data. After the adjustment of the model structure is completed, the model is retrained using the updated dataset, and the performance of the model is continuously monitored during the training process. When necessary, the parameters are adjusted and optimized to ensure that the model can achieve the expected results. Finally, the fully trained and verified model will be deployed to actual application scenarios for subsequent data value assessment tasks, so that the entire system can maintain efficient adaptability and continuous effectiveness in the face of business changes.
[0109] A data management system based on data governance and data value extraction, including:
[0110] The first construction module is used to build a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module;
[0111] The second building block is used to build a data value extraction system. The data value extraction system uses individual indicator data as the smallest unit of data to convert it into an embedded expression, which is used as the input of the deep learning model. The deep learning method is used to extract features from the embedded expression to achieve data value extraction;
[0112] Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy;
[0113] The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing;
[0114] The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification.
[0115] A computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned data management method based on data governance and data value extraction are implemented.
[0116] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned data management method based on data governance and data value extraction.
[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A data management method based on data governance and data value extraction, characterized in that: Specifically include: Constructing a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module; Construct a data value extraction system that uses individual indicator data as the smallest unit of data to convert it into an embedded expression. This system then uses deep learning methods to extract features from the embedded expression to extract the value of the data. Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy; The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing; The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification; When supplementing unshareable or missing data with external data sources or third-party analytical models, data integration and transformation are required, including: Identify and finalize any external data sources or third-party analytical models that need to be incorporated; Select the corresponding data access method to extract data based on the type and access rights of the external data source or third-party analysis model; Preprocess the extracted data; Establish field mapping relationships based on field information from different data sources or third-party analysis models; Based on the field mapping relationship, the pre-processed data is merged into a unified storage system; The data resource pool includes: Data resource architecture, which determines how data is organized, stored, and accessed; Data resource management, which ensures data integrity, consistency, and availability, supports data access and sharing, and meets business needs and data regulatory requirements; Data resource optimization, used to optimize data storage structure, query statements and indexes; Data resource security, which is used to ensure the confidentiality, integrity, and availability of data, and to prevent data leakage, tampering, and damage; The data resource architecture includes: Data model design, used to describe data entities, attributes and their relationships; Data partitioning strategy is used to split a data set into multiple data blocks, each of which is called a partition; Data indexing strategy is used to quickly access data in database tables.
2. A data management method based on data governance and data value extraction according to claim 1, characterized in that: The data collection module collects structured business data and unstructured data according to the data collection strategy, specifically including: Clarify the goals and requirements of data collection and select appropriate data sources; Collect structured data from corresponding data sources based on API interface technology or database synchronization technology; Unstructured data is collected from corresponding data sources based on text mining technology, natural language processing technology, or web crawler technology.
3. A data management method based on data governance and data value extraction according to claim 2, characterized in that: The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing, specifically including: After data collection, the data were checked for missing values, outliers, and duplicate values; Use ETL tools to monitor data quality in real time; Generate data quality reports regularly to assess trends and changes in data quality; Introducing external data sources or third-party analytical models; Supplement data that cannot be shared or is temporarily missing based on external data sources or third-party analysis models.
4. A data management method based on data governance and data value extraction according to claim 3, characterized in that: The deep learning method includes using a supervised learning model and using a manually labeled data set to train the model so that the model can automatically learn the characteristic expressions of various indicators of the enterprise based on the input, and use the input and output to update the weights in the deep learning network layer.
5. The data management method based on data governance and data value extraction according to claim 4 is characterized in that: The data management method further includes: The integrated iterative update system updates the data resource pool, indicator system and data value extraction model accordingly based on business changes and changes in data resources.
6. A data management system based on data governance and data value extraction, characterized in that: include: The first construction module is used to build a data governance system, which is used to govern data and includes a data collection module, a data quality screening and processing module, and a data resource pool construction module; The second building block is used to build a data value extraction system. The data value extraction system uses individual indicator data as the smallest unit of data to convert it into an embedded expression, which is used as the input of the deep learning model. The deep learning method is used to extract features from the embedded expression to achieve data value extraction; Wherein, the data collection module collects structured business data and unstructured data according to the data collection strategy; The data quality screening and processing module performs quality assessment and verification on the collected data according to the data quality monitoring system, and supplements the data that cannot be shared or is temporarily missing; The data resource pool construction module constructs a data resource pool and an indicator system based on the data that has completed quality assessment and verification; When supplementing unshareable or missing data with external data sources or third-party analytical models, data integration and transformation are required, including: Identify and finalize any external data sources or third-party analytical models that need to be incorporated; Select the corresponding data access method to extract data based on the type and access rights of the external data source or third-party analysis model; Preprocess the extracted data; Establish field mapping relationships based on field information from different data sources or third-party analysis models; Based on the field mapping relationship, the pre-processed data is merged into a unified storage system; The data resource pool includes: Data resource architecture, which determines how data is organized, stored, and accessed; Data resource management, which ensures data integrity, consistency, and availability, supports data access and sharing, and meets business needs and data regulatory requirements; Data resource optimization, used to optimize data storage structure, query statements and indexes; Data resource security, which is used to ensure the confidentiality, integrity, and availability of data, and to prevent data leakage, tampering, and damage; The data resource architecture includes: Data model design, used to describe data entities, attributes and their relationships; Data partitioning strategy is used to split a data set into multiple data blocks, each of which is called a partition; Data indexing strategy is used to quickly access data in database tables.
7. A computing device, characterized in that It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the data management method based on data governance and data value extraction described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the data management method based on data governance and data value extraction as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Information processing method and system based on big data
CN116932355A
Medical data management method and system based on artificial intelligence and storage medium
CN118609743A