Enterprise Management Data Governance Method Based on Big Data Analysis
By acquiring, preprocessing and building pattern recognition models, dynamic data governance strategies are generated, which solves the problem that traditional data governance models cannot be understood and flexibly used in real time, and realizes dynamic analysis and security guarantees of enterprise management data.
Patent Information
- Application Number
- CN202411386205.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The traditional enterprise data governance model focuses on batch processing and offline computing, which cannot meet the needs of real-time insights and flexible application of massive data, lacks effective means to conduct comprehensive and continuous monitoring and adjustment optimization, and cannot respond to emergencies in a timely manner.
By obtaining the initial multi-source heterogeneous data of the target enterprise, pre-processing, building a pattern recognition model, generating behavior pattern analysis results, and combining the behavior pattern frequency, a dynamic data governance strategy is generated to achieve dynamic data governance.
Real-time monitoring and dynamic analysis of enterprise management data is realized, data security is ensured, complex data scenarios can be effectively dealt with, and data governance is automated and intelligent.
Smart Images

Figure CN119441800B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to an enterprise management data governance method based on big data analysis. Background Art
[0002] With the rapid development of big data technology, the scale, variety, and complexity of enterprise data are constantly increasing. Through advanced big data processing technologies, the enterprise's management level of various types of data can be improved to adapt to the rapidly growing data scale in the big data environment and achieve more efficient and accurate data control.
[0003] However, in the actual application process, with the continuous development of big data technology, the amount of data, data types, and data processing complexity generated by enterprises are all increasing rapidly. The traditional governance mode mainly focuses on batch processing and offline computing, which cannot meet the enterprise's needs for real-time insight and flexible utilization of massive data. Moreover, most of the existing data governance methods stay at the static attributes of the data itself, do not deeply study the dynamic characteristics and behavioral laws in the data flow process, lack effective means for all-round and continuous monitoring, adjustment, and optimization, and cannot respond to emergencies in a timely manner.
[0004] Therefore, it is necessary to provide an enterprise management data governance method based on big data analysis to solve the above technical problems. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an enterprise management data governance method based on big data analysis to solve the problems that the traditional governance mode mainly focuses on batch processing and offline computing, cannot meet the enterprise's needs for real-time insight and flexible utilization of massive data, and most of the existing data governance methods stay at the static attributes of the data itself, do not deeply study the dynamic characteristics and behavioral laws in the data flow process, lack effective means for all-round and continuous monitoring, adjustment, and optimization, and cannot respond to emergencies in a timely manner.
[0006] The enterprise management data governance method based on big data analysis provided by the present invention includes:
[0007] Obtain a target enterprise and collect the initial multi-source heterogeneous data of the target enterprise;
[0008] Receive the initial multi-source heterogeneous data and preprocess the initial multi-source heterogeneous data to generate target multi-source heterogeneous data;
[0009] Obtain the behavior pattern of the target multi-source heterogeneous data, construct a pattern recognition model, and perform corresponding processing on the behavior pattern based on the pattern recognition model to generate a behavior pattern analysis result;
[0010] Generate the dynamic data governance strategy corresponding to the target enterprise based on the analysis result of the behavior pattern and in combination with the behavior pattern frequency.
[0011] Preferably, obtaining the target enterprise and collecting the initial multi-source heterogeneous data of the target enterprise specifically includes:
[0012] Connect to and access the internal database of the target enterprise to obtain internal enterprise data;
[0013] Obtain external enterprise data from the external services of the target enterprise through a data interface;
[0014] Retrieve public web pages and collect information related to the target enterprise on the public web pages to generate enterprise web page information;
[0015] Unify the data formats of the internal enterprise data, external enterprise data, and enterprise web page information to generate the initial multi-source heterogeneous data.
[0016] Preferably, receiving the initial multi-source heterogeneous data and preprocessing the initial multi-source heterogeneous data to generate target multi-source heterogeneous data specifically includes:
[0017] Detect and mark the abnormal data in the initial multi-source heterogeneous data, where the abnormal data includes missing data, duplicate data, and error data;
[0018] Fill in the missing data, and delete the duplicate data and error data to obtain the first multi-source heterogeneous data;
[0019] Clean the unstructured text in the first multi-source heterogeneous data based on natural language processing technology to obtain the second multi-source heterogeneous data;
[0020] Unify the classification of the second multi-source heterogeneous data based on the preset standardization rules to generate the target multi-source heterogeneous data.
[0021] Preferably, obtaining the behavior pattern of the target multi-source heterogeneous data, constructing a pattern recognition model, and performing corresponding processing on the behavior pattern based on the pattern recognition model to generate a behavior pattern analysis result specifically includes:
[0022] Obtain the behavior pattern of the target multi-source heterogeneous data and determine the corresponding behavior vector according to the behavior pattern;
[0023] Construct the pattern recognition model based on machine learning algorithms;
[0024] Perform recognition processing on the behavior vector according to the pattern recognition model to generate recognition behavior features;
[0025] Perform anomaly detection on the identified behavior features to generate the behavior pattern analysis result.
[0026] Preferably, the performing anomaly detection on the identified behavior features to generate the behavior pattern analysis result specifically includes:
[0027] Set the anomaly detection result as the objective function, and the prediction probability corresponding to the anomaly detection result is:
[0028]
[0029] In the formula, represents the prediction probability corresponding to the anomaly detection result; represents the anomaly detection result, that is, the objective function; represents the identified behavior features; represents the transpose of the weight vector; represents the bias term;
[0030] Obtain a preset probability threshold , and compare the preset probability threshold with the prediction probability corresponding to the anomaly detection result to generate the anomaly detection result, that is, the behavior pattern analysis result;
[0031] If , then the behavior pattern corresponding to the prediction probability is in an abnormal state; if , then the behavior pattern corresponding to the prediction probability is in a normal state.
[0032] Preferably, the generating the dynamic data governance strategy corresponding to the target enterprise based on the behavior pattern analysis result and in combination with the behavior pattern frequency specifically includes:
[0033] Retrieve a preset time period and a preset quantity, and based on the preset quantity, count the occurrence times of each behavior pattern within each preset time period, that is, the behavior pattern frequency;
[0034] Sort the behavior pattern frequencies from largest to smallest to determine the target behavior pattern corresponding to the largest behavior pattern frequency;
[0035] Predict the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm, and generate the dynamic data governance strategy corresponding to the target enterprise according to the behavior development trend and the behavior pattern analysis result.
[0036] Preferably, the predicting the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm and generating the dynamic data governance strategy corresponding to the target enterprise according to the behavior development trend and the behavior pattern analysis result specifically includes:
[0037] Predict the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm to generate a predicted behavior pattern;
[0038] Retrieve the corresponding actual behavior pattern, compare the predicted behavior pattern with the actual behavior pattern, and quantify the degree of difference between the predicted behavior pattern and the actual behavior pattern based on the variance measurement method to obtain a difference index;
[0039] Obtain a standard index, and use the target behavior pattern corresponding to the difference index exceeding the standard index as a high-impact behavior pattern;
[0040] Obtain the original data governance strategy corresponding to the target enterprise, and update the original data governance strategy based on the high-impact behavior pattern and the behavior pattern analysis result to generate the dynamic data governance strategy.
[0041] Preferably, after generating the dynamic data governance strategy, apply the dynamic data governance strategy to the target enterprise to monitor and evaluate the quality of the multi-source heterogeneous data of the target enterprise.
[0042] Compared with the related technologies, the enterprise management data governance method based on big data analysis provided by the present invention has the following beneficial effects:
[0043] The present invention obtains a target enterprise, collects the initial multi-source heterogeneous data of the target enterprise; receives the initial multi-source heterogeneous data, preprocesses the initial multi-source heterogeneous data to generate target multi-source heterogeneous data; obtains the behavior pattern of the target multi-source heterogeneous data, constructs a pattern recognition model, and performs corresponding processing on the behavior pattern based on the pattern recognition model to generate a behavior pattern analysis result; based on the behavior pattern analysis result and combined with the behavior pattern frequency, generates a dynamic data governance strategy corresponding to the target enterprise, so as to meet the needs of modern enterprises, deeply analyze and real-time monitor the dynamic behavior of data, comprehensively and efficiently analyze enterprise management data, implement the dynamic data governance strategy, and at the same time be able to ensure the security of enterprise management data, effectively respond to various complex data scenarios, and realize the automation and intelligence of enterprise data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the enterprise management data governance method based on big data analysis of the present invention;
[0045] Figure 2 It is a schematic diagram of the initial multi-source heterogeneous data acquisition process of the present invention;
[0046] Figure 3 It is a schematic diagram of the initial multi-source heterogeneous data preprocessing process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0048] Embodiment 1
[0049] As Figure 1 shown, the enterprise management data governance method based on big data analysis includes:
[0050] S1. Obtain a target enterprise and collect the initial multi-source heterogeneous data of the target enterprise;
[0051] Among them, in the process of enterprise management data governance and intelligent analysis, the target enterprise can be accurately located and obtained first. This process involves market research, competitive analysis, and screening of potential cooperation partners. The initial multi-source heterogeneous data of the target enterprise can be collected through data collection technology. These data include structured database records, unstructured documents, semi-structured data, and streaming data. Among them, unstructured documents can be emails, reports, etc., semi-structured data can be XML, JSON files, etc., and streaming data can be real-time transaction data, sensor data, etc. These data together constitute a comprehensive portrait of enterprise operations.
[0052] S2. Receive the initial multi-source heterogeneous data and preprocess the initial multi-source heterogeneous data to generate target multi-source heterogeneous data;
[0053] It can be understood that after receiving these initial multi-source heterogeneous data, data cleaning technology can be used to remove noise and redundancy in the data, different formats of data can be unified into a standardized format through data conversion, and data integration technology can be used to fuse data from multiple sources into a consistent and complete data set, that is, target multi-source heterogeneous data, so as to improve data quality and facilitate subsequent data analysis.
[0054] S3. Obtain the behavior patterns of the target multi-source heterogeneous data, construct a pattern recognition model, and perform corresponding processing on the behavior patterns based on the pattern recognition model to generate a behavior pattern analysis result;
[0055] It should be noted that the behavior patterns in the target multi-source heterogeneous data can be deeply analyzed. Through advanced data analysis technologies such as feature extraction, association rule mining, and time series analysis, key behavior sequences, abnormal behaviors, and potential trend changes in enterprise management data can be identified, and then a precise and efficient pattern recognition model can be constructed. This model can automatically identify and classify new behavior patterns, and monitor and predict the behaviors of the target enterprise in real time.
[0056] S4. Generate a dynamic data governance strategy corresponding to the target enterprise based on the behavior pattern analysis result and in combination with the behavior pattern frequency.
[0057] In practical applications, the results of these behavior pattern analyses can reflect the internal laws of enterprise management data, namely potential optimization spaces and risk points. Combining the frequencies of the appearance of behavior patterns and using dynamic data governance methods, targeted dynamic data governance strategies can be formulated. These strategies include data quality control measures, data security protection plans, data value mining plans, and business process optimization suggestions, thereby enabling the maximized utilization of enterprise data assets and the minimized control of risks.
[0058] Through the above methods, a complete and efficient enterprise data governance and analysis system can be constructed to promote the digital transformation and intelligent upgrading of enterprises.
[0059] In the specific implementation process, obtaining the target enterprise and collecting the initial multi-source heterogeneous data of the target enterprise specifically include:
[0060] Connecting to and accessing the internal database of the target enterprise to obtain enterprise internal data;
[0061] Obtaining enterprise external data from the external services of the target enterprise through data interfaces;
[0062] Retrieving public web pages and collecting information related to the target enterprise on these public web pages to generate enterprise web page information;
[0063] Unifying the data formats of the enterprise internal data, enterprise external data, and enterprise web page information to generate the initial multi-source heterogeneous data.
[0064] See Figure 2 , first, it is possible to connect to and securely access the internal database of the target enterprise. By adopting efficient data integration technologies, such as database middleware or API (Application Programming Interface) calls, without interfering with the normal operation of the enterprise, enterprise internal data can be stably and accurately obtained on the premise. These enterprise internal data include financial data, operation reports, human resources information, and supply chain management data, etc., which can directly reflect key indicators such as the internal operation status, cost structure, profitability, and employee efficiency of the enterprise.
[0065] Through preset data interfaces, enterprise external data can be obtained from the external service systems of the target enterprise. These external data may include market trends, customer feedback, competitor dynamics, and industry reports, etc. Through these data, the external environment of the enterprise can be understood, the market positioning can be evaluated, and competitive strategies can be formulated. By using data acquisition tools and API integration technologies, the efficient and automated collection of external data sources can be realized, thereby ensuring the timeliness and accuracy of the data.
[0066] In addition, public web pages can be retrieved, such as corporate official websites, social media platforms, and industry forums, etc., to collect information related to the target enterprise on these platforms. This process usually involves natural language processing (NLP) and web crawler technology. By automatically identifying, extracting, and organizing multi-dimensional web information such as enterprise news, product evaluations, user feedback, and brand exposure, a web information dataset of the enterprise can be generated.
[0067] Furthermore, the collected internal enterprise data, external enterprise data, and enterprise web information can be subjected to unified data formatting processing. This process may include steps such as data cleaning (removing duplicate, incorrect, and incomplete data), data conversion (converting data in different formats into a unified standard), and data integration (fusing multi-source data into a unified dataset), etc. Thus, initial multi-source heterogeneous data with clear structure and unified format can be generated, which is convenient for subsequent data analysis, mining, and modeling, and helps the enterprise to gain insights into market trends, optimize operation strategies, and enhance market competitiveness.
[0068] Receiving the initial multi-source heterogeneous data and preprocessing the initial multi-source heterogeneous data to generate target multi-source heterogeneous data specifically includes:
[0069] Detecting and marking abnormal data in the initial multi-source heterogeneous data, where the abnormal data includes missing data, duplicate data, and incorrect data;
[0070] Filling the missing data, and deleting the duplicate data and incorrect data to obtain first multi-source heterogeneous data;
[0071] Cleaning the unstructured text in the first multi-source heterogeneous data based on natural language processing technology to obtain second multi-source heterogeneous data;
[0072] Based on the preset standardization rules, uniformly classifying the second multi-source heterogeneous data to generate target multi-source heterogeneous data.
[0073] See Figure 3 , first, the initial multi-source heterogeneous data can be received and preprocessed to ensure the accuracy and effectiveness of the subsequent analysis process. This process specifically includes the processing of multi-dimensional and multi-type (i.e., heterogeneity) of data, so that the original data can be transformed into unified, standardized, and high-quality target multi-source heterogeneous data.
[0074] First, for the initial multi-source heterogeneous data, anomaly data detection and marking can be performed first. These anomaly data include missing values, duplicate records, and outliers that significantly deviate from the normal data distribution or business logic. Missing data refers to the parts in the dataset where necessary information is not provided, which may be caused by recording errors or omissions during the data collection process; duplicate data refers to exactly the same or highly similar records that appear repeatedly in the dataset, which may lead to data redundancy and analysis bias; incorrect data may be generated due to measurement errors or data entry errors. Through statistical analysis and machine learning algorithms, these anomaly data can be automatically identified and marked.
[0075] The missing data is filled by interpolation methods, which can be mean imputation, hot deck imputation, and model-based predictive imputation, so as to reduce the impact of data loss on the analysis results. While duplicate data and incorrect data can be directly deleted from the dataset to ensure the consistency of the dataset, generating the first multi-source heterogeneous data.
[0076] Using natural language processing techniques (NLP), such as word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis, etc., the noise in the text data, such as spelling mistakes, irrelevant symbols, etc., can be effectively cleaned. Furthermore, the key information in the data can be extracted, and the first multi-source heterogeneous data can be converted into a structured or semi-structured format, generating the second multi-source heterogeneous data.
[0077] Based on the preset standardization rules, the second multi-source heterogeneous data can be uniformly classified and formatted. These rules may involve the unification of data types, the conversion of measurement units, and the compliance with coding standards, etc., so as to eliminate the differences between data and achieve the seamless integration and sharing of data.
[0078] Through the above methods, the target multi-source heterogeneous data can be generated, retaining the multi-source and heterogeneity of the original data, and ensuring that the generated target multi-source heterogeneous data has high standardization, accuracy, and usability, which is convenient for subsequent data analysis, mining, and decision support.
[0079] The behavior pattern of obtaining the target multi-source heterogeneous data, constructing a pattern recognition model, and performing corresponding processing on the behavior pattern based on the pattern recognition model to generate a behavior pattern analysis result specifically includes:
[0080] Obtain the behavior pattern of the target multi-source heterogeneous data, and determine the corresponding behavior vector according to the behavior pattern;
[0081] Construct the pattern recognition model based on machine learning algorithms;
[0082] Perform recognition processing on the behavior vector according to the pattern recognition model to generate recognition behavior features;
[0083] Perform anomaly detection on the identified behavior features to generate the behavior pattern analysis result.
[0084] It can be understood that first, key information closely related to the target behavior pattern can be extracted from the raw data. This process includes data cleaning (removing noise and handling missing values), data integration (merging data from different sources), and data transformation (such as normalization and discretization) to ensure data consistency and comparability. Based on this preprocessed data, using feature engineering techniques, the behavior pattern can be transformed into quantifiable behavior vectors. These behavior vectors can reflect the essential characteristics of different behavior patterns and ensure the accuracy of subsequent pattern recognition.
[0085] A pattern recognition model can be constructed through machine learning algorithms. This process involves algorithm selection, parameter tuning, and model training. Through learning a large amount of sample data, this model can learn the mapping relationship between behavior vectors and specific behavior patterns, and thus can accurately classify or predict unknown behavior patterns.
[0086] In the pattern recognition stage, the behavior vectors to be analyzed can be input into the trained pattern recognition model. This model can perform recognition processing on these vectors to generate corresponding identified behavior features. These features can reflect the essential attributes of the behavior pattern and facilitate subsequent anomaly detection.
[0087] To identify potential risks or abnormal behaviors, an anomaly detection mechanism can be introduced. This mechanism, based on methods such as statistics, clustering analysis, and outlier detection, can deeply analyze the identified behavior features to determine whether they deviate from the range of normal behavior patterns. Once an anomaly is detected, the behavior pattern analysis result is generated. These results not only include specific abnormal behaviors but also include the risk level assessment of abnormal behaviors and recommended countermeasures, etc.
[0088] The performing anomaly detection on the identified behavior features to generate the behavior pattern analysis result specifically includes:
[0089] Set the anomaly detection result as the objective function, and the prediction probability corresponding to the anomaly detection result is:
[0090]
[0091] In the formula, represents the prediction probability corresponding to the anomaly detection result; represents the anomaly detection result, that is, the objective function; represents the identified behavior features; represents the transpose of the weight vector; represents the bias term;
[0092] Obtain a preset probability threshold and compare the preset probability threshold with the predicted probability corresponding to the anomaly detection result to generate the anomaly detection result, that is, the behavior pattern analysis result;
[0093] If , the behavior pattern corresponding to the predicted probability is an abnormal state; if , the behavior pattern corresponding to the predicted probability is a normal state.
[0094] Based on the behavior pattern analysis result and combined with the behavior pattern frequency, generate the dynamic data governance strategy corresponding to the target enterprise, specifically including:
[0095] Retrieve a preset time period and a preset quantity, and based on the preset quantity, count the occurrence times of the behavior pattern within each preset time period, that is, the behavior pattern frequency;
[0096] Sort the behavior pattern frequencies from largest to smallest to determine the target behavior pattern corresponding to the largest behavior pattern frequency;
[0097] Predict the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm, and generate the dynamic data governance strategy corresponding to the target enterprise according to the behavior development trend and the behavior pattern analysis result.
[0098] It can be understood that by setting a clear preset time period (such as monthly, quarterly, and annual cycles, etc.) and a preset quantity, a representative and analyzable data sample set can be screened out. Using an efficient data processing engine, the occurrence frequency of each behavior pattern within each preset time period, that is, the behavior pattern frequency, can be accurately counted, thereby ensuring the timeliness and accuracy of the data.
[0099] After obtaining the behavior pattern frequency data, a sorting algorithm can be used to sort the behavior patterns in descending order of frequency, and then the target behavior pattern that has the most significant impact on enterprise operations can be quickly identified, and the high-frequency and high-impact behavior patterns can be screened out, which is convenient for formulating targeted data governance strategies in the future.
[0100] By introducing the machine learning algorithm of decision tree analysis, the behavior development trend of the target behavior pattern can be deeply predicted. By analyzing the feature associations and pattern evolutions in historical data, the decision tree analysis algorithm can construct a corresponding prediction model to accurately predict the change trend of the target behavior pattern in the future for a period of time, including possible directions, speeds, and influence ranges, etc.
[0101] Based on the prediction results of behavioral development trends, combined with the initial behavioral pattern analysis results, and taking into account the business needs, data security specifications, and compliance requirements of the enterprise, a dynamic data governance strategy for the target enterprise can be constructed. This strategy not only includes immediate response measures for high-frequency behavioral patterns, but also incorporates preventive data management mechanisms. Through dynamic adjustment and optimization, it can ensure the security, integrity, and efficiency of enterprise data during circulation. In addition, the strategy also emphasizes the continuous monitoring and evaluation mechanism of data governance to ensure that it is adjusted in a timely manner as the internal and external environment of the enterprise changes, so as to achieve intelligent and refined data governance.
[0102] The decision tree analysis algorithm is used to predict the behavior development trend of the target behavior pattern, and the dynamic data governance strategy corresponding to the target enterprise is generated according to the behavior development trend and the behavior pattern analysis result, specifically including:
[0103] Predicting the behavior development trend of the target behavior pattern based on a decision tree analysis algorithm to generate a predicted behavior pattern;
[0104] Retrieving the corresponding actual behavior pattern, comparing the predicted behavior pattern with the actual behavior pattern, and quantifying the degree of difference between the predicted behavior pattern and the actual behavior pattern based on a variance measurement method to obtain a difference index;
[0105] Obtaining a standard indicator, and taking a target behavior pattern corresponding to a difference indicator exceeding the standard indicator as a high-impact behavior pattern;
[0106] The original data governance policy corresponding to the target enterprise is obtained, and the original data governance policy is updated based on the high-impact behavior pattern and the behavior pattern analysis result to generate the dynamic data governance policy.
[0107] In practical applications, when building an efficient and dynamic data governance strategy for the target enterprise, the decision tree analysis algorithm can accurately predict the behavioral development trend of the target behavior pattern. With its powerful classification and regression capabilities, the decision tree algorithm can conduct in-depth learning of historical behavior data and identify key feature variables and their relationship with the target behavior pattern. Through iterative splitting and pruning strategies, the decision tree model can build a prediction model and then generate predictive behavior patterns that can reflect the possible evolution path of the target behavior in the future time window.
[0108] In order to verify the accuracy and applicability of the predicted behavior pattern, the actual behavior pattern data of the same period can be retrieved and compared with the predicted behavior pattern. The variance between the two sets of data can be calculated by the variance measurement method, and the degree of difference between them can be quantified to generate the corresponding difference index. The difference index can reflect the accuracy of the prediction model and quantify the effectiveness of the evaluation of the predicted behavior pattern.
[0109] Among them, these standard metrics are formulated based on industry best practices, enterprise-specific requirements, and the basic principles of data governance. By comparing the differential metrics with the standard metrics, it is possible to quickly identify the target behavior patterns with significant differences, namely, the high-impact behavior patterns.
[0110] After obtaining the original data governance strategy corresponding to the target enterprise, it is possible to conduct targeted updates and optimizations on the original data governance strategy based on the identification of high-impact behavior patterns and the results of behavior pattern analysis. This process includes strengthening the supervision measures for specific behavior patterns and flexibly adjusting the overall data governance framework, so as to ensure that this strategy can adapt to business development and environmental changes, generate an efficient, flexible, and adaptable dynamic data governance strategy, significantly improve the enterprise's data management ability, and promote the sustainable development of the enterprise.
[0111] After generating the dynamic data governance strategy, apply the dynamic data governance strategy to the target enterprise, and monitor and evaluate the quality of the multi-source heterogeneous data of the target enterprise.
[0112] In practical applications, after obtaining the dynamic data governance strategy of the target enterprise, it can be seamlessly integrated into the data management system of the target enterprise. This process not only involves the specific deployment and execution of the strategy, but also covers the in-depth integration of the strategy with the enterprise's existing IT architecture and data processes. Through data quality monitoring tools, it is possible to comprehensively monitor the multi-source heterogeneous data of the target enterprise, and these multi-source heterogeneous data include structured, semi-structured, and unstructured data. In addition, this monitoring process needs to follow the preset data quality metrics and standards, such as metrics like accuracy, integrity, consistency, and timeliness. At the same time, a data quality evaluation mechanism can be established to conduct in-depth analysis of the monitoring results regularly, so as to evaluate the implementation effect of the dynamic data governance strategy, and optimize and adjust this strategy based on the evaluation feedback results, forming a continuously improving data governance closed-loop.
[0113] Through the introduction of the above embodiments, the present invention provides an enterprise management data governance method based on big data analysis. By obtaining a target enterprise, initial multi-source heterogeneous data of the target enterprise is collected; the initial multi-source heterogeneous data is received and preprocessed to generate target multi-source heterogeneous data; the behavior patterns of the target multi-source heterogeneous data are obtained, a pattern recognition model is constructed, and the behavior patterns are processed accordingly based on the pattern recognition model to generate a behavior pattern analysis result; based on the behavior pattern analysis result and combined with the behavior pattern frequency, a dynamic data governance strategy corresponding to the target enterprise is generated, thereby meeting the needs of modern enterprises, deeply analyzing and real-time monitoring the dynamic behavior of data, comprehensively and efficiently analyzing enterprise management data, implementing the dynamic data governance strategy, while ensuring the security of enterprise management data, effectively coping with various complex data scenarios, and realizing the automation and intelligence of enterprise data governance.
[0114] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0115] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0116] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
Claims
1. An enterprise management data governance method based on big data analysis, characterized in that, Including: Obtain a target enterprise and collect the initial multi-source heterogeneous data of the target enterprise; Receive the initial multi-source heterogeneous data and preprocess the initial multi-source heterogeneous data to generate target multi-source heterogeneous data; Obtain the behavior patterns of the target multi-source heterogeneous data, construct a pattern recognition model, and perform corresponding processing on the behavior patterns based on the pattern recognition model to generate a behavior pattern analysis result; Generate a dynamic data governance strategy corresponding to the target enterprise based on the behavior pattern analysis result and in combination with the behavior pattern frequency; The generating the dynamic data governance strategy corresponding to the target enterprise based on the behavior pattern analysis result and in combination with the behavior pattern frequency specifically includes: Retrieve a preset time period and a preset quantity, and based on the preset quantity, count the occurrence times of the behavior patterns within each preset time period, that is, the behavior pattern frequency; Sort the behavior pattern frequencies from largest to smallest, and determine the target behavior pattern corresponding to the largest behavior pattern frequency; Predict the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm, and generate a dynamic data governance strategy corresponding to the target enterprise according to the behavior development trend and the behavior pattern analysis result; The predicting the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm and generating a dynamic data governance strategy corresponding to the target enterprise according to the behavior development trend and the behavior pattern analysis result specifically includes: Predict the behavior development trend of the target behavior pattern based on the decision tree analysis algorithm to generate a predicted behavior pattern; Retrieve the corresponding actual behavior pattern, compare the predicted behavior pattern with the actual behavior pattern, and quantify the difference degree between the predicted behavior pattern and the actual behavior pattern based on the variance measurement method to obtain a difference index; Obtain a standard index, and use the target behavior pattern corresponding to the difference index exceeding the standard index as a high-impact behavior pattern; Obtain the original data governance strategy corresponding to the target enterprise, and update the original data governance strategy based on the high-impact behavior pattern and the behavior pattern analysis result to generate the dynamic data governance strategy; After generating the dynamic data governance strategy, apply the dynamic data governance strategy to the target enterprise, and monitor and evaluate the quality of the multi-source heterogeneous data of the target enterprise.
2. The enterprise management data governance method based on big data analysis according to claim 1, characterized in that, The obtaining the target enterprise and collecting the initial multi-source heterogeneous data of the target enterprise specifically includes: Connect to and access the internal database of the target enterprise to obtain enterprise internal data; Obtain enterprise external data from the external services of the target enterprise through a data interface; Retrieve public web pages, collect information related to the target enterprise on the public web pages, and generate enterprise web page information; Unify the data formats of the enterprise internal data, enterprise external data, and enterprise web page information to generate the initial multi-source heterogeneous data.
3. The enterprise management data governance method based on big data analysis according to claim 1, wherein The receiving the initial multi-source heterogeneous data and preprocessing the initial multi-source heterogeneous data to generate target multi-source heterogeneous data specifically includes: Detect and mark the abnormal data in the initial multi-source heterogeneous data, where the abnormal data includes missing data, duplicate data, and error data; Fill in the missing data, and delete the duplicate data and error data to obtain the first multi-source heterogeneous data; Clean the unstructured text in the first multi-source heterogeneous data based on natural language processing technology to obtain the second multi-source heterogeneous data; Classify the second multi-source heterogeneous data uniformly based on preset standardization rules to generate the target multi-source heterogeneous data.
4. The enterprise management data governance method based on big data analysis according to claim 1, characterized in that The act of obtaining the behavior pattern of the target multi-source heterogeneous data, constructing a pattern recognition model, and performing corresponding processing on the behavior pattern based on the pattern recognition model to generate a behavior pattern analysis result specifically includes: Obtain the behavior pattern of the target multi-source heterogeneous data, and determine the corresponding behavior vector according to the behavior pattern; Construct the pattern recognition model based on machine learning algorithms; Perform recognition processing on the behavior vector according to the pattern recognition model to generate recognition behavior features; Perform anomaly detection on the recognition behavior features to generate the behavior pattern analysis result.
5. The enterprise management data governance method based on big data analysis according to claim 4, wherein The act of performing anomaly detection on the recognition behavior features to generate the behavior pattern analysis result specifically includes: Set the anomaly detection result as the objective function, and the prediction probability corresponding to the anomaly detection result is: , In the formula, represents the predicted probability corresponding to the anomaly detection result; represents the anomaly detection result, that is, the objective function; represents the recognized behavioral characteristics; represents the transpose of the weight vector; represents the bias term; Obtain a preset probability threshold , and compare the preset probability threshold with the predicted probability corresponding to the anomaly detection result to generate the anomaly detection result, that is, the behavior pattern analysis result; If , then the behavior pattern corresponding to the predicted probability is an abnormal state; if , then the behavior pattern corresponding to the predicted probability is a normal state.
Citation Information
Patent Citations
Intelligent doorbell user behavior analysis system
CN118430111A
Enterprise gene generation method and system based on enterprise multi-dimensional data
CN118503491A