A full-link data management system for improving upper-layer application data service capability
Patent Information
- Application Number
- CN202610688626.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-25
AI Technical Summary
然而,这些方案普遍将数据采集、处理、分析、应用等环节作为独立模块分别实现,各环节之间缺乏统一的协同机制与标准化的数据流转通道,导致数据治理流程割裂,难以形成从数据接入到业务输出的端到端闭环
[0006]本发明的有益效果是:本发明通过构建数据采集层、数据处理层、数据算法层和数据应用场景层协同联动的全链路架构,形成从数据接入到业务输出的完整治理闭环。数据采集层实现多源异构数据的统一接入与校验,数据处理层对原始数据进行场景化组织与标准化转换以生成高质量数据集,数据算法层利用算法模型对数据集进行处理生成面向业务场景的分析结果,数据应用场景层将分析结果输出。通过各层之间的紧密协同与标准化数据流转,有效解决了现有数据治理方案中各环节割裂、协同不足的问题,实现了数据从采集到应用的全流程贯通,显著提升了上层应用的数据服务能力。
Smart Images

Figure CN122817673A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a full-link data governance system that enhances the data service capabilities of upper-layer applications. Background Technology
[0002] With the deepening of digital transformation, data has become a core production factor for enterprises and institutions. To unlock the value of data, various data governance solutions have emerged in the industry, covering core aspects such as data integration, quality control, and lineage analysis. However, these solutions generally implement data collection, processing, analysis, and application as independent modules, lacking a unified collaborative mechanism and standardized data flow channels between the various stages. This results in a fragmented data governance process, making it difficult to form an end-to-end closed loop from data access to business output.
[0003] This fragmented architecture makes it difficult for data governance systems to provide high-quality, traceable, and easily integrated data services to upper-layer applications, severely restricting the full realization of the value of data elements. Therefore, there is an urgent need for a governance system that enables deep collaboration and end-to-end connectivity across all stages, in order to achieve efficient empowerment of upper-layer applications by data governance capabilities. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a full-link data governance system that enhances the data service capabilities of upper-layer applications, so as to solve the above-mentioned technical problem.
[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A full-link data governance system for improving the data service capabilities of upper-layer applications, comprising: a data acquisition layer, a data processing layer, a data algorithm layer, and a data application scenario layer; the data acquisition layer is used to collect multi-source heterogeneous raw data from multiple data sources, and to perform verification processing on the collected multi-source heterogeneous raw data, and to transmit the verified multi-source heterogeneous raw data to the data processing layer; the data processing layer is used to store the verified multi-source heterogeneous raw data, and to perform data arrangement on the stored multi-source heterogeneous raw data to generate a dataset for use by the data algorithm layer; the data algorithm layer is used to obtain the dataset from the data processing layer, and to process the dataset using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios; the data application scenario layer is used to support at least one upper-layer application oriented towards business scenarios, and to provide the analysis results generated by the data algorithm layer to the upper-layer application for use.
[0006] The beneficial effects of this invention are as follows: By constructing a collaborative end-to-end architecture comprising a data acquisition layer, a data processing layer, a data algorithm layer, and a data application scenario layer, this invention forms a complete governance closed loop from data access to business output. The data acquisition layer enables unified access and verification of multi-source heterogeneous data; the data processing layer organizes and standardizes raw data according to specific scenarios to generate high-quality datasets; the data algorithm layer processes the datasets using algorithmic models to generate analysis results tailored to business scenarios; and the data application scenario layer outputs the analysis results. Through close collaboration and standardized data flow between layers, this invention effectively solves the problems of fragmented processes and insufficient coordination in existing data governance solutions, achieving seamless data flow from acquisition to application and significantly improving the data service capabilities of upper-layer applications.
[0007] Based on the above technical solution, the present invention can be further improved as follows.
[0008] Furthermore, the data processing layer includes a data center and an application database; the data center is used to store verified multi-source heterogeneous raw data, and to perform data cleaning, format unification, and classification according to business scenarios on the stored multi-source heterogeneous raw data to generate the dataset; the data center is also used to synchronize the dataset to the application database, and the application database is used to store the dataset.
[0009] The beneficial effects of adopting the above-mentioned further solution are: dividing the data processing layer into a data center and an application database, with the data center focusing on cleaning, format unification and scenario-based classification of raw data, and the application database specifically storing standardized datasets after governance, thereby decoupling data governance and data storage, facilitating independent expansion and maintenance of each module, while ensuring that the upper layer can quickly obtain high-quality, scenario-based data.
[0010] Furthermore, the data center includes a network database and a data mart; the network database is used to receive and store verified multi-source heterogeneous raw data, and synchronize the verified multi-source heterogeneous raw data to the data mart; the data mart includes multiple data sub-marts divided according to business scenarios, and each data sub-mart is used to store data that has been cleaned, formatted, and classified according to business scenarios.
[0011] The beneficial effects of adopting the above-mentioned further solution are: the data center is further subdivided into network databases and data marts. The network database retains the original data to ensure traceability, while the data marts organize the data according to business scenarios, forming a two-layer structure that separates the original data from the business data.
[0012] Furthermore, the network database employs a hybrid storage engine, which includes: pgSQL is used to store structured data from multi-source heterogeneous raw data that has passed validation. MinIO is used to store unstructured data in multi-source heterogeneous raw data that has passed verification; HBase is used to store time-series data from multi-source heterogeneous raw data that has passed verification.
[0013] The beneficial effects of adopting the above-mentioned further solutions are: pgSQL, MinIO, and HBase are selected as storage components for structured, unstructured, and time-series data respectively, giving full play to the unique advantages of each database.
[0014] Furthermore, the data acquisition layer includes lightweight acquisition agents deployed on each data source side and a unified acquisition control module; each lightweight acquisition agent is used to acquire multi-source heterogeneous raw data from each data source side according to the configured acquisition frequency; the unified acquisition control module is used to perform verification processing on the acquired multi-source heterogeneous raw data and temporarily store the data that fails verification in an exception queue.
[0015] The beneficial effects of adopting the above-mentioned further solutions are: by using lightweight acquisition agents and unified acquisition control to achieve data acquisition and verification, the reliability of the acquired data is guaranteed and the overall data quality of the system is improved.
[0016] Furthermore, multiple data source sides include the wireless network side, the cloud network side, and the computing network side.
[0017] The beneficial effects of adopting the above-mentioned further solutions are: the system supports unified access to three typical heterogeneous data sources, namely wireless network, cloud network and computing network, covering traditional network infrastructure, cloud environment and computing network, breaking down data silos between different environments.
[0018] Furthermore, the data algorithm layer includes: The feature engineering module is used to perform feature engineering processing on the dataset to obtain feature data; The model processing module is used to process the feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios.
[0019] The beneficial effects of adopting the above-mentioned further scheme are: dividing the data algorithm layer into a feature engineering module and a model processing module, the feature engineering module is responsible for data preprocessing and feature selection, and the model processing module focuses on algorithm reasoning and result generation, thereby improving the efficiency of algorithm development and the reusability of the model.
[0020] Furthermore, the feature engineering module is used to perform feature engineering processing on the dataset to obtain feature data. Specifically, it is used to: perform feature extraction processing on the dataset to obtain initial feature data; perform feature removal processing on the initial feature data based on the variance selection method to obtain preprocessed feature data; and perform feature filtering processing on the preprocessed feature data based on the random forest algorithm to generate the feature data.
[0021] The beneficial effects of adopting the above-mentioned further scheme are: by using the variance selection method combined with the random forest algorithm for feature selection, the feature dimensionality is effectively reduced, the risk of model overfitting is reduced, and the algorithm layer is more efficient in processing high-dimensional data.
[0022] Furthermore, when the data application scenario layer is used to provide the analysis results generated by the data algorithm layer to the upper-layer application, it is specifically used to: provide the analysis results generated by the data algorithm layer to the upper-layer application in the form of a visual large screen display or an API interface.
[0023] The beneficial effects of adopting the above-mentioned further solutions are: providing a visual large screen display for human decision-making, while providing a standardized API interface to facilitate integration with third-party systems, and realizing the modular and standardized output of data governance capabilities.
[0024] Furthermore, when the model processing module processes the feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios, it specifically performs the following: based on the feature data, it performs multi-dimensional aggregation statistics on the dataset, assigning a first weight to recent data and a second weight to distant data in the dataset during the statistical process to generate data statistical results; wherein, the recent data is data obtained after a preset time, and the distant data is data obtained before the preset time, and the first weight is greater than the second weight; based on the feature data, it performs time-series prediction using a pre-trained prediction model to generate predicted values of the target indicator in the future time period; and obtains the analysis results based on the data statistical results and the predicted values of the target indicator in the future time period.
[0025] The beneficial effects of adopting the above-mentioned further scheme are: by using time decay weighting, the statistical results are closer to the recent system state, and by combining it with the time series prediction model, the post-event statistics and pre-event prediction are complemented, which not only provides a snapshot of the current data distribution, but also gives a prediction of future trends, providing more comprehensive and timely data support for business decisions. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of a full-link data governance system for improving the data service capabilities of upper-layer applications according to the present invention. Detailed Implementation
[0027] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0028] According to a report by the China Academy of Information and Communications Technology (CAICT), over 80% of enterprises have recognized the importance of data governance, but only a few have achieved systematic implementation. Most still face problems such as numerous data silos, inconsistent data quality, and significant security and compliance pressures. Specifically, the lack of unified access standards for multi-source, heterogeneous data leads to inefficient cross-departmental data sharing; inconsistent data semantic definitions cause deviations in business analysis and intelligent decision-making; the lack of dynamic monitoring during data flow makes it difficult to trace quality issues in a timely manner; and the difficulty in balancing data security and open sharing hinders the full release of data value.
[0029] The purpose of this invention is to overcome the shortcomings of existing data governance solutions, such as fragmented processes, weak scenario adaptability, insufficient end-to-end collaboration, and low efficiency in enabling upper-layer applications. It achieves unified access and standardized cleaning of multi-source heterogeneous data, breaks down data silos, solves the problem of weak compatibility of existing solutions with non-cloud ecosystems and edge node data, and improves the comprehensiveness and efficiency of data collection.
[0030] like Figure 1 As shown, this embodiment provides a full-link data governance system to enhance the data service capabilities of upper-layer applications, including: a data acquisition layer, a data processing layer, a data algorithm layer, and a data application scenario layer. The data acquisition layer is used to collect multi-source heterogeneous raw data from multiple data sources, and to perform verification processing on the collected multi-source heterogeneous raw data, transmitting the verified multi-source heterogeneous raw data to the data processing layer. The data processing layer is used to store the verified multi-source heterogeneous raw data and to perform data orchestration on the stored multi-source heterogeneous raw data to generate a dataset for the data algorithm layer to call. The data algorithm layer is used to obtain the dataset from the data processing layer and to process the dataset using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios. The data application scenario layer is used to support at least one upper-layer application oriented towards business scenarios and to provide the analysis results generated by the data algorithm layer to the upper-layer application for use.
[0031] This invention constructs a collaborative end-to-end architecture comprising a data acquisition layer, a data processing layer, a data algorithm layer, and a data application scenario layer, forming a complete governance closed loop from data access to business output. The data acquisition layer enables unified access and verification of multi-source heterogeneous data; the data processing layer organizes and standardizes raw data according to specific scenarios to generate high-quality datasets; the data algorithm layer processes the datasets using algorithmic models to generate analysis results tailored to business scenarios; and the data application scenario layer outputs the analysis results. Through close collaboration and standardized data flow between these layers, the invention effectively solves the problems of fragmented processes and insufficient coordination in existing data governance solutions, achieving seamless data flow from acquisition to application and significantly improving the data service capabilities of upper-layer applications.
[0032] Specifically, the end-to-end data governance system for enhancing upper-layer application data service capabilities provided by this invention adopts a layered architecture design, comprising, from bottom to top, a data acquisition layer, a data processing layer, a data algorithm layer, and a data application scenario layer. Each layer connects with the upper layer through a standardized data synchronization and orchestration mechanism, ultimately transforming the underlying raw data into capabilities usable by the upper-layer business.
[0033] Optionally, in this embodiment, the data acquisition layer includes lightweight acquisition agents deployed on each data source side and a unified acquisition control module; each lightweight acquisition agent is used to acquire multi-source heterogeneous raw data from each data source side according to the configured acquisition frequency; the unified acquisition control module is used to perform verification processing on the acquired multi-source heterogeneous raw data and temporarily store the data that fails verification in an exception queue.
[0034] The data acquisition layer is responsible for collecting heterogeneous raw data from multiple data sources. In this embodiment, the data sources include three types: wireless network side, cloud network side, and computing network side.
[0035] For the wireless network side, lightweight data acquisition agents based on SNMP or IPMI protocols are deployed on devices such as switches and servers to capture data such as device status and port traffic in real time. For the cloud network side, data acquisition agents are deployed to call cloud platform monitoring APIs and periodically pull data such as CPU utilization, memory usage, and instance alarms of cloud resources (such as ECS and RDS). For the computing network side, traffic and packet loss rate of routing devices are collected through the NetFlow protocol, and business alarm events are obtained by connecting to business system interfaces.
[0036] The unified data acquisition and control module establishes a data acquisition scheduling center, supporting configuration of acquisition frequency based on data source type. For example, wireless network device status can be updated every minute, and cloud network instance alarms can be pushed in real time. Simultaneously, the unified data acquisition and control module performs preliminary verification on the acquired data, including numerical range verification and field integrity verification. Data that fails verification is temporarily stored in an exception queue, supporting manual retry or automatic discarding.
[0037] It should be noted that the multi-source heterogeneous raw data listed above (such as device status, port traffic, CPU utilization, memory usage, instance alarms, traffic, packet loss rate, business alarm events, etc.) are only examples. In actual applications, the data acquisition layer can collect other relevant data according to the business needs of the upper layer, and is not limited to the specific types mentioned above.
[0038] With the above configuration, the data acquisition layer can be compatible with various deployment environments, enabling unified and efficient access to multi-source heterogeneous data.
[0039] Optionally, in this embodiment, the data processing layer includes a data center and an application database; the data center is used to store the verified multi-source heterogeneous raw data, and to perform data cleaning, format unification, and classification according to business scenarios on the stored multi-source heterogeneous raw data to generate a dataset; the data center is also used to synchronize the dataset to the application database, and the application database is used to store the dataset.
[0040] Optionally, in an embodiment, the data center includes a network database and a data mart; the network database is used to receive and store verified multi-source heterogeneous raw data, and synchronize the verified multi-source heterogeneous raw data to the data mart; the data mart includes multiple data sub-marts divided according to business scenarios, and each data sub-mart is used to store data that has been cleaned, formatted, and classified according to business scenarios.
[0041] Optionally, in this embodiment, the network database employs a hybrid storage engine, which includes: pgSQL is used to store structured data from multi-source heterogeneous raw data that has passed validation. MinIO is used to store unstructured data in multi-source heterogeneous raw data that has passed verification; HBase is used to store time-series data from multi-source heterogeneous raw data that has passed verification.
[0042] The network database uses a hybrid storage engine, specifically including: pgSQL is used to store structured data, such as cloud network CPU utilization, and is partitioned into tables by "data source + year and month"; MinIO is used to store unstructured data, such as device logs, with the file naming convention being "collection source ID_timestamp"; HBase is used to store time-series data, such as network traffic, and is indexed by timestamp column families to support high-concurrency writes.
[0043] The network database receives and stores the verified multi-source heterogeneous raw data transmitted from the data acquisition layer, and synchronizes this data to the data mart.
[0044] Data marts are structured into multiple sub-data marts based on business scenarios, such as wireless network marts, cloud network marts, and computing network marts, with each sub-mart corresponding to an independent storage table. Data centers use visual orchestration tools (such as Knime) to perform data cleaning, format standardization, and classification according to business scenarios.
[0045] Specifically, data cleaning includes removing duplicate data and filling in missing fields (e.g., filling in the collection time if the fault time is empty). Format standardization includes unifying different expressions of the same business field into a preset standard format (e.g., standardizing fault levels to C1+ / C1 / C2). Data is then categorized by business scenario into corresponding data submarts.
[0046] The network database stores the raw data aggregated from the data acquisition layer, serving as the system's raw data storage center. The application database stores processed business data (such as fault work orders and differentiated business indicators), acting as the usable data carrier for upper-layer applications.
[0047] The application database, built on PostgreSQL, stores standardized datasets processed by the data center. Incremental synchronization from the data center to the application database is achieved using ETL tools (such as DataX), with synchronization rules configurable, for example, "incrementally synchronize data from the last hour every hour." The application database also provides a data access interface for the data algorithm layer to use.
[0048] The data center also provides API interfaces for the data algorithm layer to filter by dimension (such as querying wireless network fault data that distinguishes A in the past 7 days), and supports pushing data according to algorithm requirements (such as the fault analysis algorithm in the data algorithm layer subscribing to newly generated fault data, which is pushed by the data center in real time).
[0049] Optionally, in this embodiment, the data algorithm layer includes: The feature engineering module is used to perform feature engineering on the dataset to obtain feature data. The model processing module is used to process feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios.
[0050] Optionally, in the embodiments, the feature engineering module is used to perform feature engineering processing on the dataset to obtain feature data. Specifically, it is used to: perform feature extraction processing on the dataset to obtain initial feature data; perform feature removal processing on the initial feature data based on the variance selection method to obtain preprocessed feature data; and perform feature filtering processing on the preprocessed feature data based on the random forest algorithm to generate feature data.
[0051] The data algorithm layer is used to obtain standardized datasets from the data processing layer and generate business-oriented analysis results through feature engineering and algorithm models.
[0052] Taking fault analysis using a dataset as an example, based on the structured characteristics of fault data, three main categories of features are extracted: categorical features, temporal features, and statistical features, thus transforming the raw data into features usable for modeling. The feature engineering module performs feature engineering processing on the dataset, specifically including: For discrete category fields, encoding transformation is performed to meet the algorithm's input requirements. Specifically, for category fields with hierarchical or ordered relationships, such as fault level (C1+ / C1 / C2) and fault classification (e.g., network fluctuation / equipment failure / configuration anomaly), label encoding is used (e.g., urgent=4, high=3, medium=2, low=1). For unordered category fields such as the distinction (e.g., distinction A / distinction B / distinction C) and the business involved (e.g., business X / business Y / business Z), one-hot encoding is used to generate a sparse feature matrix (e.g., distinction A corresponds to [1,0,0], distinction B corresponds to [0,1,0]), avoiding the algorithm's misjudgment of ordered relationships between categories. Finally, the encoded category features are integrated into a category feature vector with unified dimensions.
[0053] Based on the failure occurrence time, multi-dimensional time features are extracted to support time occurrence rate analysis. Specifically, features are broken down into hourly segments (e.g., 0-6 AM / 7-12 PM / 13-18 PM / 19-24 PM), date types (weekdays / weekends / holidays), months (January-December), and quarters (Q1-Q4). Continuous time features such as the time difference between the failure occurrence time and midnight of the current day, and the interval between the failure occurrence time and the previous failure occurrence time are calculated to reflect the temporal patterns of failure occurrence. Then, by grouping by "region + date," statistical time features such as the failure frequency and average failure interval for each region and day are calculated.
[0054] Statistical feature construction: Derived statistical features are calculated based on single-dimensional or cross-dimensional grouping, such as calculating the average handling time and fault recurrence rate of historical faults by grouping by fault type, and calculating the proportion of fault occurrence and concentration of high-incidence periods by cross-grouping by distinction and fault type. Continuous statistical features (such as average handling time and fault interval) are normalized by Z-Score to eliminate differences in dimensions.
[0055] Finally, the variance selection method is used to remove low-discrimination features with variance less than 0.01. Then, the feature importance score is calculated based on the random forest algorithm, and the top 80% of the core features are retained to form the final feature data.
[0056] Optionally, in an embodiment, when the model processing module processes feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios, it specifically performs the following: based on the feature data, it performs multi-dimensional aggregation statistics on the dataset, and assigns a first weight to recent data in the dataset and a second weight to distant data in the dataset during the statistical process to generate data statistical results; wherein, recent data refers to data obtained after a preset time, distant data refers to data obtained before a preset time, and the first weight is greater than the second weight; based on the feature data, it performs time-series prediction using a pre-trained prediction model to generate predicted values of the target indicator in the future time period; and obtains analysis results based on the data statistical results and the predicted values of the target indicator in the future time period.
[0057] The data algorithm layer can deploy various algorithms and models according to usage requirements, such as fault analysis algorithms, delivery process control algorithms, and PCDN user identification models. Utilizing the scenario-based data provided by the data center, algorithms / models are used for analysis and computation to transform raw data into business-usable capabilities (such as fault analysis algorithms outputting root causes of faults and PCDN models identifying user characteristics).
[0058] Taking fault analysis algorithms as an example, they are mainly implemented through three core models: weighted multidimensional aggregation statistical model, time series prediction model, and association rule mining model.
[0059] The weighted multidimensional aggregation statistical model is used to perform multidimensional statistics on historical data, including the proportion analysis of the category dimension and the occurrence analysis of the time dimension, and the statistical results are made more consistent with the recent system state by using time decay weighting.
[0060] The model specifically adopts a multidimensional aggregation statistical model based on Spark SQL, which uses a distributed computing engine to quickly group and aggregate structured data, supporting real-time statistics of massive amounts of data.
[0061] Using "fault type and location" as grouping keys, an aggregation function is fitted on the training set to calculate the proportion of each type of fault and the proportion of fault distribution across different dimensions. Using "hour period and date type" as grouping keys, a time-series aggregation rule is fitted to statistically analyze the frequency, incidence, and trends of faults in each time period. Weighting coefficients are introduced; specifically, recent fault data (last 30 days) is assigned higher weights (e.g., weight coefficient 1.5), while older data (over 30 days) is assigned lower weights (e.g., weight coefficient 0.8), making the model more closely reflect recent fault patterns. This training process is completed dynamically each time statistics are executed, requiring no offline training. The final data statistics results include rankings by category dimension, incidence distribution across time dimensions, and comparisons between recent and older data.
[0062] Predictive models are used to predict target indicators (such as event frequency, resource utilization, etc.) for future periods based on historical time series data, enabling early warning. The predictive models employ either the Autoregressive Integrated Moving Average (ARIMA) model or the Long Short-Term Memory (LSTM) network. The ARIMA model is suitable for linear stationary sequences, while LSTM is suitable for nonlinear, long-term dependent sequences.
[0063] In this embodiment, the training process of the prediction model includes: Using the number of faults occurring at the hourly level as a time series, the stationarity of the training set data is tested (ADF test). If the data is non-stationary, it is made to meet the stationarity requirement through differencing.
[0064] For the ARIMA model, the autoregressive order, differencing order, and moving average order hyperparameters are optimized using the AIC criterion, and parameter tuning is performed on the validation set to minimize the prediction error.
[0065] For the LSTM model, the network structure mainly includes an input layer (time step = 24, corresponding to 24 hours of historical data), hidden layers (2 LSTM layers, with 64 and 32 neurons respectively), a fully connected layer, and an output layer (predicting the number of faults in the next hour). The Adam optimizer and MSE loss function are used for training, and early stopping is employed to prevent overfitting.
[0066] The association rule mining model uses the Apriori association rule algorithm to uncover hidden associations between fault types, distinctions, business processes, and time periods.
[0067] The training process mainly includes setting the minimum support (e.g., 0.05, meaning that the frequency of a certain association combination accounts for more than 5% of the total sample) and the minimum confidence (e.g., 0.7, meaning that when the conditions are met, the probability of the result occurring is not less than 70%).
[0068] Iteratively mine frequent itemsets on the training set: first generate 1-itemsets (e.g., {network fluctuation failure}), then gradually generate 2-itemsets (e.g., {network fluctuation failure, distinguish A}) and 3-itemsets (e.g., {network fluctuation failure, distinguish A, 18-20 points}), and remove infrequent itemsets that do not meet the minimum support.
[0069] Association rules are generated based on frequent itemsets (e.g., "distinguish between A+18-20 points → network fluctuation failure"), and valid association rules that meet the minimum confidence are selected to form a fault scenario association knowledge base.
[0070] By combining the above models, the data algorithm layer achieves the complementarity of post-event statistics and pre-event prediction of data patterns, generating analysis results oriented towards business scenarios.
[0071] Optionally, in an embodiment, when the data application scenario layer is used to provide the analysis results generated by the data algorithm layer to the upper-layer application, it is specifically used to: provide the analysis results generated by the data algorithm layer to the upper-layer application in the form of a visual large screen display or an API interface.
[0072] The data application scenario layer presents the final business logic and directly targets end users (such as operations and maintenance personnel and business personnel). This layer hosts at least one upper-layer application geared towards specific business scenarios, including a differentiated workbench, a cloud-network workbench, a production-interaction workbench, and a wireless workbench. These workbench platforms encapsulate the analysis results generated by the data algorithm layer into application services tailored to different business scenarios.
[0073] For example, the wireless workbench displays analysis results such as the proportion of network fluctuation failures, prediction of high failure periods, and failure association rules; the differentiated workbench displays the ranking of the number of failures, the distribution of failure types, and the comparison of resource usage in each region.
[0074] The data application scenario layer provides the analysis results to upper-layer applications in the form of visual dashboards or standardized API interfaces, thereby achieving modular and standardized output of governance capabilities.
[0075] The technical solution of the present invention will be described below through specific embodiments: Fault analysis was conducted using 1,000 fault data points from the data center as a sample. Among them, 350 occurred between 7 PM and midnight, 420 were network fluctuation faults, 720 occurred on weekdays, 280 occurred in zone A, and 220 occurred in zone B.
[0076] The system collects device status and fault data in real time through API interfaces, including fields such as fault ID, occurrence time, processing time, involved business, fault level, and location. The data processing layer cleans, standardizes, and classifies the raw data according to context to generate a standardized dataset.
[0077] After acquiring the dataset, the data algorithm layer performs feature engineering and then calls multidimensional aggregation statistical models, time series prediction models, and association rule models for analysis. Fault percentage analysis results: Network fluctuation faults accounted for 45% (first in percentage), and server faults accounted for 28% (second in percentage); fault type A had the most cases (28%), followed by fault type B (22%).
[0078] Time-based occurrence analysis results: 19:00-24:00 is the peak period for failures (accounting for 35%), and the failure rate on weekdays is 78%; the time-series prediction model predicts that "A+19:00-24:00" will be the peak period for failures in the next 7 days.
[0079] Association rule mining results: Association rules such as "distinguish between A+19-24 points → network fluctuation failure" (confidence > 0.7) were mined, forming a fault scenario association knowledge base.
[0080] The above analysis results are displayed on the visualization screen of the wireless workbench and the differentiated workbench through the data application scenario layer. At the same time, they are provided to third-party operation and maintenance work order systems through standardized API interfaces to realize proactive early warning and intelligent work order dispatch.
[0081] In summary, this invention constructs a full-link layered architecture to decouple each link, thereby improving system scalability and maintainability and supporting flexible adaptation to multiple scenarios. It also supports unified access and management of multiple data sources, eliminating data heterogeneity through standardized data processing workflows and achieving data asset integration. Based on this, an algorithm empowerment layer is built, deeply binding statistical analysis, correlation mining, and other models with business scenarios to achieve modularization and reusability of algorithm capabilities. Ultimately, it realizes end-to-end data flow from collection to application, optimizing data processing and algorithm invocation efficiency.
[0082] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
[0083] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.
[0084] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0085] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A full-link data governance system for enhancing the data service capabilities of upper-layer applications, characterized in that, include: Data acquisition layer, data processing layer, data algorithm layer, and data application scenario layer; The data acquisition layer is used to acquire multi-source heterogeneous raw data from multiple data sources, perform verification processing on the acquired multi-source heterogeneous raw data, and transmit the verified multi-source heterogeneous raw data to the data processing layer. The data processing layer is used to store the verified multi-source heterogeneous raw data and to perform data arrangement on the stored multi-source heterogeneous raw data to generate a dataset for the data algorithm layer to call. The data algorithm layer is used to obtain the dataset from the data processing layer and process the dataset using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios. The data application scenario layer is used to support at least one upper-layer application oriented towards business scenarios, and to provide the analysis results generated by the data algorithm layer to the upper-layer application for use.
2. The end-to-end data governance system for enhancing the data service capabilities of upper-layer applications according to claim 1, characterized in that, The data processing layer includes a data center and an application database; The data center is used to store the verified multi-source heterogeneous raw data, and to perform data cleaning, format unification and classification according to business scenarios on the stored multi-source heterogeneous raw data to generate the dataset. The data center is also used to synchronize the dataset to the application database, which is used to store the dataset.
3. The end-to-end data governance system for enhancing the data service capabilities of upper-layer applications according to claim 2, characterized in that, The data center includes a network database and a data mart; The network database is used to receive and store verified multi-source heterogeneous raw data, and to synchronize the verified multi-source heterogeneous raw data to the data mart. The data mart includes multiple data sub-marts divided according to business scenarios. Each data sub-mart is used to store data that has been cleaned, formatted, and classified according to business scenarios.
4. The end-to-end data governance system for enhancing the data service capabilities of upper-layer applications according to claim 3, characterized in that, The network database employs a hybrid storage engine, which includes: pgSQL is used to store structured data from multi-source heterogeneous raw data that has passed validation. MinIO is used to store unstructured data in multi-source heterogeneous raw data that has passed verification; HBase is used to store time-series data from multi-source heterogeneous raw data that has passed verification.
5. A full-link data governance system for enhancing the data service capabilities of upper-layer applications according to claim 1, characterized in that, The data acquisition layer includes lightweight acquisition agents deployed on each data source side and a unified acquisition control module; Each lightweight acquisition agent is used to collect multi-source heterogeneous raw data from various data sources according to the configured acquisition frequency; The unified acquisition and control module is used to verify the acquired multi-source heterogeneous raw data and temporarily store the data that fails the verification in the exception queue.
6. The end-to-end data governance system for enhancing the data service capabilities of upper-layer applications according to claim 1, characterized in that, Multiple data source sides include wireless network side, cloud network side and computing network side.
7. A full-link data governance system for enhancing the data service capabilities of upper-layer applications according to claim 1, characterized in that, The data algorithm layer includes: The feature engineering module is used to perform feature engineering processing on the dataset to obtain feature data; The model processing module is used to process the feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios.
8. A full-link data governance system for enhancing the data service capabilities of upper-layer applications according to claim 7, characterized in that, The feature engineering module is used to perform feature engineering processing on the dataset to obtain feature data, specifically for: The dataset is subjected to feature extraction processing to obtain initial feature data; Based on the variance selection method, the initial feature data is subjected to feature removal processing to obtain preprocessed feature data; Based on the random forest algorithm, feature filtering is performed on the preprocessed feature data to generate the feature data.
9. A full-link data governance system for enhancing the data service capabilities of upper-layer applications according to claim 1, characterized in that, When the data application scenario layer is used to provide the analysis results generated by the data algorithm layer to the upper-layer application, it is specifically used for: The analysis results generated by the data algorithm layer are provided to the upper-layer application in the form of a visual dashboard or an API interface.
10. A full-link data governance system for enhancing the data service capabilities of upper-layer applications according to claim 7, characterized in that, The model processing module is used to process the feature data using at least one pre-trained algorithm model to generate analysis results oriented towards business scenarios. Specifically, it is used for: Based on the feature data, multidimensional aggregation statistics are performed on the dataset, and in the statistical process, a first weight is assigned to the recent data in the dataset, and a second weight is assigned to the long-term data in the dataset, so as to generate data statistical results. Wherein, the recent data is data obtained after a preset time, the long-term data is data obtained before the preset time, and the first weight is greater than the second weight; Based on the feature data, a pre-trained prediction model is used to perform time-series prediction in order to generate predicted values of the target indicator for future time periods. The analysis results are obtained based on the statistical results of the data and the predicted values of the target indicators for the future time period.