Intelligent root cause analysis method and system for operating indicators
The intelligent root cause analysis system, built using ETL tools and causal testing algorithms, solves the problem of low efficiency in manual analysis by operators under multi-dimensional regulatory assessments. It enables rapid and accurate identification of the root causes of operational indicators, thereby improving analysis efficiency and decision support capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2026-03-31
AI Technical Summary
Faced with multi-dimensional regulatory and assessment pressures, operators find that traditional manual root cause analysis methods are inefficient and untimely, making it difficult to quickly and accurately pinpoint the root causes affecting fluctuations in financial indicators. Furthermore, these methods suffer from difficulties in cross-departmental collaboration and a strong reliance on experience.
ETL tools are used to integrate multi-source data and build a unified dimensional data model. The Apriori algorithm is used to mine association rules, and PC causal inference and Granger causality test are combined. The root cause contribution is calculated by the improved Shapley value algorithm to generate a structured diagnostic report and realize intelligent root cause analysis.
It improves analysis efficiency, reduces cross-departmental collaboration time, reduces reliance on experience, can quickly identify non-linear relationships, provides objective data-driven decision support, and shortens the time cycle from problem discovery to root cause location.
Smart Images

Figure CN121766426A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of telecommunications operation and management technology, specifically relating to an intelligent root cause analysis method and system for operational indicators. Background Technology
[0002] The regulatory and assessment pressures faced by telecom operators from the State-owned Assets Supervision and Administration Commission (SASAC) and the Ministry of Industry and Information Technology (MIIT) are multi-dimensional, covering aspects such as financial health, service quality, strategic transformation, and grassroots execution. The 2025 operating indicator system for central state-owned enterprises continued the "one profit, five ratios" framework, but replaced the "operating cash collection ratio" with the "operating cash ratio." The overall requirement is "one increase, one stability, and four improvements." Specifically, this includes stable growth in total profit ("one increase"), overall stability in the asset-liability ratio ("one stability"), and year-on-year improvements in four indicators: return on net assets, R&D expenditure intensity, overall labor productivity, and operating cash collection ratio ("four improvements").
[0003] 1. Under the dual pressure of financial stability and operational efficiency, the "one profit and five ratios" indicator system has been upgraded. With slowing revenue growth, operators need to boost profits through cost reduction (such as zero-based budgeting) and emerging businesses (such as cloud services and AI). However, the shrinking traditional business (declining voice / SMS revenue) exacerbates the difficulty of profitability, requiring improved operating cash collection efficiency and forcing operators to strengthen accounts receivable management, especially regarding the delayed payments for government and enterprise ICT projects. Revenue should be recognized based on "receipt of payment rather than invoicing," reducing "paper profits." Specifically:
[0004] (1) Stable debt-to-asset ratio: restrict aggressive investment and require infrastructure such as computing centers and 5G-A to be built "on demand" (investment after orders are placed) to avoid blind expansion;
[0005] (2) Improved return on net assets: It is necessary to optimize the capital structure and increase the profit contribution of emerging businesses (such as industrial digitalization);
[0006] (3) Increase the intensity of R&D investment: The proportion of R&D investment needs to increase year-on-year, but it needs to be quickly converted into profits (such as front-line R&D personnel going down to participate in customized projects);
[0007] (4) Improve overall labor productivity: With limited staff size, reduce management costs through AI tools and digital processes, and allocate manpower to sales / service positions.
[0008] 2. The contradiction between investment and cost control:
[0009] (1) Prudent investment in infrastructure: 5G-A, computing networks and other technologies require a mature industrial ecosystem (such as terminals and applications) before they can be promoted, which will delay the progress of universal coverage.
[0010] (2) Rigid cost reduction: The contribution of outsourcing and human resources costs to "zero-based budgeting" (such as increased revenue or R&D transformation) needs to be clearly defined, and the effectiveness needs to be verified through post-evaluation, which will increase the pressure of refined management.
[0011] Operators monitor a massive number of multi-dimensional operational metrics daily. When key operational metrics fluctuate or deviate from expectations, quickly and accurately pinpointing the root cause is crucial. Root cause analysis of financial metric fluctuations is a complex but vital task. This involves understanding the vast and dynamic telecommunications market environment, internal operational efficiency, and the impact of strategic decisions.
[0012] Therefore, it is clear that traditional artificial root cause analysis methods face significant challenges, as the macroeconomic environment affecting operators' financial indicators is complex and multifaceted.
[0013] 1. Changes in regulatory policies (tariff controls, spectrum auction rules, interconnection rates, data privacy laws, cybersecurity requirements, antitrust reviews), government subsidies (such as rural broadband construction), and geopolitical risks (international business);
[0014] 2. Changes in user needs (increased data consumption, content preferences, and requirements for service experience), changes in population structure (aging, urbanization), the digital divide, and trends in user loyalty / churn rate;
[0015] 3. Technology iteration cycle (4G->5G->6G), emergence and application of new technologies (cloud computing, AI, IoT, edge computing), changes in equipment costs, escalation of cybersecurity threats, and the impact of OTT services.
[0016] 4. The number and strength of competitors (other operators), the intensity of price wars, the competition for market share, the competition for network quality, and the competition for service innovation.
[0017] 5. The activity level of virtual network operators (MVNOs), the possibility of new players (such as tech giants) entering the communications market, and barriers to entry (spectrum, infrastructure, capital).
[0018] 6. User price sensitivity, collective bargaining power (enterprise customers), switching costs (impact of number portability policy), and information transparency.
[0019] Root cause analysis of fluctuations in telecom operators' financial indicators requires a systematic framework (PESTEL + value chain + indicator tree), combined with detailed internal and external data. Through comparison, correlation, and in-depth investigation, the core driving factors can be identified. Manually performing combined analysis presents the following difficulties:
[0020] 1. Massive Data Volume and Complex Relationships: The sheer number of indicators (hundreds or thousands) and the complex, non-linear relationships between them make manually identifying and verifying all possible connections time-consuming and laborious. For example, internal value chain analysis involves:
[0021] (1) Revenue side: This includes changes in total user growth / churn, prepaid / postpaid user ratio, individual / government / enterprise user ratio, and high-value user ratio; voice / SMS revenue decline vs. data traffic revenue growth, package pricing strategy (impact of unlimited data packages), penetration rate and revenue contribution of value-added services (content, cloud services, security); and core indicators such as broadband access fees, IPTV / OTT content fees, and enterprise leased line revenue.
[0022] (2) Cost side: energy costs, site rental (tower rental), maintenance costs, labor costs; periodic payment or auction cost amortization, inter-network settlement expenditures, advertising, channel commissions, promotional activity costs, etc.
[0023] 2. Financial indicator linkage analysis (key indicator tree): This involves an explosion of dimensions, with indicators capable of drilling down to multiple dimensions. Manually combining and analyzing these dimension combinations is extremely inefficient and prone to missing key dimension combinations. For example:
[0024] (1) Revenue decline: Analyze user numbers, ARPU, and business structure.
[0025] (2) Cost increase: Analyze OPEX (especially electricity, tower rental, labor, and inter-network settlement) / depreciation and amortization.
[0026] (3) Declining profit margin (EBITDA Margin, Net Profit Margin): Analyze whether revenue growth is slower than cost growth, whether the cost structure has deteriorated, and whether pricing power has weakened;
[0027] (4) Changes in cash flow (free cash flow): Analyze the strength of CAPEX and the performance of operating cash flow (OCF) (affected by profit and working capital).
[0028] (5) Debt ratio / leverage changes: Analyze financing strategies, sources of CAPEX investment funds, and dividend policies.
[0029] 3. Poor timeliness: The market changes rapidly, and problems require timely responses. Traditional methods of layer-by-layer reporting and manual analysis are time-consuming and often miss the best time for intervention.
[0030] 4. High dependence on experience: The analysis results are highly dependent on the analyst's experience and intuition, which can easily lead to subjective bias or knowledge blind spots, making it difficult to deal with new and complex abnormal patterns.
[0031] 5. Difficulty in cross-domain collaboration: The root cause may involve data and domain knowledge from multiple departments such as marketing, sales, network, customer service, and IT. Cross-departmental collaboration and information barriers make root cause identification more difficult.
[0032] Therefore, operators urgently need an intelligent, automated, efficient, and accurate root cause analysis solution capable of intelligently discovering correlations between indicators and automatically locating anomalies in multi-dimensional combinations. (The system automatically analyzes massive amounts of data to identify which operational indicators (such as user churn rate, network complaint volume, and revenue) have significant impacts or linkages, without relying entirely on manual experience for setting parameters. When an indicator becomes abnormal (such as a decline in revenue), the system can automatically and quickly filter and drill down from multiple dimensions (time, region, user group, service type, etc.) and their combinations to accurately pinpoint the specific dimensional combination causing the anomaly (e.g., "a sudden drop in revenue for a certain province, a young user group, and 5G packages on weekends")). It quickly locates key factors, accurately identifying the few key root causes that contribute most to indicator fluctuations from a massive pool of potential influencing factors. (The core of root cause analysis is to find the most influential factors among many influencing factors. "The few key root causes that contribute the most to the fluctuation of indicators" refers to finding the most influential factors among many influencing factors.) It improves the efficiency and timeliness of analysis, and significantly shortens the time cycle from discovering a problem to locating the cause. It lowers the experience threshold, provides data-driven objective analysis results, and reduces the over-reliance on personal experience. It supports cross-domain data fusion and effectively integrates heterogeneous data from multiple sources such as business support systems, network management systems, and customer relationship systems. Summary of the Invention
[0033] The purpose of this invention is to provide an intelligent root cause analysis method and system for business indicators in order to solve the problems mentioned above in the background art.
[0034] To achieve the above objectives, the present invention specifically adopts the following technical solution:
[0035] A method for intelligent root cause analysis of business indicators includes the following steps:
[0036] Step S1: Extract order data from the business support system, network performance data from the operation support system, and user profile data from the customer relationship management system using ETL tools. Use semantic mapping tables to convert the fields of each system into standard dimension labels and build a unified dimension data model that includes indicator number, indicator value, dimension attribute and data source fields. The dimension attributes include at least three levels: region, user type and business type.
[0037] Step S2: The Apriori algorithm is used to mine the association rules between indicators. The PC causal inference algorithm is combined with the conditional independence test and Granger causality test to quantify the nonlinear relationship and construct a weighted indicator association graph. The association weights between network performance indicators and business indicators are obtained by training with historical data.
[0038] Step S3: Calculate ±3σ as a dynamic threshold based on historical 7-day period data, and monitor key indicators such as average revenue per user and user churn rate in real time using a sliding window. When the fluctuation of the indicator exceeds the threshold, the root cause analysis process is triggered.
[0039] Step S4: Calculate the information entropy H(D) = -Σp(d)log2p(d) for each dimension, filter out the top 5 key dimensions with entropy values < 0.5, and perform combination analysis only on the selected dimensions to reduce the number of combinations of the original 20-dimensional problem from 1,048,576 to 32.
[0040] Step S5: Using the improved Shapley value algorithm that incorporates business weight factors, calculate the contribution of each factor to the anomaly of the indicator, where the weight of network-related root causes is set to 1.2 and the weight of market-related root causes is set to 1.0.
[0041] Step S6: Generate a structured diagnostic report containing root cause ranking, association evidence, and recommended measures. The report also indicates the impact path and percentage contribution of each root cause.
[0042] The causal inference algorithm in step S2 specifically includes:
[0043] Potential relationships between indicators were discovered by using conditional independence tests in the PC causal inference algorithm, and an undirected causal skeleton graph was constructed.
[0044] Granger causality test is used for indicators with time-series characteristics. A piecewise function model is established to quantify the nonlinear impact. When the price reduction of competitors is ≤30%, the impact coefficient is 0.1x, and when it is >30%, it is 0.3x+5. The impact coefficient is stored as a side weight in the correlation graph.
[0045] The dimensional pruning in step S4 specifically includes:
[0046] Calculate the information entropy values for 20 candidate dimensions, where the entropy value for the region dimension is 0.15, for the user type dimension it is 0.28, and for the terminal brand dimension it is 0.83;
[0047] Retain five key dimensions with entropy values <0.5: region, user type, business type, time period, and channel.
[0048] Only 5-dimensional combination analysis is generated, and abnormal dimensional combination points are located by drilling down layer by layer.
[0049] The improved Shapley value calculation in step S5 is specifically as follows:
[0050] The feature set N is defined to include three categories of influencing factors: network failures, marketing activities, and user behavior.
[0051] Calculate the marginal contribution of each factor.
[0052] α is the business weight factor, with 1.2 for network-type root factors and 1.0 for market-type root factors;
[0053] Output a root cause contribution ranking list and associate it with a specific event ID as evidence.
[0054] In the unified dimensional data model:
[0055] The dimension attribute field stores standardized dimension key-value pairs, including mapping the CRM system's vip_level to the standardized user_type;
[0056] The data source field identifies whether the data comes from the business support system, operations support system, or customer relationship management system, and records the timestamp and version number of the original data.
[0057] The quantification of the nonlinear relationship specifically refers to:
[0058] When the utilization rate of the base station's central processing unit is >85%, the probability of network speed degradation is ≥90%. This rule is obtained through analysis of historical fault data.
[0059] When the user churn rate increases by 1%, the average revenue per user decreases by 0.8%, a coefficient calculated through regression analysis.
[0060] It also includes step S7:
[0061] Receive confirmation or correction feedback from analysts regarding root cause diagnoses. Feedback includes root cause confirmation markers and weight adjustment suggestions.
[0062] The weight coefficients of the edges in the association graph are adjusted based on feedback. For example, the weight of base station CPU utilization and network speed is adjusted from 0.75 to 0.82.
[0063] Generative adversarial networks are used to generate synthetic anomalous data, which are then used to train models to identify novel anomalous patterns.
[0064] An intelligent root cause analysis system for business performance indicators includes:
[0065] The data fusion module is configured to access data from the business support system, operations support system, and customer relationship management system in real time via API interfaces and perform standardized transformations.
[0066] The graph construction module has a built-in PC causal inference algorithm and Granger causal test engine, and stores the association rules and influence coefficients between indicators.
[0067] The real-time detection module uses a sliding window algorithm to calculate dynamic thresholds and configures alarm rules for indicator fluctuations.
[0068] The dimensional pruning module enables dimensional filtering and combination optimization based on information entropy.
[0069] The root cause quantification module executes the improved Shapley value algorithm and outputs a contribution ranking.
[0070] The report generation module converts the analysis results into a diagnostic report that includes a root cause tree and action recommendations.
[0071] The association rules stored in the graph construction module include:
[0072] The mapping relationship between network performance indicators and service indicators, where when the base station CPU utilization rate is >85%, the probability of a decrease in related service indicators is ≥90%.
[0073] The causal relationship between marketing activity metrics and user behavior metrics, such as the impact coefficient of competitor promotional activities on user churn rate being 0.3x+5;
[0074] Each rule is labeled with its data source and confidence level.
[0075] An electronic device includes a memory, a processor, and a computer program stored in the memory. When the processor executes the program, it implements any of the steps of the method described above, specifically including:
[0076] Raw data from various operator systems is obtained through data interfaces, with a data update frequency of once per minute.
[0077] The analysis engine is invoked to execute real-time anomaly detection and root cause analysis algorithms, with analysis latency controlled within 30 seconds.
[0078] Output diagnostic reports to a visualization interface, supporting sorting by root cause contribution and drill-down analysis.
[0079] The analytical method of this invention is performed according to the following analysis:
[0080] This analytical process is a systematic, multi-layered root cause analysis framework. Its core lies in using data correlation and logical deduction to penetrate from superficial indicators (such as declining ARPU) to actionable underlying causes. The entire process can be divided into three key stages:
[0081] First, an analytical baseline is established through anomaly identification and scope definition. The macro-level problem (declining ARPU) is broken down into structural elements (reduced number of users + decreased contribution per user), and then further subdivided into specific problem dimensions (churn in East China, churn of 5G users, and decreased traffic). This layered deconstruction ensures that the analysis is both comprehensive and focused, avoiding omissions of critical paths.
[0082] Secondly, in-depth attribution was conducted for each key dimension, employing a four-dimensional cross-analysis method: "network-competition-product-user". Taking network-side analysis as an example, it not only focuses on traditional KPIs such as base station failure rate, but also emphasizes the spatiotemporal correlation between network indicators (such as PRB utilization rate) and user behavior data (complaint hotspots, APP usage logs) to verify the actual impact of network problems on user experience. This multi-source data cross-validation significantly improves the accuracy of root cause identification.
[0083] Finally, a causal chain is established through correlation analysis. A typical example is aligning the 5G user churn rate curve with the release time of competitor 5G packages, network congestion heatmaps, and user NPS survey results in time and space to identify the dual impact of "competitor's low-priced packages + deterioration of local network quality." This comprehensive evidence chain construction makes the root cause conclusions highly actionable, directly guiding adjustments to network optimization priorities or iterations of package strategies.
[0084] The implementation of the entire system relies on comprehensive data governance in the early stages, ensuring that data such as user profiles, network performance, and business operations can be dynamically correlated and analyzed across the three dimensions of time, space, and user ID. Simultaneously, it is necessary to construct a genealogy graph of metrics, making the mathematical relationships and business logic between atomic metrics (such as single-base station call drop rate) and business metrics (ARPU) explicit. This is the technical foundation for achieving automated root cause analysis.
[0085] The beneficial effects of this invention are as follows:
[0086] This invention addresses the shortcomings of existing technologies and provides the following significant advantages:
[0087] To address the issues of data silos and the inefficiency of manual analysis: ETL tools are used to automatically integrate multi-source data from BSS, OSS, and CRM systems. A unified dimensional model is adopted to standardize the data structure, shortening the traditional 3-5 day cycle of manual analysis and improving efficiency. Based on semantic mapping tables and automated data alignment mechanisms, the problem of inconsistent data standards across systems is resolved, reducing data preparation time.
[0088] To address the issues of dimensional combination explosion and missed key associations, this study innovatively employs information entropy assessment for dimensional pruning. From 20 original dimensions, five key dimensions with entropy values <0.5 are selected, reducing the number of analytical combinations from millions to 32, while maintaining analytical accuracy and reducing computational resource consumption. By combining PC causal inference algorithms with Granger tests, the study accurately identifies nonlinear threshold effects that traditional linear analysis might miss.
[0089] To address the issues of misjudgment of causal relationships and strong reliance on experience: A weighted index correlation graph is constructed to store validated causal rules. By improving the Shapley value algorithm and introducing business weight factors, the objectivity of root cause analysis results is enhanced, improving the decision-making accuracy of novice analysts. The system's built-in adversarial learning mechanism automatically generates abnormal data weekly to train the model, accelerating the identification of new anomaly patterns.
[0090] To address the issues of poor response time and difficulties in cross-departmental collaboration: Anomaly alarms are achieved through dynamic threshold monitoring (±3σ) using a sliding window. Combined with a real-time stream processing engine, this enables end-to-end control from problem detection to root cause analysis, improving the timeliness of handling major faults. Standardized diagnostic reports simultaneously present root causes from multiple domains, including network and marketing (e.g., displaying both base station faults and the impact of marketing activities), thereby improving cross-departmental collaboration efficiency and shortening decision-making meeting times. Attached Figure Description
[0091] Figure 1 This is a schematic diagram of the hierarchical and correlation modeling of indicators in this invention;
[0092] Figure 2 This is a schematic diagram of the spectrum storage (save results) of the present invention;
[0093] Figure 3 This is a schematic diagram illustrating the decrease in ARPU and related indicators in this invention;
[0094] Figure 4 This is a schematic diagram illustrating the potential causes of abnormal indicators in this invention;
[0095] Figure 5 This is a schematic diagram of the model update in this invention;
[0096] Figure 6 This is a schematic diagram of the system architecture of the present invention;
[0097] Figure 7 This is a schematic diagram of the nonlinear correlation of the present invention; Detailed Implementation
[0098] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0099] Example 1
[0100] In the fourth quarter of 2023, a certain telecom operator experienced abnormal fluctuations in its 4G user churn rate in Shanghai, with a weekly increase of 8.7%, significantly exceeding the preset warning threshold of 3%. This anomaly drew the attention of the operator's management, as historically, the churn rate in Shanghai typically remained stable between 2.5% and 3.2%. Traditional analytical methods required coordination among multiple business units, including customer service, network operations, and marketing, involving cross-departmental meetings, retrieving reports from various systems, and manual data comparison. The entire process usually took 2-4 business days to reach a preliminary conclusion. This analytical approach was not only inefficient but also often missed the optimal intervention opportunity in a rapidly changing market environment.
[0101] The present invention is implemented through the following steps:
[0102] The system automatically extracted data from three core systems using an ETL tool. User profile data obtained from the CRM system showed that the affected group was primarily the general public using basic 4G plans. The data structure was standardized using a unified dimensional model, as shown in the following example:
[0103] Typical data samples show that the current churn rate is 8.7%, with three key labels in the dimensional attributes: geographic information (Shanghai), user type (general users), and package type (basic 4G package). The data source is marked as a CRM system, and the timestamp is accurate to the millisecond.
[0104] The system synchronously accessed data from other related systems: network performance data from the OSS system showed that the base station outage rate in Pudong New Area increased by 12 percentage points year-on-year; business data from the BSS system reflected a 35% surge in the number of users whose packages were expiring this week; and the external competitor monitoring system detected that competitors were conducting "year-end data traffic promotion" marketing campaigns. This multi-source heterogeneous data was automatically converted into standard dimension labels through a semantic mapping table, such as mapping the user level field in the CRM system to a unified value grading label.
[0105] The system then performs analysis:
[0106] 1. Association Graph Construction Stage
[0107] The system, through a pre-built algorithm engine, discovered several key correlation paths. The path with the highest weight showed that base station outages (initial weight 0.78) directly lead to a decrease in network availability, which in turn increases user complaints, ultimately resulting in a higher churn rate. This causal relationship was validated through training with historical data, achieving a confidence level of 92%.
[0108] The system also identified an important special pattern: if a user's subscription expires within 3 days and a competitor's promotional activity is also implemented, the user's churn probability will surge to 65%. This non-linear relationship was discovered through Granger causality testing, and the system automatically converted it into a business rule and stored it in the knowledge base.
[0109] 2. Real-time detection trigger mechanism
[0110] Simultaneously, based on historical data from the past 7 days, the system calculates the dynamic threshold for the churn rate indicator. The calculation process shows that the historical mean of this indicator is 3.2%, and the standard deviation is 0.8%. Therefore, according to the ±3σ principle, the trigger threshold is set at 5.6%. The current monitored value of 8.7% significantly exceeds the threshold, and the system immediately triggers the root cause analysis process.
[0111] 3. Dimensional Pruning Optimization Process
[0112] The system evaluated the information entropy of 20 candidate dimensions, with the following results: region dimension entropy value 0.18, user value dimension 0.42, package type dimension 0.25, and terminal model dimension 0.91. Based on the preset entropy threshold of 0.5, the system automatically filtered out the top three key dimensions, reducing the combinatorial complexity of the original problem to a manageable scale (only three dimensions left, compared to the previous permutations and combinations of twenty candidate data).
[0113] 4. Root cause quantitative analysis
[0114] The system employs an improved Shapley value algorithm to calculate root cause contribution. For the factor "Pudong base station outage," the contribution calculation process considers a network-related root cause weight of 1.2, ultimately yielding a contribution of 62%. The combined factor of "package expiration + competitor promotion" uses a market-related weight of 1.0, resulting in a contribution of 34%. These two main root causes collectively explain 96% of the abnormal fluctuations (sufficient to cover most situations; although some analysis results are lost, efficiency is greatly improved).
[0115] The system analysis provides a diagnostic report and decision support.
[0116] First is the structured diagnostic report. The diagnostic report generated by the system includes the following core parts:
[0117] In the section on abnormal indicators, the abnormal 4G user churn rate in Shanghai was clearly identified as 8.7%, an increase of 8.7 percentage points from the baseline. Through dimensional combination analysis, the user group with the most prominent problem was identified as the general public in Shanghai using basic 4G plans, and this combination contributed 57% to the overall anomaly.
[0118] The root cause analysis table details two main root causes: the primary cause is the outage of base stations in Pudong New Area, contributing 62% and linked to specific network alarm numbers; the secondary cause is the combined effect of expired service packages and competitor promotions, contributing 34% and linked to competitor activity numbers. The system also automatically generated descriptions of the impact paths of these two root causes.
[0119] The system provides recommended measures and expected effects. Based on preset policy mapping rules, the system outputs two sets of preferred recommended measures:
[0120] The first priority measure is a network repair plan for the Pudong base station outage, which will be implemented by the network operations and maintenance department. It is expected to reduce the churn rate by 6 percentage points within 48 hours. Based on calculations from similar historical cases, the system gives a confidence interval of [5.2%, 6.8%] for this prediction.
[0121] The second priority measure is a targeted marketing program for users whose contracts are about to expire, led and implemented by the marketing department. The system recommends starting intervention 5 days before the user's contract expires, using exclusive offers to increase renewal rates, which is expected to increase the retention rate of target users by 5%.
[0122] Finally, the implementation effect needs to be evaluated to see if the specific indicators meet expectations.
[0123] Regarding network operation and maintenance, base station repairs were carried out according to the system's recommended priorities, and within 48 hours, the network availability in Pudong New Area was restored to the normal level of 99.8%. Subsequently, the observed churn rate also fell back to a reasonable range of 3.5%.
[0124] In terms of marketing, can targeted retention campaigns for users whose plans are expiring achieve the expected results (now that the network has been repaired, the churn rate has decreased)?
[0125] The system is continuously optimized. First, model parameters were adjusted. Based on the actual situation of this case, the system automatically adjusted two key parameters: the impact coefficient of base station outage on churn rate was increased from 0.78 to 0.81; and "5 days before package expiration" was added as the starting point for the early warning window. These adjustments will make the system's judgment on similar situations in the future more accurate.
[0126] The system uses an adversarial learning mechanism to generate simulated data for various abnormal scenarios, which are then used to train the model to identify more complex situations. In particular, for the combined influencing factors of "covert network failure + competitor promotion," the system has created specific detection rules to ensure faster identification of such complex problems in the future. In short, a new model for "covert network failure + competitor promotion" has been added, and similar models can be directly used for similar problems in the future.
[0127] The lesson learned in this case—that "the best time to intervene is 5 days before the package expires"—was transformed into a standard business rule and stored in the knowledge base. The system also updated the recommendation algorithm for related marketing strategies to ensure more accurate intervention suggestions in similar scenarios. This completed the iteration of the business rule and was also used to update the model.
[0128] This embodiment fully demonstrates the application effect of the intelligent root cause analysis system for operational indicators in a real business scenario. Through this typical case of abnormal 4G user churn rate in Shanghai, we can see that:
[0129] First, the system improves analysis efficiency and shortens the analysis cycle, enabling operators to truly discover problems in real time and respond quickly.
[0130] Secondly, the system's intelligent analysis capabilities far surpass human experience. It can not only handle routine problems, but also discover deeply hidden correlations such as the "overlapping effect of package expiration and competitor promotions," providing a more comprehensive basis for business decisions.
[0131] In summary, the operational logic of this technical solution revolves around data-driven automated analysis, achieving intelligent root cause localization of operator operating indicators through multi-stage collaborative work. The system first extracts heterogeneous data from multiple sources, including the business support system, operations support system, and customer relationship management system. This data is then cleaned and transformed using ETL tools, and standardized into a unified-dimensional analytical model through a semantic mapping table. This step ensures seamless data integration between different systems; for example, it maps user level fields in the customer relationship management system to standardized user type labels, laying the data foundation for subsequent analysis.
[0132] After data integration, the system employs a combination of PC causal inference algorithm and Granger causality test to construct the relationships between indicators. The PC causal inference algorithm discovers potential causal relationships between indicators through conditional independence tests, forming a preliminary causal framework diagram. The Granger causality test, on the other hand, specifically handles indicators with time-series characteristics, quantifying their non-linear impact. For example, when a competitor's price reduction exceeds 30%, the system uses a piecewise function model to calculate its specific impact coefficient on user churn rate. These relationships and impact coefficients are weighted and stored in the indicator correlation graph, forming an iteratively updatable knowledge base.
[0133] When the system calculates a dynamic threshold based on historical data and detects anomalies in indicators, it immediately triggers the root cause analysis process. For example, by analyzing the mean and standard deviation of data from the same period over the past 7 days, a fluctuation range of ±3σ is set as the threshold. Once it is found that the churn rate of users in a certain region exceeds the threshold, the system will initiate a multi-dimensional analysis process. To efficiently handle massive combinations of dimensions, the system calculates the information entropy value of each dimension and prioritizes key dimensions with entropy values below 0.5 for combination analysis. This pruning strategy reduces the original millions of possible combinations to dozens, significantly improving analysis efficiency.
[0134] In the root cause quantification phase, the system employs an improved Shapley score algorithm, comprehensively considering the marginal contribution and business weight of each factor to calculate its specific impact on indicator anomalies. Network-related factors, such as base station failures, are assigned higher weights, while market-related factors, such as competitor activities, use standard weights, ensuring that the analysis results are both objective and consistent with business realities. The final diagnostic report not only includes root cause ranking and contribution percentages but also links to specific evidence chains, such as network alarm numbers or marketing campaign records, providing decision-makers with clear operational guidelines.
[0135] The entire system is continuously optimized through a closed-loop feedback mechanism. When analysts provide corrections to the diagnostic results, the system adjusts the weight coefficients in the correlation graph accordingly. Simultaneously, generative adversarial networks are used to synthesize various anomaly scenario data to train models for identifying new anomaly patterns. This self-iterative capability enables the system to adapt to constantly changing business environments and maintain consistently high analytical accuracy. The fully automated process from data input to output not only reduces the traditional manual analysis cycle of several days to minutes but also reduces reliance on experience through quantitative models, providing operators with efficient and accurate decision support.
[0136] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0137] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the present invention.
[0138] The spirit and scope of this application. Thus, if these modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
Claims
1. A method for intelligent root cause analysis of business indicators, characterized in that The method comprises the following steps: Step S1: Extracting order data of a business support system, network performance data of an operation support system, and user portrait data of a customer relationship management system through an ETL tool to construct a unified dimension data model comprising an index number, an index value, a dimension attribute, and a data source field; Step S2: Conducting conditional independence test using a PC causal inference algorithm to construct a causal skeleton, and quantifying a nonlinear relationship by combining Granger causality; Step S3: Calculating ±3σ as a dynamic threshold based on historical 7-day same-period data to monitor the fluctuation of key indicators such as average revenue per user in real time; Step S4: Calculating the entropy of each dimension H(D) = -∑p(d)log2p(d), and screening the top 5 dimensions with an entropy value <0.5 for combined analysis; Step S5: Using an improved Shapley value algorithm with a business weight factor to calculate the contribution of each factor to the index anomaly; Step S6: Generating a structured diagnostic report comprising root cause ranking, correlation evidence, and recommended measures.
2. The method of claim 1, wherein, The causal inference algorithm in step S2 specifically comprises: Discovering potential correlation between indicators through conditional independence test in the PC causal inference algorithm; Using Granger causality test on indicators with time sequence characteristics to establish a segmented function model to quantify nonlinear influence, with an influence coefficient of 0.1x when the competitor price drop is ≤30%, and 0.3x+5 when the competitor price drop is >30%.
3. The method of claim 1, wherein, The dimension pruning in step S4 specifically comprises: Calculating the entropy values of 20 candidate dimensions; Retaining key dimensions such as region and user type with an entropy value <0.5; Generating combined analysis of only 5 dimensions, reducing the number of combinations from millions to 32.
4. The method of claim 1, wherein, The improved Shapley value calculation in step S5 specifically comprises: Defining a feature set N comprising network failure, market activity, and other influencing factors; Calculating the marginal contribution of each factor where a is a traffic weight factor.
5. The method of claim 1, wherein, In the unified dimension data model: The dimension attribute field stores standardized dimension key-value pairs, including region, user type, etc. The data source field identifies data from business support systems, operation support systems, or customer relationship management systems.
6. The method of claim 2, wherein, The nonlinear relationship quantification specifically comprises: When the base station central processor usage is >85%, the network speed reduction probability is ≥90%; When the user off-network rate increases by 1%, the average revenue per user decreases by 0.8%.
7. The method of claim 1, wherein, Further comprising step S7: Receiving analyst confirmation or correction feedback on root cause diagnosis; Adjusting the weight coefficient of the edge in the correlation graph according to the feedback; Generating synthetic abnormal data through a generative adversarial network to optimize the model.
8. An intelligent root cause analysis system for business indicators, characterized by Comprise: A data fusion module for performing step S1 of claim 1 to realize multi-source data extraction and standardization; A graph construction module for performing step S2 of claim 1 to store correlation rules and influence coefficients between indicators; A real-time detection module for performing step S3 of claim 1 to configure dynamic threshold alarm rules; A dimension pruning module for performing step S4 of claim 1 to realize dimension screening based on entropy values; A root cause quantification module for performing step S5 of claim 1 to calculate the contribution of each factor. A report generation module is configured to execute the step S6 of claim 1 to output the interpretable diagnosis report.
9. The system of claim 8, wherein, The association rules stored in the atlas construction module include: A mapping relationship between the network performance indicators and the service indicators, wherein when the base station load indicator exceeds a threshold, the associated service indicator decreases with a probability of ≥P; A causal relationship and an influence coefficient between the market activity indicators and the user behavior indicators.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, is arranged to perform the method of any one of claims 1 to 9. The processor implements the method steps of any one of claims 1-7 when executing the program, and specifically includes: Obtaining original data of each system of the operator through a data interface; Calling an analysis engine to execute a root cause analysis algorithm; Outputting a diagnosis report to a visual display interface.