An embedded operation and maintenance multi-layer decision chain layer-by-layer returning method and system
Patent Information
- Application Number
- CN202611281257.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-18
AI Technical Summary
规则匹配方法虽然响应速度快、资源开销低,但仅能处理已知故障模式,对于渐变型故障(如光模块激光器老化导致的接收光功率缓慢下降、天线端口氧化导致的驻波比逐日上升)和复杂多指标关联故障(如驻波比异常引起功放温度上升进而触发发射功率回退)缺乏有效的检测和诊断能力;基于大模型的深度分析方法虽然分析精度高、能够处理复杂根因推理,但单次调用延迟高达数百毫秒至数秒,且需要消耗大量计算资源和网络带宽,在嵌入式设备上高频使用将导致严重的性能瓶颈和响应超时
[0015] As can be seen from the above technical solution, the present invention has at least the following advantages and positive effects compared with the prior art: The present invention constructs a four-layer progressive decision chain consisting of a rule engine layer, a statistical analysis layer, a semantic cache layer, and a deep analysis layer, combined with layer-by-layer confidence threshold determination and a short-circuit return mechanism, enabling simple requests to be processed in milliseconds, and complex requests to be progressively upgraded to the deep analysis layer, taking into account the comprehensive requirements of response speed, analysis accuracy, and computational overhead in embedded resource-constrained environments; the analysis results of each layer are passed step by step through the context structure and are completely reused by subsequent layers, avoiding redundant calculations, and the L3 layer utilizes the calculations already performed by the L1 layer. The multi-indicator correlation coefficient can accurately pinpoint the root cause without starting from scratch; through delayed budget management, it automatically degrades to local fast analysis when the remaining budget is insufficient, ensuring predictable response time and avoiding timeouts, thus avoiding the impact of long-tail latency in the deep analysis layer on the overall response; when the deep analysis layer is unavailable, the decision chain is automatically reduced to a three-layer local processing mode, and the core alarm function is unaffected; the telecom equipment-specific rules and Z-Score trend analysis customized for 5G/4G small cell operation and maintenance scenarios can identify gradual faults such as antenna aging and laser degradation in advance before fixed threshold triggering, enabling preventive maintenance.
Smart Images

Figure CN122777412A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of embedded system operation and maintenance technology, specifically to a multi-layer decision chain return method and system for embedded operation and maintenance. Background Technology
[0002] In 5G / 4G mobile communication networks, Edge Units (EUs) manage multiple Remote Units (RUs). EU devices typically use ARM Cortex-A55 or RISC-V processors, with single-core or dual-core configurations, 64MB to 128MB of memory, and run embedded Linux systems. Computing and storage resources are strictly limited. The AI-powered operations and maintenance (O&M) services running on the EU side need to analyze the O&M data reported by the RUs in real time and make decisions to ensure the stable operation of the communication network. The O&M data reported by the RUs includes system-level metrics (CPU, memory, storage, temperature, network) and telecommunications service metrics (RF transmit power, VSWR, optical module transmit / receive power, fronthaul link bit error rate, air interface signal quality, etc.).
[0003] Existing analysis methods in embedded operations and maintenance systems mainly include rule-matching methods based on fixed thresholds and deep analysis methods based on large models. While rule-matching methods offer fast response times and low resource overhead, they can only handle known fault modes and lack effective detection and diagnostic capabilities for gradual faults (such as a slow decrease in received optical power due to laser aging in optical modules or a daily increase in VSWR due to antenna port oxidation) and complex multi-indicator correlated faults (such as an abnormal VSWR causing a rise in power amplifier temperature, which then triggers transmit power backoff). Deep analysis methods based on large models, while offering high analysis accuracy and capable of handling complex root cause reasoning, suffer from latency ranging from hundreds of milliseconds to several seconds per call and require significant computational resources and network bandwidth. Frequent use on embedded devices will lead to severe performance bottlenecks and response timeouts. These two methods operate independently, lacking an organically integrated decision-making mechanism. This results in both simple and complex requests following the same analysis path, causing unnecessary resource overhead or unacceptable analysis latency, making it difficult to meet the comprehensive requirements of embedded operations and maintenance scenarios regarding response time, analysis accuracy, and resource overhead. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for layer-by-layer return of a multi-layer decision chain in embedded operation and maintenance, so as to solve the technical problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: According to one aspect of the present invention, a method for layer-by-layer return of a multi-level decision chain in embedded operations and maintenance is provided, the method comprising: Receive operation and maintenance analysis requests and create a context for storing request parameters and intermediate results of analysis at each level; The execution rule engine layer matches the request features with preset rules. If the confidence level of the matching result reaches or exceeds the first threshold, the matching result is returned directly; otherwise, the matching result is stored in the context and the statistical analysis layer is executed. The statistical analysis layer is executed to perform statistical anomaly detection and trend analysis on the indicators to be analyzed in the context. If the confidence level of the statistical analysis result reaches or exceeds the second threshold, the statistical analysis result is returned directly; otherwise, the statistical analysis result is stored in the context and the semantic caching layer is executed. The semantic caching layer is executed to perform hash calculation on the query features in the context and match them with historical analysis results in the cache. If the cache is hit and the similarity confidence reaches or exceeds the third threshold, the historical results in the cache are reused; otherwise, the cached query results are stored in the context and the deep analysis layer is executed. The deep analysis layer is executed, and the analysis results of each layer stored in the context are read to construct analysis prompt words. An external large language model is called to perform deep reasoning and obtain analysis results. If the latency consumed before executing the deep analysis layer reaches or exceeds the preset total latency budget, the deep analysis is skipped and the statistical analysis layer is re-executed to obtain analysis results. The final analysis results are returned, and the execution path identifier of the decision chain, the confidence level of each layer, and the latency data are written to the database for persistent storage.
[0006] Based on the aforementioned scheme, the confidence level of the rule engine layer is calculated based on the severity score of the matching rule, and the first threshold is a preset high confidence threshold.
[0007] Based on the aforementioned scheme, the statistical analysis layer performs statistical standardization detection and linear regression trend analysis on the indicator to be analyzed using a sliding window, and calculates the confidence level of the statistical analysis results by weighting the anomaly detection intensity and the trend fit goodness.
[0008] Based on the aforementioned scheme, the semantic caching layer uses a non-cryptographic hash algorithm to generate the hash value of the query feature, and the similarity confidence is calculated based on the weighted average of the temporal proximity and numerical proximity between the current query and the cached historical results; the cache adopts a least recently used eviction policy, and cache entries have a preset lifespan.
[0009] Based on the aforementioned scheme, the deep analysis layer calls the external large language model application interface through an adapter. The adapter includes a rate limiter and a circuit breaker. The rate limiter is used to control the calling frequency. The circuit breaker switches to the open state when the error rate exceeds a preset threshold, and enters the half-open state after cooling. When the detection success rate exceeds the preset threshold in the half-open state, it returns to the closed state.
[0010] Based on the aforementioned scheme, the preset total latency budget is dynamically adjusted according to the priority of the analysis request; when the remaining latency budget before entering the deep analysis layer is less than the preset minimum execution time threshold, the deep analysis is skipped and the statistical analysis layer is re-executed.
[0011] Based on the aforementioned scheme, the preset rules include dedicated rules for communication equipment hardware indicators. These dedicated rules include one or more of the following: antenna port VSWR detection rules, power amplifier temperature detection rules, optical module received optical power detection rules, fronthaul link bit error rate detection rules, transmit power deviation detection rules, and reference signal received power detection rules.
[0012] Based on the aforementioned scheme, the rule engine layer runs in a script engine sandbox environment. The sandbox environment has resource consumption limits and prohibits the execution of system commands, file operations, and dynamic loading of unauthorized code.
[0013] Based on the aforementioned scheme, the method further includes: when the external large language model fails to be called continuously or the network connection is unavailable, a degradation mode is triggered, reducing the decision chain to a three-layer local processing mode of rule engine layer, statistical analysis layer and semantic cache layer; during the degradation period, a liveness detection request is sent to the external large language model at a preset interval, and the deep analysis layer is automatically restored after the liveness detection is successful.
[0014] According to another aspect of the present invention, a multi-layer decision chain back-to-back system for embedded operation and maintenance is provided, the system comprising: The request receiving module is used to receive operation and maintenance analysis requests and create a context for storing request parameters and intermediate results of analysis at each level. The rules engine module is used to match request features with preset rules. If the confidence level of the matching result reaches or exceeds the first threshold, the matching result is returned directly; otherwise, the matching result is stored in the context. The statistical analysis module is used to perform statistical anomaly detection and trend analysis on the indicators to be analyzed in the context. If the confidence level of the statistical analysis result reaches or exceeds the second threshold, the statistical analysis result is returned directly; otherwise, the statistical analysis result is stored in the context. The semantic caching module is used to perform hash calculations on the query features in the context and match historical analysis results in the cache. If the cache hits and the similarity confidence reaches or exceeds the third threshold, the historical results in the cache are reused; otherwise, the cached query results are stored in the context. The deep analysis module is used to read the analysis results stored in the context to construct analysis prompt words, and call an external large language model to perform deep reasoning to obtain analysis results. If the latency consumed before the deep analysis is executed reaches or exceeds the preset total latency budget, the deep analysis is skipped and the statistical analysis module is triggered to re-execute the statistical analysis to obtain analysis results. The output persistence module is used to return the final analysis results and write the execution path identifier of the decision chain, the confidence level of each layer, and the latency data into the database for persistent storage.
[0015] As can be seen from the above technical solution, the present invention has at least the following advantages and positive effects compared with the prior art: The present invention constructs a four-layer progressive decision chain consisting of a rule engine layer, a statistical analysis layer, a semantic cache layer, and a deep analysis layer, combined with layer-by-layer confidence threshold determination and a short-circuit return mechanism, enabling simple requests to be processed in milliseconds, and complex requests to be progressively upgraded to the deep analysis layer, taking into account the comprehensive requirements of response speed, analysis accuracy, and computational overhead in embedded resource-constrained environments; the analysis results of each layer are passed step by step through the context structure and are completely reused by subsequent layers, avoiding redundant calculations, and the L3 layer utilizes the calculations already performed by the L1 layer. The multi-indicator correlation coefficient can accurately pinpoint the root cause without starting from scratch; through delayed budget management, it automatically degrades to local fast analysis when the remaining budget is insufficient, ensuring predictable response time and avoiding timeouts, thus avoiding the impact of long-tail latency in the deep analysis layer on the overall response; when the deep analysis layer is unavailable, the decision chain is automatically reduced to a three-layer local processing mode, and the core alarm function is unaffected; the telecom equipment-specific rules and Z-Score trend analysis customized for 5G / 4G small cell operation and maintenance scenarios can identify gradual faults such as antenna aging and laser degradation in advance before fixed threshold triggering, enabling preventive maintenance.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram of a multi-layer decision chain return method for embedded operation and maintenance according to the present invention; Figure 2 This is a flowchart illustrating a multi-layer decision chain return method for embedded operation and maintenance according to the present invention. Detailed Implementation
[0018] To more clearly illustrate the purpose, technical solutions, and advantages of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein. On the contrary, these embodiments are provided so that the present invention will be more comprehensive and complete, and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0019] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.
[0020] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0021] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily need to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0022] The present invention will now be described in detail with reference to specific embodiments.
[0023] Example 1
[0024] like Figure 1 , 2 As shown, this embodiment provides a multi-layer decision chain return method for embedded operations and maintenance. This method runs on resource-constrained embedded devices, where computing and storage resources are limited. Each layer of decision-making is configured with an independent confidence threshold and latency budget. The multi-layer decision chain is specifically implemented as a four-layer progressive decision chain, consisting of an L0 rule engine layer, an L1 statistical analysis layer, an L2 semantic cache layer, and an L3 deep analysis layer. The specific steps of this method are as follows: S1: Receive operation and maintenance analysis requests and create a context for storing request parameters and intermediate results of analysis at each level.
[0025] In this embodiment, on the EU (Edge Unit) side, the global AI analysis engine runs in an embedded Linux operating system process; this step is executed when the decision chain module is triggered. The global AI analysis engine receives analysis requests from two sources, including: 1) system-level indicators (including CPU, memory, storage, temperature, and network) and telecommunications service indicators (including direct hardware acquisition such as RF transmit power (TX power), VSWR, EVM, CPRI / eCPRI link bit error rate (BER), optical module Tx / Rx optical power and bias current, and service software reading such as RSRP / SINR / CQI and service load indicators); 2) operation and maintenance analysis tasks initiated locally by the EU (such as periodic health checks or manually triggered diagnostic requests).
[0026] The request is delivered to the decision chain module as a structured message, serving as the input for the initial analysis request. Upon receiving the initial analysis request, the decision chain module creates a context structure within the EU's local Linux process to hold the data for the entire decision chain's execution. This context structure is defined using a C language structure, and its internal fields are divided as follows: 1) The raw analysis request storage area is used to store the complete parameters of this analysis request, including the request type (such as anomaly analysis, trend prediction, log analysis or health assessment), target device ID (corresponding to the specific RU or EU device), analysis time window and list of indicators to be analyzed; the data in this area is written once during context initialization and remains unchanged during the entire decision chain execution, and is available for reading by each layer.
[0027] 2) L0 layer result storage area, reserved for storing the output results after the L0 rule engine layer is executed, including: the matched rule ID (unique identifier of the matched rule), the rule matching degree, and the severity score (S). rule The region is filled after the L0 layer is completed, with values ranging from 1 to 100 and the L0 layer confidence score (C0).
[0028] 3) L1 layer result storage area, reserved for storing the output results after the L1 statistical analysis layer is executed, including: Z-Score anomaly detection results (including mean μ, standard deviation σ, anomaly markers), linear regression trend analysis results (including slope α, coefficient of determination R²), multi-index correlation coefficient matrix, and L1 layer confidence score (C1); this area is filled after the L1 layer is executed.
[0029] 4) L2 layer result storage area, reserved for storing the output results after the L2 semantic cache layer is executed, including: FNV-1a query hash value (H fnv ), cache hit / miss flags, historical matching result IDs, and time similarity (sim t ) and value range similarity (sim v ), and the L2 layer confidence score (C2); this region is filled after the L2 layer is completed.
[0030] 5) Accumulated intermediate result storage area, used to store all intermediate calculation results from L0 layer to the currently executed layer, as well as copies of the original analysis data passed from each layer; this area is organized in the form of key-value pairs or structured arrays to ensure that higher layers (such as L3 layer) can fully reuse the intermediate data that has been calculated in lower layers and avoid duplicate calculations.
[0031] When creating the context structure, the global AI analysis engine first allocates enough memory blocks in the Linux process's heap space to accommodate all the above fields; it then copies the original analysis request to the original analysis request storage area and sets all fields in the L0, L1, L2 result storage areas and the cumulative intermediate result storage area to zero (or the corresponding invalid state value) to form the initial context.
[0032] After initialization, the context structure, serving as the sole parameter pointer, is passed down the decision chain: first to the L0 rule engine layer; after L0 completes its execution, the results are filled into the corresponding fields and the confidence score is updated; if an upgrade is needed, the context carrying the L0 layer results is passed to the L1 layer, and so on. Each layer reads the completed calculation results from the previous layer from the context structure and appends the analysis data and confidence score generated by the current layer to the corresponding fields of the context structure. The entire process is completed within the EU's local Linux process via pointer references, without inter-process communication (IPC), thus ensuring extremely low data transfer overhead.
[0033] S2: Execute the rule engine layer, match the request features with preset rules, and if the confidence level of the matching result reaches or exceeds the first threshold, return the matching result directly; otherwise, store the matching result in the context and execute the statistical analysis layer.
[0034] In this embodiment, the L0 rule engine layer is the first layer of the four-layer progressive decision chain and also the layer with the fastest response speed. The L0 layer runs rule matching based on the Lua 5.4.6 script engine embedded on the EU side, which can accurately identify known fault modes within 1 millisecond and directly terminate the decision chain through the short-circuit return mechanism to avoid unnecessary calculations in subsequent layers.
[0035] When a context structure containing the request type, target device ID, time window, and list of analytical metrics is passed to the L0 layer, the L0 layer first reads the target device ID and the data of the metrics to be analyzed from the original analysis request field of the context structure. The L0 layer extracts key features from the context, including device type, metric name, outlier range, and alarm type, as input parameters for subsequent rule matching.
[0036] The L0 layer rule scripts are stored in the ` / etc / osmo / ai / log_rules / ` directory, categorized into `base.json` (basic system rules), `business.json` (business application rules), `hardware.json` (hardware-related rules), `msg_service.json` (message service rules), and `other.json` (other rules), with a maximum of 100 rules globally. The Lua engine loads and executes each rule for matching. Each rule contains: id: A unique identifier for the rule; pattern: POSIX regular expression matching pattern; level: Alarm level after matching (INFO / WARN / ERROR); category: rule classification; score: Severity rating (1-100); suggestion: Recommended actions after matching; actions: A list of actions to trigger; cooldown_s: Cooldown time.
[0037] The L0 layer performs regular expression matching between the extracted request features and the pattern field of the aforementioned rules. If a feature successfully matches the pattern of a rule, it is considered a hit for that rule, and the score value of that rule is read as a severity score.
[0038] The confidence score for layer L0 is calculated based on the rule severity score, using the following formula: ; in, Rate the severity of the matched Lua rules (1-100). When a rule is a perfect match, , (i.e., 100%).
[0039] The L0 layer will calculate the results. With the preset L0 short-circuit threshold Comparison, The default value is 0.95 (i.e., 95%). Specifically: if (i.e., confidence level ≥ 95%), the decision chain terminates at the L0 layer via short-circuit. The L0 layer writes the matched rule ID, severity score, confidence score, and suggested action into the L0 layer result field of the context structure, sets the short-circuit flag, and then directly returns the L0 layer analysis results; if If the confidence level is less than 95%, then the L0 layer has not met the short-circuit condition. The L0 layer writes the matching results (including the matched rule ID, severity score, and confidence score) into the L0 layer result field of the context structure and records the matching details in the cumulative intermediate data field, and the decision chain is upgraded to the L1 layer.
[0040] If the requested feature does not match any rule, i.e., the L0 layer rule is not matched, the L0 layer writes the mismatch flag to the L0 layer result field of the context structure and passes the context to the L1 layer for further processing.
[0041] The L0 layer Lua engine runs in a sandbox environment with security restrictions: a memory limit of 4MB and an instruction limit of 10. 7 The execution timeout is 5 seconds. Specific security measures include: disabling os.execute to prevent the execution of system commands; disabling io.popen to prevent the opening of process pipes; disabling io.open to prevent direct file operations; restricting loadstring and loadfile to prevent the dynamic loading of unauthorized code; and setting memory allocation limits via lua_gc.
[0042] The aforementioned sandbox mechanism ensures that the L0 layer operates safely in an embedded resource-constrained environment (EU device memory 64MB~128MB, ARM Cortex-A55 or RISC-V processor), preventing malicious or erroneous Lua rules from consuming excessive system resources and affecting the normal operation of other processes on the EU side.
[0043] Each rule includes a `cooldown_s` field. When a rule is successfully matched, the L0 layer performs a cooling-down process on the rule within the time window specified by `cooldown_s` (e.g., 60 seconds). This means that if the same device triggers the same rule again within the cooling-down period, the L0 layer will not repeat the alarm action, but will only update the matching timestamp. This mechanism can prevent alarm storms caused by RU devices repeatedly reporting the same anomaly in a short period of time, and reduce the processing pressure on the decision chain.
[0044] In 5G / 4G small cell operation and maintenance scenarios, Layer 0 pre-configures dedicated rules for telecom equipment (pRRU / RRU / HUB), managed through the hardware.json rule file, with a maximum of 64 telecom-specific rules globally (the total limit for Layer 0 rules is 100). Typical telecom-specific rules include: TELECOM-001: VSWR>3.0, ERROR level, severity score 95; TELECOM-002: PA temperature >85℃, ERROR level, severity score 100; TELECOM-003: SFP Rx optical power <-25dBm, WARN level, severity score 90; TELECOM-004: CPRI BER>1×10 -9 ERROR level, severity score 95; TELECOM-005: Transmit power deviation >3dB, WARN level, severity score 85; TELECOM-006: RSRP < -110dBm, WARN level, severity score 80.
[0045] If the device type in the request features belongs to pRRU or RRU, the L0 layer will first load the telecom-specific rules from hardware.json for matching; if the match is successful and the confidence level is ≥95%, the above short-circuit logic will be returned; this priority matching mechanism ensures that critical faults of telecom equipment (such as antenna feeder faults, power amplifier overheating, and fronthaul link degradation) can be identified and responded to as quickly as possible (<1ms).
[0046] The expected latency for layer L0 is less than 1 millisecond. This can be achieved by setting an execution timeout (5 seconds) and instruction limit (10) in the Lua sandbox. 7 Through optimization techniques such as pre-compiling rule scripts into bytecode, the actual execution time of the L0 layer is kept within 1 millisecond even in the worst case where all 100 rules are matched, thus ensuring the feasibility of the low-latency goal of the decision chain.
[0047] After completing step S2, if a short-circuit return is not triggered, the system will proceed to step S3 (L1 statistical analysis layer) with the context containing the L0 layer matching results.
[0048] S3: Execute the statistical analysis layer to perform statistical anomaly detection and trend analysis on the indicators to be analyzed in the context. If the confidence level of the statistical analysis result reaches or exceeds the second threshold, the statistical analysis result is returned directly; otherwise, the statistical analysis result is stored in the context and the semantic caching layer is executed.
[0049] In this embodiment, the L1 statistical analysis layer is the second layer of a four-layer progressive decision chain, with a response latency of less than 10 milliseconds. The L1 layer is based on a statistical algorithm engine implemented in native C language, running within the EU's native Linux process, and is used to detect gradual faults and statistical anomalies that cannot be identified by the fixed threshold rules of the L0 layer.
[0050] When the context structure carrying the L0 layer matching results is passed to the L1 layer, the L1 layer first reads the L0 layer rule matching results (including the hit rule ID, severity score, and confidence score) from the L0 layer result field of the context structure, and reads the target device ID, time window, and list of analysis indicators from the original analysis request field. Based on the cumulative intermediate data field in the context, the L1 layer obtains the historical time series data of the indicator to be analyzed (the default sliding window size is 60 historical collection points) as the data basis for subsequent statistical calculations.
[0051] The L1 layer performs 3-Sigma anomaly detection on the currently acquired values. Specifically, for a given metric to be analyzed (such as PA temperature, VSWR, SFP received optical power, etc.), the L1 layer calculates the historical mean μ and historical standard deviation σ based on N historical data points within a sliding window (N is 60 by default). The calculation formula is as follows: ; ; in, Let x be the j-th historical data point within the sliding window. Calculate the standardized Z-Score when the current collected value is x: ; When |Z|>3.0, the current value is determined to be statistically abnormal, that is, it exceeds the normal range defined by the 3-Sigma criterion; this abnormality label is used as one of the inputs for subsequent confidence calculation.
[0052] While performing Z-Score anomaly detection, the L1 layer performs linear regression trend analysis on historical time-series data to identify the direction of continuous change in the indicators. The least squares method is used to fit the historical data points to obtain the linear regression equation. , where a is the slope, b is the intercept, and t is the time series index. Calculate the coefficient of determination. Used to evaluate the goodness of linear fit. The value ranges from 0 to 1, with values closer to 1 indicating a stronger trend. The sign of the slope 'a' indicates the direction of the trend; a > 0 indicates an upward trend, and a < 0 indicates a downward trend. or When the trend fits well, the analytical conclusions have high credibility.
[0053] In scenarios with multiple outliers, the L1 layer can optionally perform multi-indicator correlation coefficient analysis to calculate the linear correlation strength between multiple indicators. Specifically, the Pearson correlation coefficient can be used to calculate the correlation coefficient r between each pair of indicators, with a value range of [-1, 1]; the closer |r| is to 1, the stronger the correlation, and the positive or negative sign indicates a positive or negative correlation. The result of this correlation coefficient calculation can help determine whether there is a correlation between multiple anomalies and serve as the input context for subsequent L3 layer LLM deep analysis.
[0054] The L1 layer combines the Z-Score anomaly detection strength with the goodness of fit of the linear regression trend to calculate the L1 layer confidence score. The calculation formula is as follows: ; in, , is the Z-Score anomaly intensity normalization function. When hour, f (Z)=1.0 indicates that the anomaly intensity has reached its maximum value; To determine the goodness of fit to the trend, the coefficient of determination of the linear regression is directly taken. ; and These are the weighting coefficients. ,default , The above formula ensures that when the indicator simultaneously exhibits a significant statistical anomaly (|Z| is large) and a clear trend direction (…), the indicator will be effective. (High) A higher value indicates a higher degree of reliability in the analytical conclusions.
[0055] The calculated results With the preset L1 short-circuit threshold (Default 0.90, i.e., 90%) for comparison: like Furthermore, if the analysis conclusions are clear (e.g., a clear trend is detected such as "CPU temperature continues to rise, predicting it will exceed the threshold in 30 minutes"), the decision chain terminates at the L1 layer. The L1 layer will then perform statistical analysis (including mean μ, standard deviation σ, Z-score, outlier markers, linear regression slope a, and coefficient of determination). and confidence level Write the L1 layer result field into the context structure, set the short-circuit flag, and directly return the L1 layer analysis result; like , or although If the target is met but the statistical anomaly lacks historical pattern matching (i.e. the cause of the anomaly cannot be determined), then the L1 layer has not met the short-circuit condition; the L1 layer writes all the above statistical results into the L1 layer result field and the cumulative intermediate data field of the context structure, and the decision chain is upgraded to the L2 layer.
[0056] In 5G / 4G small cell operation and maintenance scenarios, Layer 1 performs trend analysis on key indicators of telecommunications equipment, which can detect gradual faults that cannot be captured by the fixed thresholds of Layer 0. For example, typical scenarios include: (1) VSWR Trend Analysis (Antenna Aging Detection): For the VSWR index, the mean μ and standard deviation σ are calculated using 60 historical acquisition points as a sliding window. When |Z|>3.0, it is judged as a statistical anomaly. Further linear regression analysis is performed on the slope a: if a>0 and R 2 A value >0.8 indicates a continuously rising VSWR, presumably due to antenna port aging or connector oxidation causing performance degradation of the antenna feed system. At this point, the L1 layer can output "VSWR trend increasing (slope a = +0.02 / day, R..." 2 The analysis conclusion, "VSWR = 0.85), predicts that VSWR will exceed the 3.0 threshold in 10 days, and recommends preventive maintenance," has a confidence level of over 90%.
[0057] (2) SFP optical module Rx optical power trend analysis (laser aging test): The Z-Score of the received optical power of the SFP optical module is measured. When the Z-Score is continuously negative (|Z|>3.0) and the linear regression slope a<0, R 2 A value >0.7 indicates a continuous downward trend in received optical power, presumably due to aging of the optical module laser or contamination of the fiber optic connector. The L1 layer can output "SFP Rx" optical power with a decreasing trend (slope a = -0.15 dBm / day, R...). 2 The analysis concluded that "the value is 0.78, and it is predicted that it will fall below the -25dBm threshold in 30 days, so it is recommended to clean the fiber optic connector or replace the optical module."
[0058] (3) PA Temperature Trend Analysis (Degradation Detection of Heat Dissipation System): The Z-Score of the power amplifier module temperature is measured. When the temperature Z-Score is consistently positive (|Z|>3.0) and the linear regression slope a>0, R 2 A value >0.75 indicates that the power amplifier temperature is showing a continuous upward trend, but has not yet reached the fixed threshold of 85℃ for L0 layer. L1 layer can identify heat dissipation system degradation (reduced fan speed, dust accumulation on heat sinks, etc.) in advance, outputting "PA temperature trend rising (slope a = +0.3℃ / day, R..."). 2 =0.82), currently at 78℃, predicted to reach the 85℃ threshold in 23 days, and the warning conclusion is: "Check the cooling fan and ventilation channel".
[0059] In the three scenarios described above, the L1 layer, through Z-Score anomaly detection combined with linear regression trend analysis, can detect gradual-type faults before the fixed threshold rules of the L0 layer are triggered. When the confidence level... It returns via a short circuit at L1 layer with a latency of less than 10 milliseconds, providing preventative maintenance decision support for telecommunications equipment operation and maintenance.
[0060] The L1 layer supports model inference via the ONNX Runtime. This feature is controlled by the AIOPS_ENABLE_ONNX conditional compilation macro and is disabled by default. When the ONNX inference engine is available, the L1 layer can load pre-trained ONNX models for more complex statistical analysis. When ONNX inference is unavailable, it automatically falls back to the 3-Sigma statistical detection method mentioned above, ensuring that the core functionality of the L1 layer is not affected.
[0061] The L1 layer is implemented in C language within a local Linux process on the EU, with no external network calls; all calculations are performed within the process. Through optimized algorithms (such as incrementally updating the mean and standard deviation, and avoiding repeated traversal of historical data), it is ensured that even when performing Z-Score and linear regression analysis on multiple indicators simultaneously, the total latency remains within 10 milliseconds, thus meeting the decision chain's expected target for the response time of the second layer.
[0062] After completing step S3, if a short-circuit return is not triggered, the system will proceed to step S4 (L2 semantic cache layer) with the context containing all results from L0 and L1 layers to continue execution.
[0063] S4: Execute the semantic caching layer, perform hash calculation on the query features in the context and match the historical analysis results in the cache. If the cache is hit and the similarity confidence reaches or exceeds the third threshold, the historical results in the cache are reused; otherwise, the cached query results are stored in the context and the deep analysis layer is executed.
[0064] In this embodiment, the L2 semantic cache layer is the third layer of a four-layer progressive decision chain, with a response latency of less than 5 milliseconds. The L2 layer is implemented based on semantic caching and runs locally in the EU. It reuses historical analysis results through fast hash matching, avoiding repeated analysis of the same or highly similar failure modes.
[0065] When the context structure carrying the analysis results of L0 and L1 layers is passed to L2 layer, L2 layer first reads the rule matching results (including the hit rule ID, device type, and anomaly indicator name) from the L0 layer result field of the context structure, and then reads the statistical analysis results (including Z-Score anomaly marker, trend slope, and coefficient of determination R) from the L1 layer result field. 2 It also reads the target device ID and time window information from the original analysis request fields.
[0066] Based on the above information, the L2 layer extracts query features, which include: device type (such as pRRU, RRU), abnormal indicator name (such as VSWR, PA temperature, SFP received optical power), abnormal value range (such as VSWR value, temperature value, optical power value), alarm type (such as antenna feeder fault, over-temperature alarm, optical module abnormality), and fault mode signature (a combined identifier formed by combining L0 rule ID and L1 statistical abnormality type).
[0067] Layer 2 uses the FNV-1a fast unencrypted hash algorithm to calculate the hash value of the query content. The FNV-1a algorithm uses a 64-bit hash output with an initial offset of 0xcbf29ce484222325.
[0068] The specific process of hash calculation is as follows: The extracted query features are concatenated into a byte sequence in a fixed order, including the device type string, the abnormal indicator name string, the alarm level string, and the integer value of the L0 rule ID; then, the byte sequence is iteratively hashed using the FNV-1a algorithm to finally generate a 64-bit unsigned integer as the query hash value H. fnv Unlike cryptographic hash algorithms (such as MD5 and SHA-1), FNV-1a is a non-cryptographic hash algorithm. It is fast and has extremely low overhead, making it suitable for cache lookup scenarios with millisecond-level latency requirements in embedded resource-constrained environments. Its hash collision probability is extremely low, which is sufficient to guarantee the uniqueness of the query results with a cache capacity of 256 entries.
[0069] The L2 layer searches for the semantic cache and matches H. fnv Historical analysis results of the matches. The semantic cache is stored in the EU local SQLite database ai_ops.db, with a cache capacity of 256 entries, and uses an LRU (Least Recently Used) eviction policy to manage the cache space.
[0070] When a new entry is added, if the cache is full (reaching 256 entries), the least recently used entry is evicted. Each cached entry contains the following fields: query hash value H. fnv (As a primary key index), historical analysis results summary (including conclusion text and suggested actions), and timestamp of the results generation. outliers And the type of equipment it belongs to.
[0071] L2 layer with H fnv The key is used to perform an exact lookup in the cache. If the lookup is successful (i.e., a cache entry with the same hash value exists), it is marked as a cache hit; if the lookup fails (i.e., no matching hash value exists), it is marked as a cache miss.
[0072] When a cache hit occurs, the L2 layer further compares the current query with historical cached entries to assess the applicability of historical results to the current query. The similarity comparison includes two dimensions: (1) Time window similarity The formula for comparing the proximity of the current query time to the time when the cached historical results were generated is as follows: ; in, This is the current query timestamp. For the historical results timestamps of cached entries, This is the maximum time window (default 300 seconds, consistent with the cache TTL). When the time difference is 0... When the time difference reaches or exceeds hour .
[0073] (2) Value range similarity The formula for comparing the current outlier value with the cached historical outlier values is as follows: ; in, This is an exception value in the current request. For historical outliers recorded in the cache entries, The maximum permissible deviation range (different indicators have different ranges) , such as VSWR The value is 1.0, representing the temperature. At 10℃, the optical power (5dBm). When the deviation is 0. =1.0; when the deviation reaches or exceeds hour =0.
[0074] Combining the two similarities mentioned above, the L2 layer calculates the cache matching confidence C. 2: ; in, and These are the weighting coefficients. ,default , Value range similarity is given higher weight because numerical similarity is more critical for fault mode matching than temporal similarity.
[0075] The calculated C2 is compared with the preset L2 short-circuit threshold T2 (default 0.85, i.e., 85%): If the cache hits and C2≥T2, the decision chain is short-circuited and terminated at the L2 layer. The L2 layer returns the historically cached analysis result (including the conclusion text and suggested operation) of the cache hit as the current analysis result, and writes the query hash value H fnv , cache hit flag, ID of the matched historical result, time similarity sim t , value range similarity sim v and confidence C2 into the L2 layer result field of the context structure, sets the short-circuit flag, and directly returns the L2 layer analysis result; If the cache misses, or the cache hits but C2<T2 (that is, the similarity between the historical result and the current context is insufficient), the L2 layer does not meet the short-circuit condition. At this time, the L2 layer writes the query hash value H fnv , cache hit / miss flag and confidence C2 into the L2 layer result field and the cumulative intermediate data field of the context structure, and the decision chain is upgraded to the L3 layer.
[0076] In the 5G / 4G small base station operation and maintenance scenario, the L2 layer semantic cache establishes a cache of historical analysis results for common telecommunication equipment failure modes, and quickly matches similar failure signatures through FNV-1a hashing.
[0077] For example, in a scenario where the cache hits for the VSWR antenna feeder failure mode, when a VSWR>3.0 antenna feeder failure occurs on the RU-3 device, after matching the TELECOM-001 rule at the L0 layer, the analysis context of this failure (device type pRRU, indicator name VSWR, abnormal value 3.2, antenna port 0, suggested operation) is calculated by FNV-1a hashing and stored in the semantic cache. When the RU-5 device subsequently has an antenna feeder failure with the same mode of VSWR>3.0, the L2 layer matches the historical cache entry of RU-3 through FNV-1a hashing, and the time window similarity sim t = 0.95 (the failure occurred within 5 minutes), the value range similarity sim v = 0.92 (the VSWR values are 3.2 and 3.15 respectively), the comprehensive matching confidence C2 = 0.4×0.95 + 0.6×0.92 = 0.93 ≥ 85%, so the L2 layer performs short-circuit return, reuses the historical analysis result of "antenna feeder failure, it is recommended to check the antenna connector and feeder", and the delay is less than 5 milliseconds.
[0078] In scenarios involving SFP optical module cache hits, when the RU-7 device experiences a reception anomaly with SFP Rx optical power <-25dBm, the context cache is analyzed and stored in the L2 layer after matching the TELECOM-003 rule at the L0 layer. If other RU devices (RU-8, RU-9) under the same HUB subsequently experience the same SFP optical power anomaly within 30 seconds, the L2 layer uses FNV-1a hashing to match the historical cache entry of RU-7. If the overall similarity is greater than 85%, the L2 layer short-circuit and returns, quickly reusing the analysis conclusion of "optical module reception anomaly, it is recommended to check the HUB-side optical port and fiber optic link."
[0079] For scenarios where the PA over-temperature fault mode cache is hit, when the RU-12 device generates an over-temperature alarm with a PA temperature > 85℃, after matching the TELECOM-002 rule at the L0 layer, the analysis context is stored in the semantic cache via FNV-1a hash. When other RU devices (RU-13, RU-14) at the same site successively experience PA temperature exceeding the limit within 60 seconds, the L2 layer cache is hit, the overall similarity is 0.88 ≥ 85%, and the historical analysis result "Power amplifier overheating, it is recommended to check the computer room air conditioning and cooling fans" is returned.
[0080] In the above scenario, the L2 layer reduces the identification latency of similar faults from 200-5000 milliseconds in the L3 layer to less than 5 milliseconds through FNV-1a hashing, while avoiding repeated LLM analysis for the same fault mode and reducing API call costs. In extreme scenarios where 16 RUs report the same type of fault simultaneously (such as a HUB-side optical port fault causing all lower-level RUs to report SFP optical power anomalies), the L2 layer semantic cache can complete the fault identification of all 16 RUs within one second, while calling the L3 layer LLM analysis one by one would take approximately 16 × 800 milliseconds ≈ 12.8 seconds.
[0081] Semantic caching employs management strategies, including: Capacity limit: The maximum number of cache entries is 256 to prevent unlimited growth from consuming the EU device's limited memory and storage resources; Eviction policy: The LRU (Least Recently Used) policy is adopted. When the cache reaches its capacity limit, the entry that has not been accessed for the longest time is evicted. Time to Live (TTL): Each cached entry has a TTL of 300 seconds. After the timeout, it will automatically expire and be removed from the cache, ensuring that historical results do not become outdated over time. Cache update: After the L3 layer completes in-depth analysis of the current query, if the confidence level of the analysis result is high (>90%), the hash value of the current query and the analysis result will be added to the cache as a new entry for reuse in subsequent similar queries.
[0082] The L2 semantic cache and the query hash cache in the MCP adapter share the FNV-1a hash algorithm and LRU eviction policy, but their caching granularity and purpose differ. The L2 semantic cache focuses on the semantic matching of the entire analysis request, with a coarser granularity, caching complete fault analysis conclusions (including root cause diagnosis and remediation suggestions); the cache in the MCP adapter focuses on the query results of LLM calls, with a finer granularity, caching the raw response content of the LLM API. They operate independently and complementaryly, jointly reducing the overall analysis latency of the decision chain.
[0083] All operations in Layer 2 are performed locally within the EU, including hash calculations, cache lookups, and similarity comparisons, without involving external network calls. FNV-1a hash calculations are completed in microseconds, SQLite database queries with index support have a lookup latency of less than 1 millisecond, and similarity comparisons are simple arithmetic operations with negligible latency. Through these optimizations, the total latency of Layer 2 is controlled to within 5 milliseconds, meeting the decision chain's response time requirements for Layer 3.
[0084] After completing step S4, if a short-circuit return is not triggered, the system will proceed to step S5 (L3 LLM deep analysis layer) with the context containing all results from layers L0, L1, and L2.
[0085] S5: Execute the deep analysis layer, read the analysis results of each layer stored in the context to construct analysis prompt words, call an external large language model to perform deep reasoning, and obtain analysis results; wherein, before executing the deep analysis layer, if the consumed latency has reached or exceeded the preset total latency budget, the deep analysis is skipped and the statistical analysis layer is re-executed to obtain analysis results.
[0086] Execute the L3 LLM deep analysis layer, read the complete conclusions of L0, L1, and L2 in the context to construct prompt words, and call the external large model API through the MCP adapter to obtain the analysis results; and before executing S5, if the accumulated consumption delay has caused the remaining budget to be less than 200ms, skip S5, downgrade and repeat the fast statistics of S3 and return.
[0087] In this embodiment, the L3 LLM deep analysis layer is the highest layer in the four-layer progressive decision chain, and also the most powerful but with the highest latency, ranging from 200 milliseconds to 5000 milliseconds. The L3 layer calls the external DeepSeek large language model API through the MCP adapter to perform deep inference analysis for complex scenarios where the L0, L1, and L2 layers cannot provide high-confidence conclusions.
[0088] When the context structure carrying the analysis results of layers L0, L1, and L2 is passed to layer L3, layer L3 first reads all existing analysis conclusions from the result fields of each layer in the context structure, including: rule matching results (hit rule ID, match degree, severity score) and L0 confidence score from the L0 layer result fields; and 3-Sigma anomaly detection results (mean μ, standard deviation σ, anomaly marker) and linear regression trend analysis results (slope a, R²) from the L1 layer result fields. 2 The L2 layer results field contains the following data: coefficient of determination, multi-index correlation coefficient matrix, and L1 confidence score; query hash value, cache hit / miss flag, historical matching result ID, time similarity, and L2 confidence score; and request type, target device ID, time window, and list of analysis indicators.
[0089] Layer 3 (L3) obtains device topology information based on the `ai_topology_deps` topology dependency table, including the connection relationships between the current device and other devices, and hierarchical relationships (such as the HUB to which the RU belongs, and the EU to which the HUB belongs). In cross-device analysis scenarios involving multiple devices simultaneously experiencing anomalies, topology information helps Layer 3 determine the anomaly propagation path and root cause localization.
[0090] Layer L3 selects a matching analysis template from a pre-built analysis template library. Template types include: anomaly_analysis (anomaly analysis template): Input anomaly indicators and time windows, output anomaly classification, severity assessment and recommendations; trend_forecast (Trend Forecast Template): Input historical time series data, output future trend prediction and inflection point prediction; log_analysis (log analysis template): Input log fragments and context, output anomaly pattern identification and suggestions; health_assessment (health assessment template): Input indicators from four dimensions: CPU, memory, storage, and temperature, and output a comprehensive health score and risk warning.
[0091] Template selection is based on the request type field in the context: if the request type is ANOMALY, select the anomaly_analysis template; if it is TREND, select the trend_forecast template; and if it is HEALTH, select the health_assessment template.
[0092] Layer L3 combines the complete conclusions from L0, L1, and L2 with the selected template to construct LLM analysis prompts. These prompts include: a list of outliers and their current values, historical mean, standard deviation, and Z-score; the rule ID, alarm level, and severity score matched by layer L0; and the trend direction (slope a) and goodness-of-fit (R²) detected by layer L1. 2 (and prediction conclusions); strong correlations in the multi-index correlation coefficient matrix where |r|>0.5; topological dependencies of equipment; and pre-set analysis prompts.
[0093] Through the above construction method, the L3 layer provides sufficient contextual information for the LLM, enabling in-depth reasoning based on existing analysis conclusions, rather than starting the analysis from scratch.
[0094] Layer 3 initiates HTTPS requests to the DeepSeek API through the MCP adapter, invoking the deepseek-chat model. The MCP adapter supports multiple LLM providers, including DeepSeek, OpenAI, Ollama, and Custom Providers, and performs failover in priority order. The MCP adapter provides protection mechanisms, including: (1) Token bucket rate limiting (rate limiter). The token bucket has a capacity of 60 and a replenishment rate of 1 token / second. Before each API call, a token is obtained from the token bucket. If the bucket is empty, it waits for tokens to be replenished to prevent the LLM API from being over-called and causing the quota to be exhausted.
[0095] (2) Three-state fuse. The fuse has three states: CLOSED (closed), OPEN (open), and HALF_OPEN (half-open). In the CLOSED state, API calls are made normally. When the error rate exceeds 50%, the fuse switches to the OPEN state. In the OPEN state, all API calls are rejected, and an error is returned directly. After a 30-second cooldown, it enters the HALF_OPEN state, allowing a limited number of probe requests. If the success rate of the probe requests exceeds 80%, it returns to the CLOSED state; if it fails, it returns to the OPEN state.
[0096] The MCP adapter sends the constructed prompt to the DeepSeek API endpoint via an HTTPS POST request, with a timeout of 30 seconds.
[0097] After receiving the analysis results returned by the LLM, the L3 layer performs post-processing operations, including: formatting the JSON string returned by the LLM, such as fixing missing quotes, extra or missing commas, and adding missing parentheses to ensure that the output is in a valid JSON format; desensitizing sensitive fields in the analysis results, such as replacing IP addresses with placeholders, replacing password fields with [REDACTED], and replacing the last 4 digits of the device serial number; when the analysis results returned by the LLM exceed 16384 characters, a [TRUNCATED] flag is added at the excess position and the result is truncated to prevent excessively long results from consuming the limited memory resources of the EU device.
[0098] After post-processing, the L3 layer writes the LLM analysis conclusions into the L3 layer results field of the context structure; it writes the complete analysis results into the database (to the ai_analysis_results table) and pushes them to the front-end console via the WebSocket protocol for real-time viewing by operations and maintenance personnel.
[0099] When an LLM call fails or times out (e.g., three consecutive DeepSeek API calls time out, with each timeout exceeding 30 seconds; or the MCP adapter vendor health check state machine determines FAILED; or the EU device network connection is lost), the L3 layer automatically triggers a degradation process: downgrading to the L1 layer to re-execute fast statistical analysis and skipping the deep analysis of the L3 layer. The degradation and recovery process is transparent to the upper-layer business logic, and degradation notifications are pushed to the front-end console via WebSocket. During degradation, a liveness probe request is sent to the DeepSeek API every 60 seconds, and the L3 layer is automatically restored upon successful liveness probe.
[0100] In 5G / 4G small cell operation and maintenance scenarios, L3 layer DeepSeek LLM deep analysis is mainly aimed at complex telecom fault scenarios where L0, L1, and L2 layers cannot provide high-confidence conclusions.
[0101] For example, in a multi-indicator-related root cause analysis scenario for telecommunications faults, when the RU-12 device simultaneously exhibits three abnormal indicators—VSWR > 3.0 (antenna feeder fault), PA temperature > 85℃ (power amplifier overheating), and transmit power decrease by 4dB (abnormal transmit power)—each rule in the L0 layer can be matched independently, but the confidence level decreases due to the concurrency of multiple indicators (the confidence level of each single indicator is <95%). The L1 layer multi-indicator correlation coefficient analysis shows a correlation coefficient r = 0.82 (strong positive correlation) between VSWR and PA temperature, and a correlation coefficient r = -0.75 (strong negative correlation) between VSWR and transmit power, but the causal relationship cannot be determined. The L2 layer cache has no similar historical patterns. The decision chain is upgraded to the L3 layer, which receives the complete context (L0 rule matching results, L1 statistics and correlation analysis, topology information), calls the DeepSeek API through the MCP adapter, and constructs the analysis prompt by selecting the anomaly_analysis template. The LLM analysis concluded: "The increased VSWR at the antenna port of RU-12 caused a power amplifier load mismatch, leading to increased reflected power and a rise in power amplifier temperature. The power amplifier protection mechanism automatically reduced the transmit power after activation. The root cause is a fault in the antenna feeder system (loose antenna connector or water ingress into the feeder). Recommendations: 1) Check the tightness of the coaxial connector at antenna port 0; 2) Check the feeder for water ingress or oxidation; 3) Repair the antenna feeder system and recalibrate the power amplifier output power." The total delay is approximately 800 milliseconds, and the decision path is L0→L1→L2→L3→return.
[0102] For cross-RU telecommunications fault correlation analysis scenarios, when the four pRRU devices (RU-5, RU-6, RU-7, RU-8) connected to the HUB simultaneously experience SFP optical module Rx optical power <-25dBm and CPRI BER>1×10⁻⁶ within 30 seconds, -9When an anomalies occur, the TELECOM-003 and TELECOM-004 rules for each RU in L0 layer are triggered independently, but it cannot explain why all four RUs are anomaly simultaneously. The L2 layer cache hit rate is low (each RU has a different combination of failures), so an upgrade to L3 layer is implemented. L3 layer constructs a complete cross-RU context (L0 rule matching results for the four RUs, L1 statistical anomaly results, and HUB topology dependencies), and uses the MCP adapter to call the DeepSeekAPI for cross-RU correlation analysis. The LLM analysis concluded that: "The SFP optical module at HUB-side optical port 1 (Tx optical power -3dBm normal, but Rx optical power -28dBm abnormal) has a one-way reception fault, causing the four downstream pRRU devices to simultaneously experience abnormal optical power and increased bit error rate. The root cause is a fault in the SFP receiver at HUB-side optical port 1. Recommendations: 1) Replace the SFP optical module at HUB-side optical port 1; 2) Check if the HUB-side fiber optic connector is clean; 3) Verify the recovery of optical power and BER of the four pRRUs after replacement." The total delay is approximately 1200 milliseconds, and the decision path is L0→L1→L2→L3→return.
[0103] In both scenarios mentioned above, the L3 layer DeepSeek LLM receives the complete context from the L0, L1, and L2 layers (including rule matching results, statistical anomaly detection, trend analysis, multi-indicator correlation coefficients, and topological dependencies), avoiding starting the analysis from scratch. It completes multi-indicator correlation root cause reasoning and cross-RU fault correlation analysis within 800 to 1200 milliseconds, providing operation and maintenance personnel with actionable root cause diagnosis and remediation suggestions.
[0104] Before executing step S5, the L3 layer first checks whether the current accumulated latency is close to the total latency budget. The total latency budget B_total is 5000 milliseconds by default, and the expected latency for each layer is: less than 1 millisecond for L0, less than 10 milliseconds for L1, and less than 5 milliseconds for L2.
[0105] Calculate the remaining budget: B_remain = B_total - (D_0 + D_1 + D_2), where D_0, D_1, and D_2 are the actual latency consumed by layers L0, L1, and L2, respectively. If B_remain ≥ 200 milliseconds (minimum execution time for layer L3), then proceed normally to layer L3 for LLM deep analysis; if B_remain < 200 milliseconds, then the remaining budget for layer L3 is insufficient, layer L3 is skipped, and the process is downgraded to layer L1 for fast statistical analysis, with the reason for the downgrade recorded in the context.
[0106] The latency budget can be dynamically adjusted according to the priority of the analysis request: B_total = 8000 milliseconds for CRITICAL priority requests, B_total = 5000 milliseconds for HIGH priority requests, B_total = 3000 milliseconds for MEDIUM priority requests, and B_total = 1000 milliseconds for LOW priority requests (not entering L3 layer).
[0107] S6: Return the final analysis results and write the execution path identifier of the decision chain, the confidence level of each layer, and the delay data to the database for persistent storage.
[0108] In this embodiment, when the decision chain terminates due to a short circuit at layer L0, layer L1 or layer L2 and jumps to this step, or enters this step after completing in-depth analysis at layer L3, the final analysis result is returned to the caller, and the complete decision chain execution record is written to the database for post-event auditing, strategy optimization and operation and maintenance traceability.
[0109] First, the final analysis results are read from the context structure. The source of the final analysis results varies depending on the actual execution path of the decision chain: if the decision chain terminates at layer L0 (short-circuit), the final analysis results come from the rule matching conclusions in the L0 layer result field of the context structure, including the matched rule ID, alarm level, severity score, suggested action, and confidence score; if the decision chain terminates at layer L1 (short-circuit), the final analysis results come from the statistical analysis conclusions in the L1 layer result field of the context structure, including anomaly detection results, trend prediction conclusions, and confidence scores; if the decision chain terminates at layer L2 (short-circuit), the final analysis results come from the cache reuse conclusions in the L2 layer result field of the context structure, including the reused historical analysis conclusion text and confidence scores; if the decision chain completes in-depth analysis at layer L3 and then proceeds to step S6, the final analysis results come from the LLM analysis conclusions in the L3 layer result field of the context structure; if step S5 triggers a degradation due to insufficient latency budget and repeats the L1 layer analysis before proceeding to step S6, the final analysis results come from the re-executed L1 layer analysis conclusions.
[0110] The final analysis results and auxiliary information are assembled into a complete return payload. The auxiliary information includes: a unique request identifier (req_id, taken from the context header); the target device ID (taken from the original analysis request field); and a decision path identifier (decision_path), indicating the actual sequence of levels traversed in this decision chain, with possible values including "L0→return", "L0→L1→return", "L0→L1→L2→return", "L0→L1→L2→L3→return" or "L0→L1→L2→degraded L1→return"; the final return... The final layer (final_layer) can be L0, L1, L2, or L3; the final confidence score (final_confidence) is the confidence value of the final returned layer; the total latency (total_latency) is the cumulative time from receiving the request in step S1 to returning the result in step S6, in milliseconds; whether to trigger degradation is a boolean value, indicating whether this request was triggered by the unavailability of the L3 layer or insufficient budget; the cumulative intermediate analysis data (intermediate_data) contains all intermediate calculation results from the L0 layer to the final returned layer, for use in post-event auditing.
[0111] The returned payload is sent back to the caller (such as the upper-layer network management system or local console) through the inter-process communication interface on the EU side, and is also pushed to the front-end console for operation and maintenance personnel to view in real time via the WebSocket protocol.
[0112] The complete execution record of this decision chain is written to the database in JSON format. The core fields included in the execution record are shown in Table 1:
[0113] The aforementioned JSON record fully preserves the data from the moment the request is initiated to the end of the analysis, ensuring that the entire decision-making process is traceable and auditable.
[0114] Persistent records are written to the ai_analysis_results table in the local SQLite database ai_ops.db in the EU; this table contains the following: id (primary key, auto-incrementing integer), req_id (string, unique index), timestamp (integer, indexed to support time range queries), device_id (string, indexed to support device-level queries), decision_path (string), result_json (JSON format for storing the complete record), and created_at (record creation time).
[0115] The database uses the following configuration to balance write performance and data security: WAL (Write-Ahead Logging) mode: Allows concurrent read and write operations, improving write throughput; synchronous=NORMAL: Reduce disk synchronization overhead while ensuring data security; journal_size_limit=4MB: Limits the size of the WAL log file to prevent the log from growing indefinitely; cache_size=2000 pages: Caches 2000 database pages (4KB per page by default), approximately 8MB of cache, adapted to the limited memory of EU devices.
[0116] Each write operation is a single INSERT operation with a latency of less than 10 milliseconds, which does not affect the overall response time of the decision chain.
[0117] To prevent the database from growing uncontrollably and consuming the limited storage resources of the EU device (the storage space of the EU device is typically 256MB to 512MB), the data cleanup mechanism is as follows: each analysis record is retained for 7 days; a cleanup task is executed when the system starts or at 2:00 AM every day to delete records whose created_at field is older than 7 days; the cleanup operation is performed in batch deletion mode (1000 records are deleted each time, and the process is repeated until all expired records are cleared) to avoid locking the database in a single large transaction.
[0118] When the decision chain terminates due to a short circuit at a certain layer, the output not only includes the analysis conclusions of that layer but also all accumulated intermediate analysis data. Specifically, the intermediate_data field in the returned payload stores the complete intermediate data extracted from the context structure, including: analysis results from all layers from L0 to the final returned layer (e.g., when L2 is short-circuited, it also includes intermediate results from L0 and L1 layers); confidence scores for each layer; actual latency consumed by each layer; and short-circuit layer identifier.
[0119] Through the above mechanism, even if the decision chain is short-circuited and terminated at a lower level, the operation and maintenance personnel can still obtain complete analysis process data by querying the database, which can be used for post-event auditing and decision chain strategy optimization.
[0120] When step S5 triggers a degradation due to insufficient latency budget or LLM unavailability, mark `degraded=true` in the persistent record and record the specific reason in the `degraded_reason` field (e.g., "L3 remaining budget less than 200ms" or "DeepSeek API timed out 3 times consecutively"). The `decision_path` field records "L0→L1→L2→Degrade to L1→return", the `final_layer` field records L1 (the result returned by the L1 layer after degradation), and the `final_confidence` field records the confidence value after re-performing statistical analysis at the L1 layer. These markings facilitate subsequent statistical analysis of degradation trigger frequency and cause distribution, providing data support for latency budget optimization and LLM availability monitoring.
[0121] Example 2
[0122] This embodiment exemplifies a multi-layered decision chain return system for embedded operations and maintenance. The system runs in the embedded Linux environment of an EU (Edge Computing Unit) device and consists of six functional modules. Data is transferred between modules via context structures, and the call relationships between modules form a progressive decision chain from rapid response to in-depth analysis. Specifically, it includes: a request receiving module, a rule engine module, a statistical analysis module, a semantic caching module, a deep analysis module, and an output persistence module.
[0123] The request receiving module receives operation and maintenance analysis requests from RU devices, EU local scheduled tasks, or upper-layer network management systems. It parses parameters such as device identifier, time window, indicator list, and request type from the request and creates a context structure in the heap area of the EU's local Linux process to store request parameters and intermediate analysis results from each layer. Specifically, the request receiving module employs a pooled memory management strategy. At system startup, a fixed number of context structure pools (with capacity matching the number of RUs managed by the EU) are pre-allocated. When a request is received, a structure is retrieved from the idle queue and initialized. Request parameters are written to the corresponding storage location in the context, and the hierarchical cursor and delay budget counter are initialized simultaneously, enabling the context structure to be passed to subsequent modules with zero-copy functionality.
[0124] The rules engine module, as the first-level processing unit in the decision chain, embeds a Lua 5.4.6 script engine and runs within the EU's local Linux process. This module reads the target device's metric data and device type information from the context, extracts key request features, and matches them against pre-built Lua rule scripts. These rule scripts are categorized and stored in the EU's local file system's rules directory, including system rules, business rules, hardware rules, and message service rules. When a request feature matches a rule's matching pattern, the module calculates the matching confidence based on the rule's severity score and compares it to a preset high-confidence threshold. If the threshold is reached, the matching conclusion is returned directly as the final analysis result, and the decision path is marked as a short-circuit return. If the threshold is not reached or no rule is matched, the matching result is written to the context, triggering the statistical analysis module to continue processing. The rules engine module runs in a sandbox environment with security limitations on memory, instruction count, and execution time.
[0125] The statistical analysis module, as the second-level processing unit in the decision chain, runs as a statistical algorithm engine implemented in C within the EU's local Linux process. This module reads the L0 layer matching results and raw indicator data from the context, and performs sliding window-based statistical standardization detection and linear regression trend analysis on the indicators of interest. Using historical data points within the sliding window as samples, the module calculates the mean and standard deviation, performs standardization on the currently collected values, and marks statistical anomalies when the standardized values exceed a preset threshold. It then performs least-squares linear fitting on the historical time-series data, calculating the trend slope and coefficient of determination. The module calculates the confidence level of the statistical analysis by weighting the anomaly strength and the trend fit goodness of fit, and compares it with a preset medium-to-high confidence threshold. If the threshold is reached, the module directly returns the statistical analysis conclusion (including anomaly marking, trend direction, and prediction information); otherwise, the statistical results are written to the context and the semantic caching module is triggered to continue processing.
[0126] The semantic caching module, as the third-level processing unit in the decision chain, runs locally on the EU and uses an SQLite database for cache storage. This module reads the analysis results from layers L0 and L1 of the context, extracting query features such as device type, abnormal indicator name, alarm type, and fault mode signature, and generates a query hash value using a non-encrypted hash algorithm. The module performs a precise lookup in the cache table using the hash value as the key. If a match is found, it further calculates the temporal and numerical similarity between the current query and historical cache entries, weighting the similarity confidence score. If the similarity reaches a preset threshold, it directly reuses historical analysis results from the cache (including conclusion text and suggested actions), keeping the recognition latency to the millisecond level. If there is no match or insufficient similarity, the query result is written to the context and the deep analysis module is triggered to continue processing. The cache uses a least recently used eviction policy to manage capacity, and cache entries have a preset lifespan, automatically becoming invalid after expiration.
[0127] The deep analysis module, as the highest-level processing unit in the decision chain, calls the external large language model API for deep inference through the MCP adapter. This module reads all existing analysis conclusions from layers L0, L1, and L2 (including rule matching results, statistical anomaly detection, trend analysis, multi-indicator correlation coefficients, topological dependencies, etc.) from the context. Based on the request type, it selects the corresponding analysis template, combines the complete context with the template to construct analysis prompts, and sends these prompts to the large language model server via HTTPS. Before executing deep analysis, this module first checks whether the currently consumed latency has reached or exceeded the preset total latency budget. If the remaining latency budget is insufficient to complete a deep analysis (less than the preset minimum execution time threshold), the deep analysis is skipped, and the statistical analysis module is degraded to re-execute the statistical analysis and return the results. The reason for the degradation is marked in the context. If the budget is sufficient, the large language model API is called normally. After receiving the returned results, post-processing operations such as format repair, sensitive information desensitization, and length truncation are performed, and the analysis conclusions are written into the context. The MCP adapter has a built-in current limiter and fuse: the current limiter uses the token bucket algorithm to control the frequency of API calls and prevent quota exhaustion; the fuse has three states: closed, open and half open, and automatically switches states when the error rate exceeds a preset threshold to achieve fault isolation and automatic recovery.
[0128] The output persistence module returns the final analysis results and writes the complete execution record of this decision chain to persistent storage in the database. This module reads the final analysis conclusion from the context (potentially from any of the L0, L1, L2, or L3 layers), assembles the conclusion text, decision path identifiers, confidence scores for each layer, latency data for each layer, and degradation flags into a complete output payload, and returns it to the caller via an inter-process communication interface. This module writes all the above data in JSON format to the analysis results table of the local SQLite database in the EU. The database uses WAL mode to support concurrent read and write operations, and each record is automatically cleaned up after a preset time. When the decision chain terminates due to a short circuit at a lower layer, the persistent record output by this module still contains all accumulated intermediate analysis data up to the short-circuit layer, for use by operations personnel for post-event auditing and strategy optimization.
[0129] The six modules described above are sequentially linked to form a complete decision chain: after the request receiving module creates the context, it sequentially calls the rule engine module, statistical analysis module, semantic caching module, and deep analysis module (skipping when the latency budget is insufficient or the external API is unavailable); any module short-circuites and returns when its confidence level reaches the corresponding threshold, and the output persistence module completes the output and record storage of the results. There is no inter-process communication between the modules; all modules pass data within the same process through the context structure, ensuring that the overall latency of the decision chain is predictable and controllable.
[0130] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims. It should be understood that the invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for layer-by-layer return of a multi-level decision chain in embedded operations and maintenance, characterized in that, The method includes: Receive operation and maintenance analysis requests and create a context for storing request parameters and intermediate results of analysis at each level; The execution rule engine layer matches the request features with preset rules. If the confidence level of the matching result reaches or exceeds the first threshold, the matching result is returned directly; otherwise, the matching result is stored in the context and the statistical analysis layer is executed. The statistical analysis layer is executed to perform statistical anomaly detection and trend analysis on the indicators to be analyzed in the context. If the confidence level of the statistical analysis result reaches or exceeds the second threshold, the statistical analysis result is returned directly; otherwise, the statistical analysis result is stored in the context and the semantic caching layer is executed. The semantic caching layer is executed to perform hash calculation on the query features in the context and match them with historical analysis results in the cache. If the cache is hit and the similarity confidence reaches or exceeds the third threshold, the historical results in the cache are reused; otherwise, the cached query results are stored in the context and the deep analysis layer is executed. The deep analysis layer is executed, and the analysis results of each layer stored in the context are read to construct analysis prompt words. An external large language model is called to perform deep reasoning and obtain analysis results. If the latency consumed before executing the deep analysis layer reaches or exceeds the preset total latency budget, the deep analysis is skipped and the statistical analysis layer is re-executed to obtain analysis results. The final analysis results are returned, and the execution path identifier of the decision chain, the confidence level of each layer, and the latency data are written to the database for persistent storage.
2. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The confidence level of the rule engine layer is calculated based on the severity score of the matching rule, and the first threshold is a preset high confidence threshold.
3. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The statistical analysis layer performs statistical standardization detection and linear regression trend analysis based on a sliding window on the indicator to be analyzed, and calculates the confidence level of the statistical analysis results by weighting the anomaly detection intensity and the trend fit goodness.
4. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The semantic caching layer uses a non-cryptographic hash algorithm to generate the hash value of the query features. The similarity confidence is calculated based on the weighted average of the temporal and numerical proximity between the current query and the cached historical results. The cache adopts a least recently used eviction policy, and cache entries have a preset lifespan.
5. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The deep analysis layer calls the external large language model application interface through an adapter. The adapter includes a rate limiter and a circuit breaker. The rate limiter is used to control the calling frequency. The circuit breaker switches to the open state when the error rate exceeds a preset threshold, and enters the half-open state after cooling. When the detection success rate exceeds the preset threshold in the half-open state, it returns to the closed state.
6. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The preset total latency budget is dynamically adjusted according to the priority of the analysis request; when the remaining latency budget before entering the deep analysis layer is less than the preset minimum execution time threshold, the deep analysis is skipped and the statistical analysis layer is re-executed.
7. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The preset rules include dedicated rules for communication equipment hardware specifications. These dedicated rules include one or more of the following: antenna port VSWR detection rules, power amplifier temperature detection rules, optical module received optical power detection rules, fronthaul link bit error rate detection rules, transmit power deviation detection rules, and reference signal received power detection rules.
8. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The rule engine layer runs in a script engine sandbox environment, which has resource consumption limits and prohibits the execution of system commands, file operations, and dynamic loading of unauthorized code.
9. The embedded operation and maintenance multi-layer decision chain step-by-step return method according to claim 1, characterized in that, The method further includes: when the external large language model fails to be called continuously or the network connection is unavailable, triggering a degradation mode to reduce the decision chain to a three-layer local processing mode of rule engine layer, statistical analysis layer and semantic cache layer; during the degradation period, sending a liveness detection request to the external large language model at a preset interval, and automatically restoring the deep analysis layer after the liveness detection is successful.
10. A multi-layer decision chain back-to-back system for embedded operation and maintenance, used to implement the method as described in any one of claims 1-9, characterized in that, include: The request receiving module is used to receive operation and maintenance analysis requests and create a context for storing request parameters and intermediate results of analysis at each level. The rules engine module is used to match request features with preset rules. If the confidence level of the matching result reaches or exceeds the first threshold, the matching result is returned directly; otherwise, the matching result is stored in the context. The statistical analysis module is used to perform statistical anomaly detection and trend analysis on the indicators to be analyzed in the context. If the confidence level of the statistical analysis result reaches or exceeds the second threshold, the statistical analysis result is returned directly; otherwise, the statistical analysis result is stored in the context. The semantic caching module is used to perform hash calculations on the query features in the context and match historical analysis results in the cache. If the cache hits and the similarity confidence reaches or exceeds the third threshold, the historical results in the cache are reused; otherwise, the cached query results are stored in the context. The deep analysis module is used to read the analysis results stored in the context to construct analysis prompt words, and call an external large language model to perform deep reasoning to obtain analysis results. If the latency consumed before the deep analysis is executed reaches or exceeds the preset total latency budget, the deep analysis is skipped and the statistical analysis module is triggered to re-execute the statistical analysis to obtain analysis results. The output persistence module is used to return the final analysis results and write the execution path identifier of the decision chain, the confidence level of each layer, and the latency data into the database for persistent storage.