Optical network unit fault diagnosis method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610997792.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-10-02
AI Technical Summary
当设备发生故障导致网络中断后,运维人员才介入处理,这种方式存在响应滞后、用户体验差、运维成本高等问题
[0009]上述光网络单元故障诊断方法、装置、设备及存储介质,在云端为每一台光网络单元构建高保真数字孪生体,在虚拟空间完成物理设备全维度运行映射,完整复现光网络单元运行全过程,得到运行数据,将运行数据输入故障诊断模型,区别于传统固定阈值告警、人工规则匹配的诊断方式,能够基于设备长期运行趋势识别渐进式隐性故障,实现故障风险提前预判,将运维模式从故障发生后被动抢修转变为故障爆发前主动预防,大幅降低用户网络中断概率。同时故障诊断模型自动识别多类故障风险并匹配专属故障处理策略,自动下发执行对应的处置操作,无需运维人员人工分析故障、现场处置,实现故障诊断、策略下发、自动修复的全流程闭环。
Smart Images

Figure CN122870271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis, and in particular to a method, apparatus, device, and storage medium for diagnosing faults in optical network units. Background Technology
[0002] With the rapid popularization of Fiber to the Home (FTTH) networks, the number of optical network unit gateways, as key access devices on the user side, has exploded. Statistics show that by 2026, the number of FTTH users worldwide had exceeded 1 billion, and the number of deployed optical network unit devices had reached billions. Operating in users' home environments for extended periods, optical network unit gateways face multiple challenges, including temperature fluctuations, unstable power supplies, optical module aging, and storage media wear, resulting in a persistently high equipment failure rate.
[0003] Traditional optical network unit (ONU) maintenance primarily relies on user reports of faults or periodic inspections, representing a typical reactive approach. Maintenance personnel only intervene after a device malfunctions and causes a network outage, resulting in delayed response, poor user experience, and high maintenance costs. Furthermore, the widespread distribution and large number of ONU devices make traditional manual inspections insufficient for comprehensive coverage and prevent the timely detection of potential faults. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for diagnosing optical network unit (ONU) faults, thereby overcoming the shortcomings of existing ONU maintenance technologies and enabling early prediction of ONU fault risks.
[0005] In a first aspect, embodiments of the present invention provide a method for diagnosing optical network unit faults, comprising: Create a digital twin corresponding to the optical network unit; The operation process of an optical network unit is simulated using a digital twin to obtain operational data; The running data is input into the pre-trained fault diagnosis model, and the fault diagnosis model is used to diagnose different types of fault risks to obtain fault diagnosis results. Obtain the fault handling strategy corresponding to the fault diagnosis result, and execute the fault handling strategy.
[0006] Secondly, embodiments of the present invention provide an optical network unit fault diagnosis device, comprising: A module is created to generate digital twins corresponding to optical network units; The simulation module is used to simulate the operation of optical network units using a digital twin to obtain operational data; The fault diagnosis module is used to input the running data into the pre-trained fault diagnosis model, and use the fault diagnosis model to diagnose different types of fault risks and obtain fault diagnosis results. The acquisition and execution module is used to acquire the fault handling strategy corresponding to the fault diagnosis result and execute the fault handling strategy.
[0007] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described optical network unit fault diagnosis method.
[0008] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described optical network unit fault diagnosis method.
[0009] The aforementioned optical network unit (ONU) fault diagnosis methods, devices, equipment, and storage media construct a high-fidelity digital twin for each ONU in the cloud. This virtual system maps the physical device's operation across all dimensions, fully reproducing the entire ONU's operational process and generating operational data. This data is then input into the fault diagnosis model. Unlike traditional methods relying on fixed threshold alarms and manual rule matching, this approach identifies progressive, latent faults based on long-term operational trends, enabling early prediction of fault risks. This transforms the maintenance model from reactive repair after a fault occurs to proactive prevention before a fault occurs, significantly reducing the probability of network outages for users. Simultaneously, the fault diagnosis model automatically identifies multiple fault risks and matches them with specific fault handling strategies, automatically issuing and executing corresponding actions. This eliminates the need for maintenance personnel to manually analyze faults or handle them on-site, achieving a closed-loop process for fault diagnosis, strategy issuance, and automatic repair. Attached Figure Description
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] Figure 1 This is a schematic diagram of an application environment for an optical network unit fault diagnosis method according to an embodiment of the present invention; Figure 2 This is a flowchart of a method for diagnosing optical network unit faults according to an embodiment of the present invention; Figure 3 This is another flowchart of a method for diagnosing optical network unit faults in one embodiment of the present invention; Figure 4 yes Figure 2 A flowchart of step S103; Figure 5 yes Figure 2 A flowchart of step S104; Figure 6 This is a flowchart of an optical network unit fault diagnosis device according to an embodiment of the present invention; Figure 7 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0012] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0013] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0014] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0015] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0016] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0017] References to "one embodiment" or "some embodiments" as used in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0018] Currently, there are three main mature and implemented technical solutions for fault monitoring, diagnosis, and maintenance of optical network units in fiber optic access networks: network management systems based on SNMP / TR069 protocols, fault diagnosis systems based on manual rules, and manual preventive maintenance solutions with fixed cycles. The specific implementation logic of these three solutions is as follows.
[0019] Network management system based on SNMP / TR069 protocol: It relies on standardized operation and maintenance communication protocol to realize device data acquisition, and the core adopts SNMP simple network management protocol or TR069 remote device management protocol to periodically complete the data interaction of optical network units.
[0020] Technical principle: The operation and maintenance platform sends data query commands to all optical network units according to a preset polling cycle. After receiving the command, the optical network unit sends back basic operating indicators, mainly including shallow status information such as online / offline status, transmit / receive power, and link error rate. The network management platform has a built-in unified judgment threshold. It compares the real-time collected indicators with the fixed threshold. Once the indicator exceeds the threshold boundary, the platform immediately generates an alarm message and pushes it to the operation and maintenance personnel. The operation and maintenance personnel can then use the alarm information to conduct remote troubleshooting or dispatch offline maintenance work orders.
[0021] Implementation method: The operation and maintenance platform, as the server, initiates polling requests periodically, and multiple optical network units, as terminals, passively respond and report instantaneous single-point data; the platform only makes independent threshold judgments on single indicators, does not associate them with historical operation records, and only triggers the alarm process when the indicator exceeds the limit. Subsequent fault handling relies entirely on manual intervention.
[0022] A fault diagnosis system based on an expert rule base: This system establishes a rule base mapping fault characteristics to their causes, relying on historical fault cases accumulated by industry operations and maintenance experts. When equipment malfunctions, the system locates and diagnoses the fault based on rule matching.
[0023] Technical principle: Operation and maintenance experts review a large number of past fault handling cases, extract the correspondence between abnormal indicators and fault types, solidify them into judgment rules, and establish a static mapping relationship between fault characteristics and fault causes; when the network management system collects abnormal data from the device, the system traverses all preset rules to complete the matching, and outputs the corresponding fault diagnosis conclusion and standardized handling suggestions based on the matched rules.
[0024] Implementation: The system pre-configures a massive number of if-then decision rules. All diagnostic logic relies entirely on manually written rules. When monitoring data meets the rule conditions, the system automatically provides diagnostic conclusions and handling suggestions. It lacks self-learning and trend inference capabilities, and can only identify fault scenarios defined within the rules.
[0025] Fixed-cycle manual preventive maintenance plan: Adopting a standardized batch inspection mode, offline manual operation and maintenance of optical network units in the region is carried out by setting a unified time cycle to proactively reduce the probability of failure.
[0026] Technical principle: The operator's operation and maintenance team conducts preventive inspections and maintenance on optical network units at fixed intervals, including firmware upgrades, configuration optimization, hardware testing, etc., in order to reduce the probability of failure.
[0027] Implementation: The operations and maintenance team develops an annual maintenance plan and conducts regular inspections of optical network unit equipment by region and batch. Inspections include equipment status checks, log analysis, performance testing, etc., with any issues addressed promptly.
[0028] The above three types of traditional optical network unit operation and maintenance diagnosis solutions all have significant shortcomings in the large-scale deployment of FTTH fiber optic networks, and cannot meet the business needs of refined, predictive, and automated operation and maintenance of massive terminals. The specific defects of each solution are described below: The main disadvantages of traditional SNMP / TR069 network management systems: (1) Delayed response: Alarms are only generated after a fault occurs or when the indicator exceeds the limit, making it impossible to provide early warning of potential faults. (2) Difficulty in setting thresholds: Fixed thresholds are difficult to adapt to the differences between different devices and environments. Thresholds that are too high will lead to missed alarms, while thresholds that are too low will lead to a flood of false alarms. (3) Information silos: Each monitoring indicator is independent of the others, making it impossible to comprehensively analyze the overall health status of the equipment and making it difficult to detect complex faults that are related to multiple factors. (4) Lack of predictive ability: It is impossible to predict the future status of the equipment based on historical data, making it impossible to achieve true preventive maintenance.
[0029] The main drawbacks of rule-based fault diagnosis systems are: (1) Incomplete rule coverage: Fault modes are diverse and constantly evolving, making it difficult for manual rules to cover all fault scenarios and new fault types cannot be identified. (2) High maintenance costs: The rule base requires continuous maintenance and updates by experts, and the workload of rule maintenance is huge as the number of equipment models increases and fault modes evolve. (3) Lack of adaptive capability: The rules are based on historical experience and cannot automatically adapt to dynamic factors such as equipment aging and environmental changes, resulting in a decrease in diagnostic accuracy over time. (4) Inability to predict unknown faults: The system cannot provide early warnings for fault types that have not occurred before.
[0030] The main disadvantages of regular preventative maintenance programs: Serious waste of resources: A large number of normal equipment undergoes unnecessary inspections, wasting manpower and resources; while equipment that is about to fail may fail between two inspections. (2) Poor timeliness: The maintenance cycle is fixed and cannot be dynamically adjusted according to the actual status of the equipment, making it difficult to intervene at the best time. (3) Limited coverage: Faced with a large number of optical network unit equipment, manual inspection can only cover a small part, and most of the equipment is out of control. (4) High cost: A large number of maintenance personnel and vehicles are required, and the maintenance cost increases linearly with the number of equipment, making it difficult to scale up.
[0031] In summary, existing technologies generally lack a unified data mapping carrier for all dimensions of equipment, making it impossible to accumulate time-series operational data throughout the entire equipment lifecycle. They also lack machine learning time-series prediction and multi-dimensional comprehensive evaluation capabilities, relying solely on instantaneous indicators, static rules, or manual offline troubleshooting for fault handling. This fails to achieve fully automated closed-loop operation and maintenance, making it difficult to meet the low-cost, high-efficiency, and high-accuracy early warning requirements for the operation and maintenance of massive optical network units. Therefore, this invention proposes an optical network unit fault diagnosis method to address the various shortcomings of existing technologies.
[0032] This invention provides a method for diagnosing optical network unit faults, which can be applied to, for example... Figure 1 The application environment is shown. Specifically, this fault diagnosis method is applied in a fault diagnosis system, which includes a client and a server. The client and server communicate via a network to create a digital twin, run a fault diagnosis model, output and execute fault handling strategies, etc. The client, also known as the user terminal, refers to the program that provides local services to the client, corresponding to the server. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0033] In one embodiment, please refer to Figure 2 A method for diagnosing faults in an optical network unit is provided, comprising the following steps: Step S101: Create a digital twin corresponding to the optical network unit; Step S102: Simulate the operation of the optical network unit using a digital twin to obtain operational data; Step S103: Input the running data into the pre-trained fault diagnosis model, use the fault diagnosis model to diagnose different types of fault risks, and obtain fault diagnosis results.
[0034] Step S104: Obtain the fault handling strategy corresponding to the fault diagnosis result and execute the fault handling strategy.
[0035] In this embodiment of the invention, an independent digital twin instance is created for each on-network optical network unit device in the cloud. The digital twin is a precise mapping of the physical optical network unit in virtual space, containing a multi-dimensional data model including the complete configuration information, real-time operating status, historical performance data, and environmental parameters of the optical network unit.
[0036] As an example, in step S101, when an optical network unit (ONU) goes online for the first time, the system independently constructs a dedicated digital twin for each ONU connected to the network in the cloud, initializes the basic information of the ONU, synchronizes data on device configuration, online status, and performance indicators across all dimensions, and builds a data model that maps the physical ONU to the virtual model in real time. The basic information includes static attributes such as device model, serial number, hardware version, firmware version, and geographical location.
[0037] The operational data is divided into two categories: one is the real indicator data uploaded by the physical optical network unit to the digital twin in real time; the other is the simulated operating condition data generated by the digital twin simulation. It uniformly includes flash memory operation data, optical module operation data, central processing unit or memory operation data, etc., which serve as the input basis for subsequent fault diagnosis models to perform fault analysis.
[0038] As an example, in step S102, the operating conditions of the optical network unit are simulated and reproduced using a digital twin, and future operating trends are predicted, generating complete operating data including real acquisition indicators and simulation prediction indicators.
[0039] Among them, the fault diagnosis result is a standardized conclusion output by the fault diagnosis model after processing and analyzing the operating data. It includes the fault type, fault risk level, expected fault occurrence time, and prediction confidence level, which is used to match subsequent fault handling strategies.
[0040] As an example, in step S103, the operating data output by the digital twin is used as input parameters to import into the fault diagnosis model. By using multiple built-in models, the hidden degradation trend of optical network units is identified, and multiple types of fault risks are determined respectively. The fault diagnosis results, including fault category, risk level, expected occurrence time of fault, and confidence level, are output.
[0041] Among them, the fault handling strategy is an automated handling solution that is pre-stored in the fault handling strategy library and matched one by one with various fault risks. It covers operation instructions that can be remotely issued and executed, such as configuration parameter tuning, traffic load transfer, and hardware parameter compensation.
[0042] As an example, in step S104, based on the fault type and risk level in the fault diagnosis results, the fault handling strategy library is retrieved, a suitable fault handling strategy is matched, and a strategy instruction is remotely sent to the optical network unit to automatically execute the fault handling strategy.
[0043] Please see Figure 3 In one embodiment, after creating the digital twin, the method further includes: step S001, when a change is detected in the configuration data and status data of the optical network unit, synchronizing the changed configuration data and status data to the digital twin.
[0044] The configuration data consists of static or semi-static service parameters for the optical network unit, including WAN / LAN configuration, WiFi parameters, QoS policies (Quality of Service policies), firewall rules, etc. The status data describes whether the device or hardware component is currently in a normal / abnormal operating condition, distinguishing only between "yes / no," "optical module inserted / removed," and "online / offline."
[0045] As an example, in step S001, the system monitors optical network unit (ONU) configuration change events in real time via the TR069 protocol (CPE WAN management protocol). When the ONU configuration changes, the system immediately synchronizes the changes to the digital twin, ensuring that the digital twin's configuration is completely consistent with the physical ONU's configuration. The ONU device has a built-in lightweight status monitoring agent. When key statuses in the status data change (such as online / offline, optical module insertion / removal, temperature exceeding limits, etc.), the agent proactively pushes status change events to the cloud, and the digital twin updates the corresponding status fields in real time. In this embodiment, by synchronizing the changed configuration and status data to the digital twin in real time via the cloud, the system ensures complete consistency between the virtual model and the physical device configuration.
[0046] Please see Figure 4 In one embodiment, the fault diagnosis model includes a lifetime diagnosis model, an optical module performance degradation diagnosis model, and a memory resource overload diagnosis model; step S103 involves inputting the running data into the pre-trained fault diagnosis model, using the fault diagnosis model to diagnose different types of fault risks, and obtaining fault diagnosis results including: Step S1031: Extract flash memory operation data from the operation data, input the flash memory operation data into the life diagnosis model for life diagnosis, and obtain the remaining life of flash memory; Step S1032: Extract optical module operation data from the operation data, input the optical module operation data into the optical module performance degradation diagnosis model for degradation diagnosis, and obtain the optical module degradation rate; Step S1033: Extract memory operation data from the operation data, input the memory operation data into the resource overload diagnostic model, and obtain the memory utilization rate.
[0047] The system deploys multiple diagnostic models to predict different types of fault risks. These models include a lifespan diagnostic model, an optical module performance degradation diagnostic model, and a memory resource overload diagnostic model. The lifespan diagnostic model diagnoses the aging progress of the flash memory in the embedded multimedia card of the optical network unit based on the rated lifespan parameters of the flash memory. The optical module performance degradation diagnostic model is trained using a time-series analysis algorithm, continuously fitting the curves of optical power and current changes over time to identify long-term slow degradation trends and distinguish between instantaneous fluctuations and permanent hardware aging. The resource overload diagnostic model uses a time-series prediction model (LSTM / Prophet), taking long-term memory and CPU load data as input to identify peak load patterns and abnormal occupancy peaks, predicting future resource congestion trends.
[0048] As an example, in step S1031, EMMC flash memory operation data is extracted from the operation data output by the digital twin. This flash memory operation data is then imported into a pre-trained lifetime diagnostic model to predict the degree of aging, and the remaining usable lifetime of the embedded multimedia card flash memory in the optical network unit is output. This embodiment extracts dedicated flash memory operation indicators and uses a dedicated lifetime diagnostic model, differing from traditional single-threshold alarms. It quantifies the remaining lifespan of storage, rather than only issuing alarms after the flash memory is completely damaged, enabling early warning of storage aging several months in advance. This addresses the shortcomings of traditional maintenance systems that cannot predict the progressive wear and tear of EMMC flash memory.
[0049] Among them, the optical module operation data is the optical link hardware indicator extracted from the overall operation data, which includes continuous time-series values such as the optical module's transmit optical power, receive optical power, operating temperature, bias current, optical bit error rate, and fiber loss. The optical module attenuation rate is the decrease in optical receive power per unit period, representing how fast the optical module hardware ages; the higher the attenuation rate, the shorter the optical link will reach the failure threshold, causing disconnections, interruptions, etc.
[0050] As an example, in step S1032, the optical module operating data is separated from the device operating data, and the optical module operating data is input into the optical module performance degradation diagnostic model to fit the long-term power change curve, thereby calculating the performance degradation rate of the optical module. This embodiment calculates the long-term degradation rate of the optical module using the optical module performance degradation diagnostic model, identifying slow and continuous optical power degradation trends. Unlike traditional network management systems that can only detect whether the current optical power exceeds limits and cannot predict optical attenuation exceeding standards several months later, this solution can achieve early warning of optical module failures.
[0051] Among them, memory operation data are indicators reflecting memory usage status, including time-series data such as real-time memory utilization, peak memory usage, memory leak increment, free memory capacity, and background process memory usage. Memory utilization is the percentage of currently occupied memory relative to the total physical memory; a persistently high utilization rate can cause packet buffer overflow, process freezes, and overall network lag.
[0052] As an example, in step S1033, memory operation data is extracted from the overall operation data, and the memory operation data is sent to the resource overload diagnostic model for load trend analysis. Real-time memory utilization is output, and historical data is used to predict long-term overload risks. This embodiment relies on the resource overload diagnostic model to analyze memory time-series load data, not only outputting real-time memory utilization but also identifying periodic load peaks and increasing memory leak trends, thus predicting future resource congestion in advance.
[0053] Please see Figure 5 In one embodiment, step S104, obtaining the fault handling strategy corresponding to the fault diagnosis result, includes: Step S1041: If the remaining lifespan of the flash memory is lower than the lifespan threshold, the fault diagnosis strategy is determined to be adjusting the flash memory configuration parameters; Step S1042: If the optical module attenuation rate is higher than the attenuation rate threshold, the fault diagnosis strategy is determined to be to switch to a backup optical module or adjust the transmit power of the optical module. Step S1043: If the memory usage rate is higher than the usage rate threshold, the fault diagnosis strategy is determined to be adjusting the service quality strategy or limiting the execution frequency of background tasks.
[0054] Among them, the lifetime threshold is a pre-set critical value for the remaining lifetime of the EMMC flash memory, representing the warning boundary when the flash memory is nearing failure. The flash memory configuration parameters are device configuration items that control the flash memory erase and write frequency and write volume of the optical network unit, including parameters such as log output level, log reporting cycle, and preset log debugging functions.
[0055] As an example, in step S1041, if the remaining lifespan of the flash memory obtained from the diagnosis is less than a preset lifespan threshold, the corresponding fault handling strategy is to adjust the relevant configuration parameters of the flash memory. For example, reducing the device log level (e.g., from DEBUG to INFO), reducing the log write frequency, and disabling non-preset debugging functions, thereby reducing the storage write rate and extending the lifespan of the EMMC. This embodiment sets a unified lifespan threshold as a risk assessment standard, replacing the traditional fixed optical power and instantaneous bit error rate thresholds. This allows for the identification of potential flash memory aging hazards months in advance, solving the problem that traditional devices can only alarm after storage failure and have no early intervention measures.
[0056] The attenuation rate threshold is a preset critical standard for the aging of optical module performance, representing the maximum allowable decrease in optical received power per cycle. If the actual attenuation rate of the optical module exceeds this threshold, it indicates that the hardware is aging too rapidly, and low optical power, network packet loss, and network outages will occur in the short term. The transmit power of the optical module is the power value of the optical signal transmitted from the optical module to the optical fiber, which can be dynamically adjusted. Appropriately increasing the transmit power can compensate for optical fiber line loss, offset the natural attenuation of the optical module, and extend the normal service life of the optical module.
[0057] As an example, in step S1042, if the calculated optical module attenuation rate is greater than a preset attenuation rate threshold, the corresponding fault handling strategy is to switch to a backup optical module or dynamically increase the optical module's transmit power to compensate for optical link loss. This embodiment uses the attenuation rate as the criterion, rather than relying solely on instantaneous optical power values, which can distinguish between instantaneous fiber jitter and permanent hardware aging, significantly reducing the false alarm rate of optical module failures. Simultaneously, by employing both methods—adjusting transmit power to compensate for optical attenuation and automatically switching the hardware optical path when using a backup optical module—the failure time of the optical module can be delayed to the greatest extent possible. Furthermore, the handling operation can be automatically scheduled to be performed during off-peak hours in the early morning, avoiding peak periods for internet access and IPTV, thus preventing video buffering and network drops, balancing fault prevention and user experience.
[0058] The utilization threshold is a preset critical percentage of memory usage, serving as a standard for determining system resource overload. Prolonged memory usage exceeding this threshold can lead to cache overflow, process freezes, and network congestion, necessitating the implementation of traffic and task management measures. The Quality of Service (QoS) policy consists of built-in traffic scheduling rules within the optical network unit, used to prioritize services, limit single-stream bandwidth, and schedule packet forwarding queues. Adjusting this policy can reduce the memory and CPU resources consumed by high-traffic non-critical services, alleviating overall system load pressure. The background task execution frequency is the interval between non-business processes such as background log collection, traffic statistics, scheduled monitoring, and remote debugging. Reducing the execution frequency can decrease the frequency of background programs consuming memory, alleviating the problem of sustained high memory load.
[0059] As an example, in step S1043, if the real-time memory usage exceeds a preset usage threshold, the corresponding fault handling strategy is matched, such as adjusting the service quality strategy, reducing the execution frequency of background tasks, and optimizing the caching strategy. This embodiment identifies resource overload risks based on memory usage thresholds and uses multiple differentiated software optimization methods to simultaneously reduce memory usage from multiple dimensions, including traffic scheduling, background processes, and caching strategies. Compared with a single management method, the load mitigation effect is more significant.
[0060] As an example, for issues such as resource leaks and process crashes caused by software anomalies, the system automatically restarts the relevant services or devices within an appropriate time window (such as 2-5 AM). In scenarios that support load balancing, the system diverts some traffic to adjacent devices to reduce the load on the local device.
[0061] In one embodiment, step S103, which involves inputting running data into a pre-trained fault diagnosis model, using the fault diagnosis model to diagnose different types of fault risks, and obtaining fault diagnosis results, further includes: Step S002: Based on the fault diagnosis results, and combined with the operating time and fault history of the optical network unit, obtain the health score of the optical network unit.
[0062] Among them, runtime is the total online operating time of the optical network unit gateway from its initial power-on to the present moment. The longer the equipment operates, the higher the probability of EMMC flash memory wear, optical module natural aging, and memory process leakage. It is a basic dimension for assessing the overall aging degree of the equipment. Fault history is a complete record of past faults of the equipment archived in the digital twin, including historical storage life alarms, optical module attenuation alarms, memory overload records, various self-healing records, fault occurrence time and handling results. It is used to determine whether the equipment is prone to failure or has high potential risks. Health score is a score ranging from 0 to 100, quantified from multiple dimensions of indicators. It reflects the overall status of the equipment and integrates multiple factors such as remaining flash memory life, optical module attenuation rate, memory load, cumulative runtime of the equipment, and historical fault frequency. The lower the score, the greater the overall operational risks of the equipment and the higher the probability of fault outbreak. It serves as the basis for determining graded operation and maintenance and priority handling. When the score is below the threshold, an early warning is triggered.
[0063] As an example, in step S002, the fault diagnosis results are combined with multiple factors such as the cumulative runtime of the optical network unit, historical fault records, and ambient temperature to comprehensively calculate the device health score corresponding to the optical network unit. This embodiment uses multiple dimensions such as runtime and fault history to make a comprehensive score, rather than viewing instantaneous fault data in isolation. This avoids biased misjudgments caused by a single indicator, and the overall equipment operating condition assessment is more objective and comprehensive.
[0064] In one embodiment, after obtaining the fault handling strategy corresponding to the fault diagnosis result in step S104, the method further includes: Step S105: When the fault handling strategy fails to be obtained, generate an operation and maintenance work order corresponding to the fault diagnosis result and issue an alarm notification.
[0065] The maintenance work order is a standardized electronic document that carries complete equipment fault information, including equipment number, installation address, fault type, risk level, equipment health status, fault cause, suggested handling plan, and estimated fault occurrence time. It is automatically sent to the maintenance personnel's management terminal as a warrant for on-site repairs and hardware replacements. Alarm notifications refer to real-time message pushes to maintenance personnel, which can be sent via maintenance platform pop-ups, SMS, and background messages to quickly alert them to high-risk faults or devices with failed self-healing mechanisms, ensuring immediate fault detection.
[0066] As an example, in step S105, for risks that cannot be resolved through automatic configuration adjustments (such as impending hardware failure), the system automatically generates a maintenance work order and sends an alarm notification to the maintenance personnel. The alarm information includes detailed information such as device location, fault type, risk level, recommended measures, and estimated failure time, facilitating maintenance personnel to prepare spare parts and arrange replacements in advance. This embodiment achieves layered processing by distinguishing fault handling boundaries: software-optimizable storage, memory, and minor optical decay faults are resolved through fault handling strategies; fault handling strategies are ineffective, and hardware aging faults trigger manual intervention processes, avoiding indiscriminate dispatching of all work orders, significantly reducing unnecessary on-site visits, and saving maintenance manpower and transportation costs.
[0067] Please see Figure 6 In one embodiment, an optical network unit (ONU) fault diagnosis device 300 is provided, which corresponds one-to-one with the ONU fault diagnosis method described in the above embodiments. The ONU fault diagnosis device includes: Module 301 is used to create a digital twin corresponding to an optical network unit; The simulation module 302 is used to simulate the operation of an optical network unit using a digital twin to obtain operational data; The fault diagnosis module 303 is used to input the running data into the pre-trained fault diagnosis model, use the fault diagnosis model to diagnose different types of fault risks, and obtain fault diagnosis results. The acquisition and reception module 304 is used to acquire the fault handling strategy corresponding to the fault diagnosis result and execute the fault handling strategy.
[0068] Optionally, in one embodiment, the optical network unit fault diagnosis device further includes a data synchronization module, which is used to synchronize the changed configuration data and status data to the digital twin when a change is detected in the configuration data and status data of the optical network unit.
[0069] Optionally, in one embodiment, the fault diagnosis model includes a lifetime diagnosis model, an optical module performance degradation diagnosis model, and a memory resource overload diagnosis model; the fault diagnosis module 303 is further used for: Extract flash memory operation data from the operation data, input the flash memory operation data into the lifespan diagnostic model for lifespan diagnostics, and obtain the remaining lifespan of the flash memory; The optical module's operating data is extracted from the operating data, and then input into the optical module's performance degradation diagnostic model for degradation diagnosis to obtain the optical module's degradation rate. Extract memory usage data from the runtime data, input the memory usage data into the resource overload diagnostic model, and obtain the memory utilization rate.
[0070] Optionally, in one embodiment, the acquisition and reception module 304 is further configured to: If the remaining lifetime is less than the lifetime threshold, the fault diagnosis strategy is determined to be adjusting the flash memory configuration parameters; If the attenuation rate is higher than the attenuation rate threshold, the fault diagnosis strategy is determined to be to switch to the backup optical module or adjust the transmit power of the optical module. If memory usage exceeds the usage threshold, the fault diagnosis strategy is determined to be adjusting the quality of service strategy or limiting the execution frequency of background tasks.
[0071] Optionally, in one embodiment, the optical network unit fault diagnosis device further includes a health score unit, which is used to obtain a health score of the optical network unit based on the fault diagnosis results and in combination with the runtime and fault history of the optical network unit.
[0072] Optionally, in one embodiment, the optical network unit fault diagnosis device further includes an alarm notification module, which is used to generate an operation and maintenance work order corresponding to the fault diagnosis result and issue an alarm notification when the fault handling strategy fails.
[0073] Specific limitations regarding the optical network unit fault diagnosis device 300 can be found in the limitations of the optical network unit fault diagnosis method described above, and will not be repeated here. Each module in the aforementioned optical network unit fault diagnosis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0074] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores runtime data, configuration data, and terminal device information. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for diagnosing optical network unit faults.
[0075] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the optical network unit fault diagnosis method described in the above embodiment. To avoid repetition, it will not be described again here. The computer-readable storage medium can be non-volatile or volatile.
[0076] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0077] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0078] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for diagnosing faults in an optical network unit, characterized in that, include: Create a digital twin corresponding to the optical network unit; The operation process of the optical network unit is simulated using the digital twin to obtain operational data; The running data is input into a pre-trained fault diagnosis model, and the fault diagnosis model is used to diagnose different types of fault risks to obtain fault diagnosis results. Obtain the fault handling strategy corresponding to the fault diagnosis result, and execute the fault handling strategy.
2. The optical network unit fault diagnosis method as described in claim 1, characterized in that, The method further includes: when a change is detected in the configuration data and status data of the optical network unit, synchronizing the changed configuration data and status data to the digital twin.
3. The optical network unit fault diagnosis method as described in claim 1, characterized in that, The fault diagnosis model includes a lifespan diagnosis model, an optical module performance degradation diagnosis model, and a memory resource overload diagnosis model. The operational data is input into the pre-trained fault diagnosis model, which is then used to diagnose different types of fault risks, yielding fault diagnosis results including: Flash memory operation data is extracted from the operation data, and the flash memory operation data is input into the lifetime diagnosis model for lifetime diagnosis to obtain the remaining lifetime of the flash memory. The optical module operation data is extracted from the operation data, and the optical module operation data is input into the optical module performance degradation diagnosis model for degradation diagnosis to obtain the optical module degradation rate. Memory operation data is extracted from the operation data, and the memory operation data is input into the resource overload diagnostic model to obtain the memory utilization rate.
4. The optical network unit fault diagnosis method as described in claim 3, characterized in that, Obtaining the fault handling strategy corresponding to the fault diagnosis result includes: If the remaining lifespan of the flash memory is lower than the lifespan threshold, then the fault diagnosis strategy is determined to be adjusting the flash memory configuration parameters; If the attenuation rate of the optical module is higher than the attenuation rate threshold, then the fault diagnosis strategy is determined to be to switch to a backup optical module or to adjust the transmission power of the optical module. If the memory usage rate is higher than the usage rate threshold, the fault diagnosis strategy is determined to be either adjusting the quality of service strategy or limiting the execution frequency of background tasks.
5. The optical network unit fault diagnosis method as described in claim 4, characterized in that, The flash memory configuration parameters include at least one of the following: the log level of the optical network unit, the log write frequency, and the preset debugging function.
6. The optical network unit fault diagnosis method as described in claim 1, characterized in that, The method further includes inputting runtime data into a pre-trained fault diagnosis model, using the model to diagnose different types of fault risks, and obtaining fault diagnosis results. Based on the fault diagnosis results, and combined with the runtime and fault history of the optical network unit, a health score of the optical network unit is obtained.
7. The optical network unit fault diagnosis method as described in claim 1, characterized in that, After obtaining the fault handling strategy corresponding to the fault diagnosis result, the method further includes: When the fault handling strategy fails to be obtained, an operation and maintenance work order corresponding to the fault diagnosis result is generated and an alarm notification is issued.
8. A fault diagnosis device for an optical network unit, characterized in that, The device includes: A module is created to generate digital twins corresponding to optical network units; The simulation module is used to simulate the operation of the optical network unit using the digital twin to obtain operational data; The fault diagnosis module is used to input the running data into a pre-trained fault diagnosis model, and use the fault diagnosis model to diagnose different types of fault risks to obtain fault diagnosis results. The acquisition and execution module is used to acquire the fault handling strategy corresponding to the fault diagnosis result and execute the fault handling strategy.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the optical network unit fault diagnosis method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the optical network unit fault diagnosis method as described in any one of claims 1-7.