L1 Software Fail-Safe Recovery in Wireless Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current wireless communication systems face significant downtime and service outages due to unexpected software errors in the Layer 1 (L1) software, leading to crashes and prolonged reboot times, which compromise network availability and Key Performance Indicators (KPIs) such as call drop rates.
Innovation Solution
Implementing a system where the Layer 1 (L1) module monitors discrepancies in information processed by hardware accelerators, resets queues, and performs request-response communication to deactivate and restart the L1 module within the expiry time of the Radio-Link Failure Timer (T310) and Radio Link Re-establishment Timer (T311), ensuring minimal outage time and maintaining user equipment connections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If L1 software encounters unexpected errors and crashes, then software reliability is improved through restart, but network availability deteriorates due to prolonged downtime
Solution Approach 1:
The patent segments the L1 software into multiple independent processes (L1 processes) that can operate in parallel. When one L1 process crashes due to unexpected errors, only that specific process needs to be restarted rather than the entire L1 software or DU. This segmentation isolates failures and enables faster recovery, resolving the contradiction between maintaining software reliability through restarts and minimizing network downtime.
Solution Approach 2:
The patent implements preliminary actions by pre-configuring backup L1 processes and establishing monitoring mechanisms that detect crashes before they propagate. The L2 software continuously monitors L1 processes and has pre-prepared recovery procedures ready to execute immediately upon detecting a crash. This preliminary preparation enables rapid response and minimizes downtime while maintaining reliability.
2Stability of the object's composition
If L1 software crashes and DU reboots to recover, then system stability is improved, but service continuity deteriorates due to cell outage
Solution Approach 1:
The patent divides the L1 functionality into multiple independent processes that can fail and recover independently. When one L1 process crashes, the system maintains stability by keeping other L1 processes running, and service continuity is preserved by rapidly restarting only the failed process rather than rebooting the entire DU. This segmentation allows localized recovery without system-wide disruption.
Solution Approach 2:
The L2 software acts as an intermediary between the crashed L1 process and the core network. It detects the crash, manages the restart of the failed L1 process, and handles communication with the core network throughout the recovery process. This intermediary role ensures smooth transition and maintains service continuity while restoring system stability.
3Measurement precision
If hardware accelerator takes excessive cycles due to out-of-range attributes, then processing accuracy is improved, but L1 software reliability deteriorates due to crashes
Solution Approach 1:
The patent implements feedback mechanisms where the L1 software monitors the performance and status of hardware accelerators. When out-of-range attributes cause excessive processing cycles or failures, the feedback system detects these conditions and triggers appropriate responses such as adjusting attributes or restarting the affected L1 process. This feedback loop maintains processing accuracy while preventing reliability deterioration through timely corrections.
Solution Approach 2:
The system performs preliminary validation of attributes before they are processed by hardware accelerators. By checking attributes in advance and correcting out-of-range values before processing, the system prevents crashes caused by invalid inputs while maintaining the accuracy needed for proper hardware accelerator operation.
Data Source
AI summary
The present invention provides an efficient and reliable systems and methods for facilitating FAIL SAFE possibilities in a Network by exploiting 3GPP defined Radio Resource Control (RRC) T310 (Radio-Link Failure Timer), N310 (Radio-Link Failure Counter), T311 (Radio Link Re-establishment Timer), N311 (Radio Link Re-establishment Counter) Timers and associated Counters to enable the L1 to recover within the combined duration of the sum of T310 and T311 timers, for example, typically, 100 msec following a L1 SW exception event.

