Recovering system from uncorrectable memory errors

WO2026206498A1PCT designated stage Publication Date: 2026-10-01QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/015907
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-02-19
Publication Date
2026-10-01

Smart Images

  • Figure US2026015907_01102026_PF_FP_ABST
    Figure US2026015907_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Aspects presented herein may improve the overall performance and efficiency related to recovering a system / subsystem from uncorrectable memory error(s), such as a system / subsystem with resource constraints. For example, an interrupt-based mechanism is provided to handle uncorrectable memory error(s) in a more efficient manner. In one aspect, a device detects an uncorrectable error in memory of a process domain (PD). The device increments an uncorrectable error counter based on detection of the uncorrectable error in the memory. The device performs, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuing a restart for the PD if the uncorrectable error counter does not exceed the counter threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Qualcomm Ref. No. 2501028WO 1 / 53RECOVERING SYSTEM FROM UNCORRECTABLE MEMORY ERRORSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Non-Provisional Patent Application No.19 / 092,873, entitled “RECOVERING SYSTEM FROM UNCORRECTABLE MEMORY ERRORS” and filed on March 27, 2025, which is expressly incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates generally to system recovery, and more particularly, to recovering system from uncorrectable memory errors with resource constraints.INTRODUCTION

[0003] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, and broadcasts. Typical wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple-access technologies include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.

[0004] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate on a municipal, national, regional, and even global level. An example telecommunication standard is 5G New Radio (NR). 5G NR is part of a continuous mobile broadband evolution promulgated by Third Generation Partnership Project (3 GPP) to meet new requirements associated with latency, reliability, security, scalability (e.g., with Internet of Things (IoT)), and other requirements. 5G NR includes services associated with enhanced mobile broadband (eMBB), massive machine type communications (mMTC), and ultra-reliable low latency communications (URLLC). Some aspects of 5G NR may be based on the 4G Long129025-2565WO01Qualcomm Ref. No. 2501028WO 2 / 53Term Evolution (LTE) standard. There exists a need for further improvements in 5G NR technology. These improvements may also be applicable to other multi-access technologies and the telecommunication standards that employ these technologies.BRIEF SUMMARY

[0005] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects. This summary neither identifies key or critical elements of all aspects nor delineates the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.

[0006] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus detects an uncorrectable error in memory of a process domain (PD). The apparatus increments an uncorrectable error counter based on detection of the uncorrectable error in the memory. The apparatus performs, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or(3)issuinga re start for the PD if the uncorrectable error counter does not exceed the counter threshold.

[0007] To the accomplishment of the foregoing and related ends, the one or more aspects may include the features hereinafter fully described and particularly pointed outin the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects may be employed.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG. 1 is a diagram illustrating an example of a wireless communications system and an access network.

[0009] FIG. 2A is a diagram illustrating an example of a first frame, in accordance with various aspects of the present disclosure.129025-2565WO01Qualcomm Ref. No. 2501028WO 3 / 53

[0010] FIG. 2B is a diagram illustrating an example of downlink (DL) channels within a subframe, in accordance with various aspects of the present disclosure.

[0011] FIG. 2C is a diagram illustrating an example of a second frame, in accordance with various aspects of the present disclosure.

[0012] FIG. 2D is a diagram illustrating an example of uplink (UL) channels within a subframe, in accordance with various aspects of the present disclosure.

[0013] FIG. 3 is a diagram illustrating an example of a base station and user equipment (UE) in an access network.

[0014] FIG. 4 is a diagram illustrating an example of a vehicle performing road object detection using different types of sensors in accordance with various aspects of the present disclosure.

[0015] FIG. 5 is a diagram illustrating an example of types of errors in caches in accordance with various aspects of the present disclosure.

[0016] FIG. 6 is a diagram illustrating an example of an interrupt-based mechanism that is capable of handling uncorrectable memory error(s) in a more efficient manner in accordance with various aspects of the present disclosure.

[0017] FIG. 7 is a flow chart illustrating an example of the interrupt-based mechanism with an uncorrectable error counter in accordance with various aspects of the present disclosure.

[0018] FIG. 8 is a communication flow illustrating an example system programming sequence forthe interrupt-basedmechanism in accordance with various aspects of the present disclosure.

[0019] FIG. 9 is a flowchart of a method of recovering system from uncorrectable memory error(s).

[0020] FIG. 10 is a flowchart of a method of recovering system from uncorrectable memory error(s).

[0021] FIG. 11 is a diagram illustrating an example of a hardware implementation for an example apparatus and / or network entity.DETAILED DESCRIPTION

[0022] Aspects presented herein may improve the overall performance and efficiency related to recovering a system / sub system from uncorrectable memory error(s), such as a system / sub system with resource constraints. In one aspect of the present disclosure,129025-2565WO01Qualcomm Ref. No. 2501028WO 4 / 53an interrupt-based mechanism is provided to handle uncorrectable memory errors) in a more efficient manner. When an uncorrectable error in memory, such as vector tightly coupled memory (VTCM), is encountered, the error instance is logged in the error logger region of the memory, and an interrupt is raised to the subsystem or a high-level operation system (HLOS) (e.g., an application processor) in order to take specified action(s). For example, the process identification (ID) (e.g., the victim process domain (PD)) associated with the error may be identified based on software (SW) thread ID programmed in hardware (HW)thread / core on which error correction code (ECC) fault is encountered. The faulty region or cache line may be recorded using a register or ECC reserved bits. The hardware threads associated with the process ID may be interrupted. Faulty cache line / region may be restricted or made unusable by other PDs by setting the corresponding invalid bit to high. An additional interruptto higher operating system (OS) or subsystem may be configuredto be raised to high once the number of faults exceeds a threshold. Such an interrupt may trigger a complete system restart. The victim PD may be restarted once a recovery from the error has been made.

[0023] The detailed description set forth below in connection with the drawings describes various configurations and does not represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, these concepts may be practiced without these specific details. In some instances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.

[0024] Several aspects of telecommunication systems are presented with ref erenceto various apparatus and methods. These apparatus and methods are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as “elements”). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0025] By way of example, an element, or any portion of an element, or any combination of elements may be implemented as a “processing system” that includes one or more129025-2565WO01Qualcomm Ref. No. 2501028WO 5 / 53processors. When multiple processors are implemented, the multiple processors may perform the functions individually or in combination. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on a chip (SoC), baseband processors, field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system may execute software. Software, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, or any combination thereof.

[0026] Accordingly, in one or more example aspects, implementations, and / or use cases, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a computer. By way of example, such computer-readable media can include a random-access memory (RAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the types of computer- readable media, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.

[0027] While aspects, implementations, and / or use cases are describedin this application by illustration to some examples, additional or different aspects, implementations and / or use cases may come about in many different arrangements and scenarios. Aspects, implementations, and / oruse cases described herein may be implemented across many differingplatform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects, implementations, and / or use cases may come about via integrated chip implementations and other non-module-component based devices129025-2565WO01Qualcomm Ref. No. 2501028WO 6 / 53(e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, artificial intelligence (Al)-enabled devices, etc.). While some examples may or may not be specifically directed to use cases or applications, a wide assortment of applicability of described examples may occur. Aspects, implementations, and / or use cases may range a spectrum from chip-level or modular components to non-modular, non-chip- level implementations and further to aggregate, distributed, or original equipment manufacturer (OEM) devices or systems incorp oratingone or more techniques herein. In some practical settings, devices incorporating described aspects and features may also include additional components and features for implementation and practice of claimed and described aspect. For example, transmission and reception of wireless signals necessarily includes a number of components for analog and digital purposes (e.g., hardware components including antenna, radio frequency (RF)-chains, power amplifiers, modulators, buffer, processor(s), interleaver, adders / summers, etc.). Techniques described herein may be practiced in a wide variety of devices, chip-level components, systems, distributed arrangements, aggregated or disaggregated components, end-user devices, etc. of varying sizes, shapes, and constitution.

[0028] Deployment of communication systems, such as 5GNR systems, may be arranged in multiple manners with various components or constituent parts. In a 5G NR system, or network, a network node, a network entity, a mobility element of a network, a radio access network (RAN) node, a core network node, a network element, or a network equipment, such as a base station (BS), or one or more units (or one or more components) performing base station functionality, may be implemented in an aggregated or disaggregated architecture. For example, a BS (such as a Node B (NB), evolved NB (eNB), NR BS, 5GNB, access point (AP), a transmission reception point (TRP), or a cell, etc.) may be implemented as an aggregated base station (also known as a standalone BS or a monolithic BS) or a disaggregated base station.

[0029] An aggregated base station may be configured to utilize a radio protocol stack that is physically or logically integrated within a single RAN node. A disaggregated base station may be configured to utilize a protocol stack that is physically or logically distributed among two or more units (such as one or more central or centralized units (CUs), one or more distributed units (DUs), or one or more radio units (RUs)). In some aspects, a CU may be implemented within a RAN node, and one or more DUs129025-2565WO01Qualcomm Ref. No. 2501028WO 7 / 53may be co-located with the CU, or alternatively, may be geographically or virtually distributed throughout one or multiple other RAN nodes. The DUs may be implemented to communicate with one or more RUs. Each of the CU, DU and RU can be implemented as virtual units, i.e., a virtual central unit (VCU), a virtual distributed unit (VDU), or a virtual radio unit.

[0030] Base station operation or network design may consider aggregation characteristics of base station functionality. For example, disaggregated base stations may be utilized in an integrated access backhaul (IAB) network, an open radio access network (O- RAN (such as the network configuration sponsored by the 0-RAN Alliance)), or a virtualized radio access network (vRAN, also known as a cloud radio access network (C-RAN)). Disaggregation may include distributing functionality across two or more units at various physical locations, as well as distributing functionality for at least one unit virtually, which can enable flexibility in network design. The various units of the disaggregated base station, or disaggregated RAN architecture, can be configured for wired or wireless communication with at least one other unit.

[0031] FIG. 1 is a diagram 100 illustrating an example of a wireless communications system and an access network. The illustrated wireless communications system includes a disaggregated base station architecture. The disaggregated base station architecture may include one or more CUs 110 that can communicate directly with a core network 120 via a backhaul link, or indirectly with the core network 120 through one or more disaggregated base station units (such as a Near-Real Time (Near-RT) RAN Intelligent Controller (RIC) 125 via an E2 link, or a Non-Real Time (Non-RT)RIC 115 associated with a Service Management and Orchestration (SMO) Framework 105, or both). A CU 110 may communicate with one or more DUs 130 via respective midhaul links, such as an Fl interface. The DUs 130 may communicate with one or more RUs 140 via respective fronthaul links. The RUs 140 may communicate with respective UEs 104 via one or more radio frequency (RF) access links. In some implementations, theUE 104 may be simultaneously served by multiple RUs 140.

[0032] Each of the units, i.e., the CUs 110, the DUs 130, the RUs 140, as well as the Near- RT RIC s 125, the Non-RT RICs 115, and the SMO Framework 105, may include one or more interfaces or be coupled to one or more interfaces configured to receive or to transmit signals, data, or information (collectively, signals) via a wired or wireless transmission medium. Each of the units, or an associated processor or controller129025-2565WO01Qualcomm Ref. No. 2501028WO 8 / 53providing instructions to the communication interfaces of the units, can be configured to communicate with one or more of the other units via the transmission medium. For example, the units can include a wired interface configured to receive or to transmit signals over a wired transmission medium to one or more of the other units. Additionally, the units can include a wireless interface, which may include a receiver, a transmitter, or a transceiver (such as an RF transceiver), configured to receive or to transmit signals, or both, over a wireless transmission medium to one or more of the other units.

[0033] In some aspects, the CU 110 may host one or more higher layer control functions.Such control functions can include radio resource control (RRC), packet data convergence protocol (PDCP), service data adaptation protocol (SDAP), or the like. Each control function can be implemented with an interface configured to communicate signals with other control functions hosted by the CU 110. The CU 110 may be configured to handle user plane functionality (i.e., Central Unit - User Plane (CU-UP)), control plane functionality (i.e., Central Unit - Control Plane (CU-CP)), or a combination thereof. In some implementations, theCU 110 can be logically split into one or more CU-UP units and one or more CU-CP units. The CU-UP unit can communicate bidirectionally with the CU-CP unit via an interface, such as an El interface when implemented in an 0-RAN configuration. The CU 110 can be implemented to communicate with the DU 130, as necessary, for network control and signaling.

[0034] The DU 130 may correspond to a logical unit that includes one or more base station functions to control the operation of one or more RUs 140. In some aspects, the DU 130 may host one or more of a radio link control (RLC) layer, a medium access control (MAC) layer, and one or more high physical (PHY) layers (such as modules for forward error correction (FEC) encoding and decoding, scrambling, modulation, demodulation, or the like) depending, at least in part, on a functional split, such as those defined by 3 GPP. In some aspects, the DU 130 may further host one or more low PHY layers. Each layer (or module) can be implemented with an interface configured to communicate signals with other layers (and modules) hosted by the DU 130, or with the control functions hosted by the CU 110.

[0035] Lower-layer functionality can be implemented by one or more RUs 140. In some deployments, an RU 140, controlled by a DU 130, may correspond to a logical node129025-2565WO01Qualcomm Ref. No. 2501028WO 9 / 53that hosts RF processing functions, or low-PHY layer functions (such as performing fast Fourier transform (FFT), inverse FFT (iFFT), digital beamforming, physical random access channel (PRACH) extraction and filtering, or the like), or both, based at least in part on the functional split, such as a lower layer functional split. In such an architecture, the RU(s) 140 can be implemented to handle over the air (OTA) communication with one or more UEs 104. In some implementations, real-time and non-real-time aspects of control and user plane communication with the RU(s) 140 can be controlled by the corresponding DU 130. In some scenarios, this configuration can enable the DU(s) 130 and the CU 110 to be implemented in a cloud-based RAN architecture, such as a vRAN architecture.

[0036] The SMO Framework 105 may be configured to support RAN deployment and provisioning of non-virtualizedandvirtualizednetwork elements. Fornon-virtualized network elements, the SMO Framework 105 may be configured to support the deployment of dedicated physical resources for RAN coverage requirements that may be managed via an operations and maintenance interface (such as an 01 interface). For virtualized network elements, the SMO Framework 105 may be configured to interact with a cloud computing platform (such as an open cloud (O-Cloud) 190) to perform network element life cycle management (such as to instantiate virtualized network elements) via a cloud computing platform interface (such as an 02 interface). Such virtualized network elements can include, but are not limited to, CUs 110, DUs 130, RUs 140 andNear-RTRICs 125. In some implementations, the SMO Framework 105 can communicate with a hardware aspect of a 4G RAN, such as an open eNB (O- eNB) 111, via an 01 interface. Additionally, in some implementations, the SMO Framework 105 can communicate directly with one or more RUs 140 via an 01 interface. The SMO Framework 105 also may include aNon-RTRIC 115 configured to support functionality of the SMO Framework 105.

[0037] The Non-RT RIC 115 may be configured to include a logical function that enables non-real-time control and optimization of RAN elements and resources, artificial intelligence (Al) / machine learning (ML) (AI / ML) workflows including model training and updates, or policy-based guidance of applications / features in the Near- RT RIC 125. The Non-RT RIC 115 may be coupled to or communicate with (such as via an Al interface) the Near-RT RIC 125. TheNear-RTRIC 125 may be configured to include a logical function that enables near-real-time control and optimization of129025-2565WO01Qualcomm Ref. No. 2501028WO 10 / 53RAN elements and resources via data collection and actions over an interface (such as via an E2 interface) connecting one or more CUs 110, one or more DUs 130, or both, as well as an O-eNB, with the Near-RT RIC 125.

[0038] In some implementations, to generate AI / ML models to be deployed in the Near-RT RIC 125, the Non-RT RIC 115 may receive parameters or external enrichment information from external servers. Such information may be utilized by the Near-RT RIC 125 and may be received at the SMO Framework 105 or the Non-RT RIC 115 from non-network data sources or from network functions. In some examples, the Non-RTRIC 115 or the Near-RT RIC 125 may be configured to tune RAN behavior or performance. For example, the Non-RT RIC 115 may monitor long-term trends and patterns for performance and employ AI / ML models to perform corrective actions through the SMO Framework 105 (such as reconfiguration via 01) or via creation of RAN management policies (such as Al policies).

[0039] At least one of the CU 110, the DU 130, and the RU 140 may be referred to as a base station 102. Accordingly, abase station 102 may include one ormore ofthe CU 110, the DU 130, and the RU 140 (each component indicated with dotted lines to signify that each component may or may not be included in the base station 102). The base station 102 provides an access point to the core network 120 for aUE 104. The base station 102 may include macrocells (high power cellular base station) and / or small cells (low power cellular base station). The small cells include femtocells, picocells, and microcells. A network that includes both small cell and macrocells may be known as a heterogeneous network. A heterogeneous network may also include Home Evolved Node Bs (eNBs) (HeNBs), which may provide service to a restricted group known as a closed subscriber group (CSG). The communication links between the RUs 140 and the UEs 104 may include uplink (UL) (also referred to as reverse link) transmissions fromaUE 104 to an RU 140 and / or downlink (DL) (also referred to as forward link) transmissions from an RU 140 to aUE 104. The communication links may use multiple-input and multiple-output (MIMO) antenna technology, including spatial multiplexing, beamforming, and / or transmit diversity. The communication links may be through one or more carriers. The base station 102 / UEs 104 may use spectrum up to fMHz (e.g., 5, 10, 15, 20, 100, 400, etc. MHz) bandwidth per carrier allocated in a carrier aggregation of up to a total of Fx MHz (x component carriers) used for transmission in each direction. The carriers may or may not be adjacent to129025-2565WO01Qualcomm Ref. No. 2501028WO 11 / 53each other. Allocation of carriers may be asymmetric with respecttoDL andUL (e.g., more or fewer carriers may be allocated for DL than for UL). The component carriers may include a primary component carrier and one or more secondary component carriers. A primary component carrier may be referred to as a primary cell(PCell) and a secondary component carrier may be referred to as a secondary cell (SCell).

[0040] Certain UEs 104 may communicate with each other using device-to-device (D2D) communication link 158. The D2D communication link 158 may use the DL / UL wireless wide area network (WWAN) spectrum. TheD2D communication link 158 may use one or more sidelink channels, such as a physical sidelink broadcast channel (PSBCH), a physical sidelink discovery channel (PSDCH), a physical sidelink shared channel (PSSCH), and a physical sidelink control channel (PSCCH). D2D communication may be through a variety of wireless D2D communications systems, such as for example, Bluetooth™ (Bluetooth is a trademark of the Bluetooth Special Interest Group (SIG)), Wi-Fi™ (is a trademark of the Wi-Fi Alliance) based on the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard, LTE, orNR.

[0041] The wireless communications system may further include a Wi-Fi AP 150 in communication with UEs 104 (also referred to as Wi-Fi stations (STAs)) via communication link 154, e.g., in a 5 GHz unlicensed frequency spectrum orthe like. When communicating in an unlicensed frequency spectrum, the UEs 104 / AP 150 may perform a clear channel assessment (CCA) prior to communicating in order to determine whether the channel is available.

[0042] The electromagnetic spectrum is often subdivided, based on frequency / wavelength, into various classes, bands, channels, etc. In 5GNR, two initial operating bands have been identified as frequency range designations FR1 (410 MHz - 7.125 GHz) and FR2 (24.25 GHz - 52.6 GHz). Although a portion of FR1 is greater than 6 GHz, FR1 is often referred to (interchangeably) as a “sub-6 GHz” band in various documents and articles. A similar nomenclature issue sometimes occurs with regard to FR2, which is often referred to (interchangeably) as a “millimeter wave” bandin documents and articles, despite being different from the extremely high frequency (EHF) band (30 GHz - 300 GHz) which is identified by the International Telecommunications Union (ITU) as a “millimeter wave” band.

[0043] The frequencies between FR1 andFR2 are often referred to as mid-band frequencies.Recent 5G NR studies have identified an operating band for these mid-band129025-2565WO01Qualcomm Ref. No. 2501028WO 12 / 53frequencies as frequency range designation FR3 (7.125 GHz - 24.25 GHz). Frequency bands falling within FR3 may inherit FR1 characteristics and / or FR2 characteristics, and thus may effectively extend features of FR1 and / or FR2 into midband frequencies. In addition, higher frequency bands are currently being explored to extend 5GNR operation beyond 52.6 GHz. For example, three higher op erating bands have been identified as frequency range designations FR2-2 (52.6 GHz - 71 GHz), FR4 (71 GHz- 114.25 GHz), andFR5 (114.25 GHz- 300 GHz). Each of these hi^ier frequency bands falls within the EHF band.

[0044] With the above aspects in mind, unless specifically stated otherwise, the term “sub-6GHz” or the like if used herein may broadly represent frequencies that may be less than 6 GHz, may be within FR1 , or may include mid-band frequencies. Further, unless specifically stated otherwise, the term “millimeter wave” or the like if used herein may broadly represent frequencies that may include mid-band frequencies, may be within FR2, FR4, FR2-2, and / or FR5, or may be within the EHF band.

[0045] The base station 102 and the UE 104 may each include a plurality of antennas, such as antenna elements, antenna panels, and / or antenna arrays to facilitate beamforming The base station 102 may transmit a beamformed signal 182 to the UE 104 in one or more transmit directions. The UE 104 may receive the beamformed signal from the base station 102 in one or more receive directions. The UE 104 may also transmit a beamformed signal 184 to the base station 102 in one or more transmit directions. The base station 102 may receive the beamformed signal from the UE 104 in one or more receive directions. The base station 102 / UE 104 may perform beam training to determine the best receive and transmit directions for each of the base station 102 / UE 104. The transmit and receive directions for the base station 102 may or may not be the same. The transmit and receive directions for the UE 104 may or may not be the same.

[0046] The base station 102 may include and / or be referred to as a gNB, Node B, eNB, an access point, a base transceiver station, a radio base station, a radio transceiver, a transceiver function, a basic service set (BSS), an extended service set (ESS), a TRP, network node, network entity, network equipment, or some other suitable terminology. The base station 102 can be implemented as an integrated access backhaul (IAB) node, a relay node, a sidelink node, an aggregated (monolithic) base station with a baseband unit (BBU) (including a CU and a DU) and an RU, or as a129025-2565WO01Qualcomm Ref. No. 2501028WO 13 / 53disaggregated base station including one or more of a CU, a DU, and / or an RU. The set of base stations, which may include disaggregated base stations and / or aggregated base stations, may be referred to as next generation (NG) RAN (NG-RAN).

[0047] The core network 120 may include an Access and Mobility Management Function (AMF) 161, a Session Management Function (SMF) 162, a User Plane Function (UPF) 163, a Unified Data Management (UDM) 164, one or more location servers 168, and other functional entities. The AMF 161 is the control node that processes the signaling between the UEs 104 and the core network 120. The AMF 161 supports registration management, connection management, mobility management, and other functions. The SMF 162 supports session management and other functions. The UPF 163 supports packet routing, packet forwarding, and other functions. The UDM 164 supports the generation of authentication and key agreement (AKA) credentials, user identification handling, access authorization, and sub scription management. The one or more location servers 168 are illustrated as including a Gateway Mobile Location Center (GMLC) 165 and a Location Management Function (LMF) 166. However, generally, the one or more location servers 168 may include one or more location / positioning servers, which may include one or more of the GMLC 165, the LMF 166, a position determination entity (PDE), a serving mobile location center (SMLC), a mobile positioning center (MPC), or the like. The GMLC 165 and the LMF 166 support UE location services. The GMLC 165 provides an interface for clients / applications (e.g., emergency services) for accessing UE positioning information. The LMF 166 receives measurements and assistance information from the NG-RAN and the UE 104 via the AMF 161 to compute the position of the UE 104. The NG-RAN may utilize one ormore positioningmethods in orderto determine the position of the UE 104. Positioningthe UE 104 may involve signal measurements, a position estimate, and an optional velocity computation based on the measurements. The signal measurements may be made by the UE 104 and / or the base station 102 serving the UE 104. The signals measured may be based on one ormore of a satellite positioning system (SPS) 170 (e.g., one or more of a Global Navigation Satellite System (GNSS), global position system (GPS), non-terrestrial network (NTN), or other satellite position / location system), LTE signals, wireless local area network (WLAN) signals, Bluetooth signals, a terrestrial beacon system (TBS), sensor-based information (e.g., barometric pressure sensor, motion sensor), NR enhanced cell ID129025-2565WO01Qualcomm Ref. No. 2501028WO 14 / 53(NRE-CID) methods, NRsignals(e.g., multi-round trip time (Multi-RTT), DL angle- of-departure (DL-AoD), DL time difference of arrival (DL-TDOA), UL time difference of arrival (UL-TDOA), and UL angle-of-arrival (UL-AoA) positioning), and / or other systems / signals / sensors.

[0048] Examples of UEs 104 include a cellular phone, a smart phone, a session initiation protocol (SIP) phone, a laptop, a personal digital assistant (PDA), a satellite radio, a global positioning system, a multimedia device, a video device, a digital audio player (e.g., MP3 player), a camera, a game console, a tablet, a smart device, a wearable device, a vehicle, an electric meter, a gas pump, a large or small kitchen appliance, a healthcare device, an implant, a sensor / actuator, a display, or any other similar functioning device. Some of the UEs 104 may be referred to as loT devices (e.g, parking meter, gas pump, toaster, vehicles, heart monitor, etc.). TheUE 104 may also be referred to as a station, a mobile station, a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communications device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, or some other suitable terminology. In some scenarios, the term UE may also apply to one or more companion devices such as in a device constellation arrangement. One or more of these devices may collectively access the network and / or individually access the network.

[0049] Referring again to FIG. 1, in certain aspects, the UE 104 may have a memory error recover component 198 that may be configured detect an uncorrectable error in memory of a process domain (PD); increment an uncorrectable error counter based on detection of the uncorrectable error in the memory; and perform, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or(3)issuinga restart forthePD if the uncorrectable error counter does not exceed the counter threshold. In certain aspects, the base station 102 or the one or more location servers 168 may have a memory error recover configuration component 199 that may be configured to provide memory recovery related configuration(s) to the UE 104.129025-2565WO01Qualcomm Ref. No. 2501028WO 15 / 53

[0050] FIG. 2 A is a diagram 200 illustrating an example of a first subframe within a 5G R frame structure. FIG. 2B is a diagram 230 illustrating an example of DL channels within a 5G NR subframe. FIG. 2C is a diagram 250 illustrating an example of a second subframe within a 5G NR frame structure. FIG. 2D is a diagram 280 illustrating an example of UL channels within a 5 G NR subframe. The 5 G NR frame structure may be frequency division duplexed (FDD) in which for a particular set of subcarriers (carrier system bandwidth), subframes within the set of subcarriers are dedicated for either DL or UL, or may be time division duplexed (TDD) in which for a particular set of subcarriers (carrier system bandwidth), subframes within the set of subcarriers are dedicated for both DL andUL. In the examples provided by FIGs. 2A, 2C, the 5G NR frame structure is assumed to be TDD, with subframe 4 being configured with slot format 28 (with mostly DL), where D is DL, U is UL, and F is flexible for use between DL / UL, and subframe 3 being configured with slot format 1 (with all UL). While subframes 3, 4 are shown with slot formats 1, 28, respectively, any particular subframe may be configured with any of the various available slot formats 0-61. Slot formats 0, 1 are all DL, UL, respectively. Other slot formats 2-61 include a mix of DL, UL, and flexible symbols. UEs are configured with the slot format (dynamically through DL control information (DCI), or semi- statically / statically through radio resource control (RRC) signaling) through a received slot format indicator (SFI). Note that the description infra applies also to a 5G NR frame structure that is TDD.

[0051] FIGs. 2 A-2D illustrate a frame structure, and the aspects of the present disclosure may be applicable to other wireless communication technologies, which may have a different frame structure and / or different channels. A frame (10 ms) may be divided into 10 equally sized subframes (1 ms). Each subframe may include one or more time slots. Subframes may also include mini-slots, which may include 7, 4, or 2 symbols. Each slot may include 14 or 12 symbols, depending on whether the cyclic prefix (CP) is normal or extended. For normal CP, each slot may include 14 symbols, and for extended CP, each slot may include 12 symbols. The symbols on DL may be CP orthogonal frequency division multiplexing (OFDM) (CP-OFDM) symbols. The symbols on UL may be CP-OFDM symbols (for high throughput scenarios) or discrete Fourier transform (DFT) spread OFDM (DFT-s-OFDM) symbols (for power limited scenarios; limited to a single stream transmission). The number of slots within129025-2565WO01Qualcomm Ref. No. 2501028WO 16 / 53a subframe is based on the CP and the numerology. The numerology defines the subcarrier spacing (SCS) (see Table 1). The symbol length / duration may scale with 1 / SCS.Table 1: Numerology, SCS, and CP

[0052] For normal CP (14 symbols / slot), different numerologies p 0 to 4 allowfor 1, 2, 4, 8, and 16 slots, respectively, per subframe. For extended CP, the numerology 2 allows for 4 slots per subframe. Accordingly, for normal CP and numerology p, there are 14 symbols / slot and 2.Llsi ots / sub frame. The subcarrier spacing may be equal to 2^ * 15 kHz, where . is the numerology 0 to 4. As such, the numerology p=0 has a subcarrier spacing of 15 kHz and the numerology p=4 has a subcarrier spacing of 240 kHz. The symbol length / duration is inversely related to the subcarrier spacing. FIGs.2A-2D provide an example of normal CP with 14 symbols per slot and numerology p=2 with 4 slots per subframe. The slot duration is 0.25 ms, the subcarrier spacing is 60 kHz, and the symbol duration is approximately 16.67 ps. Within a set of frames, there may be one or more different bandwidth parts (BWPs) (see FIG. 2B) that are frequency division multiplexed. Each BWP may have a particular numerology and CP (normal or extended).

[0053] A resource grid may be used to represent the frame structure. Each time slot includes a resource block (RB) (also referred to as physical RBs (PRBs)) that extends 12 consecutive subcarriers. The resource grid is divided into multiple resource elements (REs). The number of bits carried by each RE depends on the modulation scheme.

[0054] As illustrated in FIG. 2 A, some of the REs carry reference (pilot) signals (RS) for the UE. The RS may include demodulation RS (DM-RS) (indicated as Rfor one particular129025-2565WO01Qualcomm Ref. No. 2501028WO 17 / 53configuration, but other DM-RS configurations are possible) and channel state information reference signals (CSI-RS) for channel estimation attheUE. The RS may also include beam measurement RS (BRS), beam refinement RS (BRRS), and phase tracking RS (PT-RS).

[0055] FIG. 2B illustrates an example of various DL channels within a subframe of a frame.The physical downlink control channel (PDCCH) carries DCI within one or more control channel elements (CCEs) (e.g., 1, 2, 4, 8, or 16 CCEs), each CCE including six RE groups (REGs), each REG including 12 consecutive REs in an OFDM symbol of an RB. A PDCCH within one BWP may be referred to as a control resource set (CORESET). A UE is configured to monitor PDCCH candidates in a PDCCH search space (e.g., common search space, UE-specific search space) during PDCCH monitoring occasions on the CORESET, where the PDCCH candidates have different DCI formats and different aggregation levels. Additional BWPs may be located at greater and / or lower frequencies across the channel bandwidth. A primary synchronization signal (PSS) may be within symbol 2 of particular subframes of a frame. The PSS is used by a UE 104 to determine subframe / symbol timing and a physical layer identity. A secondary synchronization signal (SSS) may be within symbol 4 of particular subframes of a frame. The SSS is used by a UE to determine a physical layer cell identity group number and radio frame timing. Based on the physical layer identity and the physical layer cell identity group number, the UE can determine a physical cell identifier (PCI). Based on the PCI, the UE can determine the locations of the DM-RS. The physical broadcast channel (PBCH), which carries a master information block (MIB), may be logically grouped with the PSS and SSS to form a synchronization signal (SS) / PBCH block (also referred to as SS block (SSB)). The MIB provides a number of RBs in the system bandwidth and a system frame number (SFN). The physical downlink shared channel (PDSCH) carries user data, broadcast system information not transmitted through the PBCH such as system information blocks (SIBs), and paging messages.

[0056] As illustrated in FIG. 2C, some of the REs carry DM-RS (indicated as R for one particular configuration, but other DM-RS configurations are possible) for channel estimation at the base station. The UE may transmit DM-RS for the physical uplink control channel (PUCCH) and DM-RS for the physical uplink shared channel (PUSCH). The PUSCH DM-RS may be transmitted in the first one or two symbols of129025-2565WO01Qualcomm Ref. No. 2501028WO 18 / 53the PUSCH. The PUCCH DM-RS may be transmitted in different configurations depending on whether short or long PUCCHs are transmitted and depending on the particular PUCCH format used. The UE may transmit sounding reference signals (SRS). The SRS may be transmitted in the last symbol of a subframe. The SRS may have a comb structure, and a UE may transmit SRS on one of the combs. The SRS may be used by a base station for channel quality estimation to enable frequencydependent scheduling on the UL.

[0057] FIG. 2D illustrates an example of various UL channels within a subframe of a frame.The PUCCH may be located as indicated in one configuration. The PUCCH carries uplink control information (UCI), such as scheduling requests, a channel quality indicator (CQI), a precoding matrix indicator (PMI), a rank indicator (RI), and hybrid automatic repeat request (HARQ) acknowledgment (ACK) (HARQ-ACK) feedback (i.e., one or more HARQ ACK bits indicating one or more ACK and / or negative ACK (NACK)). The PUSCH carries data, and may additionally be used to carry a buffer status report (BSR), a power headroom report (PHR), and / or UCI.

[0058] FIG. 3 is a block diagram of a base station 310 in communication with a UE 350 in an access network. In the DL, Internet protocol (IP) packets may be provided to a controller / processor 375. The controller / processor 375 implements layer 3 and layer 2 functionality. Layer 3 includes a radio resource control (RRC) layer, and layer 2 includes a service data adaptation protocol (SDAP) layer, a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, and a medium access control (MAC) layer. The controller / processor 375 provides RRC layer functionality associated with broadcasting of system information (e.g., MIB, SIBs), RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), inter radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functionality associated with header compression / decompression, security (ciphering, deciphering, integrity protection, integrity verification), and handover support functions; RLC layer functionality associated with the transfer of upper layer packet data units (PDUs), error correction through ARQ, concatenation, segmentation, and reassembly of RLC service data units (SDUs), re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and129025-2565WO01Qualcomm Ref. No. 2501028WO 19 / 53transport channels, multiplexing of MAC SDUs onto transport blocks (TBs), demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.

[0059] The transmit (TX) processors 16 and the receive (RX) processor 370 implement layer 1 functionality associated with various signal processing functions. Layer 1, which includes a physical (PHY) layer, may include error detection on the transport channels, forward error correction (FEC) coding / decoding of the transport channels, interleaving, rate matching, mapping onto physical channels, modulation / demodulation of physical channels, andMIMO antenna processing The TX processor 316 handles mapping to signal constellations based on various modulation schemes (e.g., binary phase-shift keying (BPSK), quadrature phase-shift keying (QPSK), M-phase-shift keying (M-PSK), M-quadrature amplitude modulation (M-QAM)). The coded and modulated symbols may then be split into parallel streams. Each stream may then be mapped to an OFDM subcarrier, multiplexed with a reference signal (e.g., pilot) in the time and / or frequency domain, and then combined together using an Inverse Fast Fourier Transform (IFFT) to produce a physical channel carryingatime domain OFDMsymbol stream. The OFDM stream is spatially precoded to produce multiple spatial streams. Channel estimates from a channel estimator 374 may be used to determine the coding and modulation scheme, as well as for spatial processing. The channel estimate may be derived from a reference signal and / or channel condition feedback transmitted by the UE 350. Each spatial stream may then be provided to a different antenna 320 via a separate transmitter 318Tx. Each transmitter 318Tx may modulate a radio frequency (RF) carrier with a respective spatial stream for transmission.

[0060] At the UE 350, each receiver 354Rx receives a signal through its respective antenna 352. Each receiver 354Rx recovers information modulated onto an RF carrier and provides the information to the receive (RX) processor 356. The TX processor 368 and the RX processor 356 implement layer 1 functionality associated with various signal processing functions. TheRX processor 356 may perform spatial processing on the information to recover any spatial streams destined fortheUE350. If multiple spatial streams are destined for the UE 350, they may be combined by the RX processor 356 into a single OFDM symbol stream. The RX processor 356 then converts the OFDM symbol stream from the time-domain to the frequency domain129025-2565WO01Qualcomm Ref. No. 2501028WO 20 / 53using a Fast Fourier Transform (FFT). The frequency domain signal includes a separate OFDM symbol stream for each subcarrier of the OFDM signal. The symbols on each subcarrier, and the reference signal, are recovered and demodulated by determining the most likely signal constellation points transmitted by the base station 310. These soft decisions may b e based on channel estimates computed by the channel estimator 358. The soft decisions are then decoded and deinterleaved to recover the data and control signals that were originally transmitted by the base station 310 on the physical channel. The data and control signals are then provided to the controller / processor 359, which implements layer 3 and layer 2 functionality.

[0061] The controller / processor 359 can be associated with at least one memory 360 that stores program codes and data. The at least one memory 360 may be referred to as a computer-readable medium. In the UL, the controller / processor 359 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, and control signal processing to recover IP packets. The controller / processor 359 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.

[0062] Similar to the functionality described in connection with the DL transmission by the base station 310, the controller / processor 359 provides RRC layer functionality associated with system information (e.g., MIB, SIBs) acquisition, RRC connections, and measurement reporting; PDCP layer functionality associated with header compression / decompression, and security (ciphering, deciphering, integrity protection, integrity verification); RLC layer functionality associated with the transfer of upper layer PDUs, error correction through ARQ, concatenation, segmentation, and reassembly of RLC SDUs, re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.

[0063] Channel estimates derived by a channel estimator 358 from a reference signal or feedback transmitted by the base station 310 may be used by the TX processor 368 to select the appropriate coding and modulation schemes, and to facilitate spatial processing. The spatial streams generated by the TX processor 368 may be provided129025-2565WO01Qualcomm Ref. No. 2501028WO 21 / 53to different antenna 352 via separate transmitters 354Tx. Each transmitter 354 Tx may modulate an RF carrier with a respective spatial stream for transmission.

[0064] The UL transmission is processed at the base station 310 in a manner similar to that described in connection with the receiver function atthe UE 350. Each receiver 318Rx receives a signal through its respective antenna 320. Each receiver 318Rx recovers information modulated onto an RF carrier and provides the information to a RX processor 370.

[0065] The controller / processor 375 can be associated with at least one memory 376 that stores program codes and data. The at least one memory 376 may be referred to as a computer-readable medium. In the UL, the controller / processor 375 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, control signal processing to recover IP packets. The controller / processor 375 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.

[0066] At least one of the TX processor 368, the RX processor 356, and the controller / processor 359 may be configured to perform aspects in connection with the memory error recover component 198 of FIG. 1.

[0067] At least one of the TX processor 316, the RX processor 370, and the controller / processor 375 may be configured to perform aspects in connection with the memory error recover configuration component 199 of FIG. 1.

[0068] In recentyears, vehicle manufacturers have been developing vehicles with assisted driving and / or autonomous driving capabilities. Assisted driving, which may also be called advanced driver assistance systems (ADAS), may refer to a set of technologies designed to enhance vehicle safety and improve the driving experience by providing assistance and automation to the driver. These technologies may use various sensor(s), such as camera(s), radar(s), light detection and ranging (lidar(s) or lidar sensor(s)), etc., and other components to monitor a vehicle’s surroundings and assist the driver of the vehicle with certain driving tasks. For example, some features of assisted driving systems may include: (1) adaptive cruise control (ACC) (e.g., a system that automatically adjusts a vehicle’s speed to maintain a safe following distance from the vehicle ahead), (2) lane-keeping assist (LKA) (e.g., a system that uses cameras to detect lane markings and helps keep the vehicle centered within the lane, and provides steering inputs to prevent unintentional lane departure), (3), autonomous emergency129025-2565WO01Qualcomm Ref. No. 2501028WO 22 / 53braking (AEB) (e.g., a system that detects potential collisions with obstacles or pedestrians and automatically apply the brakes to avoid or mitigate the impact), (4) blind spot monitoring (BSM) (e.g., a system that uses sensors to detect vehicles in a driver’s blind spots and provides visual or audible alerts to avoid potential collisions during lane changes), (5) parking assistance (e.g., a system that assists drivers in parking their vehicles by using camera(s) and sensor(s) to help with parallel parking or maneuvering into tight spaces), and / or traffic sign recognition (e.g., camera(s) and image processing are used to recognize and display traffic signs such as speed limits, stop signs, and other road regulations on the vehicle’s dashboard).

[0069] Autonomous driving (AD), which may also be referred to as the autonomous driving system (ADS), self-driving, and / or driverless technology, may refer to the ability of a vehicle to navigate and operate itself without specifying human intervention (e.g, travelling from one place to another place without a human controlling the vehicle). The goal of the autonomous drivingisto create vehicles that are capable of perceiving their surroundings, making decisions, and controlling their movements, all without the direct involvement of a human driver. To achieve or improve the autonomous driving, a vehicle may be specified to use a map (or map data) with detailed information, such as a high-definition (HD) map. An HD map may refer to a highly detailed and accurate digital map designed for use in autonomous driving and ADAS. In one example, HD maps may typically include one or more of: (1) geometric information (e.g., preciseroad geometry, including lane boundaries, curvature, slopes, and detailed 3D models of the surrounding environment), (2) lane-level information (e.g., information about individual lanes on the road, such as lane width, lane type (e.g., driving, turning, or parking lanes), and lane connectivity), (3) road attributes (e.g., data on road features like traffic signs, signals, traffic lights, speed limits, and road markings), (4) topology (e.g., information about the relationships between different roads, intersections, and connectivity patterns), (5) static objects (e.g., locations and details of fixed objects alongthe road, such as buildings, traffic barriers, and poles), (6) dynamic objects (e.g., real-time or frequently updated data about moving objects, like other vehicles, pedestrians, and cyclists), and / or (7) localization and positioning: precise reference points and landmarks that help in accurate vehicle localization on the map, etc.129025-2565WO01Qualcomm Ref. No. 2501028WO 23 / 53

[0070] Note while some assisted / autonomous driving systems may demand the use of HD map data, there are also assisted / autonomous driving systems and information systems that may be configured not to use HD map data (e.g., due to costs). For example, the Society of Automotive Engineers (SAE) has defined six levels of driving automation, from Level 0 (no automation) to Level 5 (full automation). For Level 0 (no automation), the human driver may be responsible for all aspects of driving and the system may provide warnings or momentary assistance but does not take control of the vehicle. Example features for SAE Level 0 may include automatic emergency braking, blind spot warnings, and lane departure warnings, etc. As such, SAE Level 0 may not specify using HD map data. For Level 1 (driver assistance), the vehicle may assist with either steering or acceleration / deceleration (but may not perform both simultaneously). The human driver is still responsible for most driving tasks and may need to be ready to take over at any time. Example features for SAE Level 1 may include adaptive cruise control or lane-keeping assistance (e.g., lane centering), etc. For Level 2 (partial automation), the vehicle may control both steering and acceleration / deceleration under certain conditions, but the human driver is requested to remain engaged and monitor the driving environment at all times. Example features for SAE Level 2 may include ADAS, adaptive cruise control and lane-keeping assistance atthe same time, etc. For Level 3 (conditional automation), the vehicle may perform all driving tasks under specific conditions, and the human driver may not be specified to monitor the environment but may need to be ready to take over when requested by the system. Example featuresfor SAE Level 3 may include traffic jam chauffeur, where the vehicle is capable of handling driving in traffic jams without driver intervention. For Level 4 (high automation), the vehicle is capable of handling all driving tasks within certain conditions or environments (geofenced areas). The system may operate without human intervention but may specify a human driver outside its operational domain. Example features for SAE Level 4 may include local driverless taxi and pedals / steering, etc. For Level 5 (full automation), the vehicle is capable of performing all driving tasks under all conditions, and does not specify the human driver at any time. Example features for SAE Level 5 may include fully autonomous vehicles with no steering wheel or pedals. In summary, SAE Level 0 may be defined as features to provide warnings and assistance. ADAS is usually SAE Level 1 and 2, while AD is considered SAE level 3 to 5.129025-2565WO01Qualcomm Ref. No. 2501028WO 24 / 53

[0071] To enable a vehicle to be capable of providing assisted driving and / or autonomous driving, the vehicle may be configured to use various machine learning (ML) and / or neural network (NN) frameworks. An ML / NN framework may refer to a set of tools, libraries, and / or software components that are configured to provide a structured way to design, build, and deploy ML / NN models and applications. These frameworksmay be able to simplify the process of developing ML / NN algorithms and applications by providing a foundation of pre-built functions, algorithms, and utilities. They may typically include features for data preprocessing, model training, evaluation, and / or deployment, etc. ML / NN frameworks may come in various programming languages, and they may be configured to cater to different types of machine learning tasks, including supervised learning, unsupervised learning, and / or reinforcement learning etc. An ML / NN model may refer to a mathematical representation of a real-world process or problem, created using ML / NN algorithms and techniques. These ML / NN models may be configured to make predictions, classify data, and / or solve specific tasks based on patterns and relationships learned from input data. A deep learning framework may refer to a specialized software library or toolset that provides specified components and abstractions for building, training, and deploying deep neural networks. Deep learning frameworks may be designed to facilitate the development of complex neural network models, especially deep neural networks with multiple layers. These frameworks may offer a wide range of pre-implemented layers, optimizers, loss functions, and other components, making it easier for researchers and developers to work with deep learning models.

[0072] FIG. 4 is a diagram 400 illustrating an example of a vehicle performing road object detection using different types of sensors in accordance with various aspects of the present disclosure. In some implementations, a vehicle system may be configured to perform road object detections using multiple types of sensors (and also one or more ML / NN models). For purposes of the present disclosure, a road object or a traffic participant may refer to an object that is related to roads and driving, and is typically / commonly used / considered by the vehicle system in providing assisted driving or performing autonomous driving. In some examples, the road object / traffic participant may also be referred to as a “traffic object” ora “traffic-related object.” For example, a road object / traffic participant may be another vehicle, a pedestrian, a cyclist / bicycle, an animal, a traffic cone, a traffic sign, a traffic light, traffic, a traffic129025-2565WO01Qualcomm Ref. No. 2501028WO 25 / 53lane, a traffic line, a vulnerable road user (VRU), an object that is within a threshold distance of the vehicle, and / or any objects that may typically present on the roads (e.g., on the driving paths of vehicles), etc. On the other hand, a non-road object or a non-traffic participant (which may also be referred to as a non-traffic related object) may refer to an object that is not related to roads and driving, and is typically / commonly notused / consideredby the vehicle system in providing assisted driving or performing autonomous driving. For example, a non-road object / non- traffic participant may be an object that is not within a threshold distance of the vehicle (e.g., a house on the side of the road, a mountain that is far away), an object that is not typically presented on a driving path / road (an airplane, a fire hydrant, a tree, etc.), a structure that is typically not traversed by vehicles (e.g., a pedestrian bridge), etc. An ML / NN model may be trained to identify whether an object is a road object or a non-road object.

[0073] For example, as shown by the diagram 400, a vehicle or a vehicle system (collectively as a “UE 402”) may be configured to use different types of sensors, such as a set of cameras 404 and / or a set of radars 406 for detecting road objects. For purposes of the present disclosure, the term “radar” may broadly refer to a device / componentthatis capable of detecting at least the presence and / or the distance of a physical object. Examples of radar may include an RF radar, a sonar, an ultrasonic sensor, a light detection and ranging (lidar), etc. In some implementations, the UE 402 may also use different MN / NN models for identifying different types of road objects. For example, a first ML / NN model may be trained / used to detect and track polylines from sensor output(s) (e.g., images captured by the camera(s) of the vehicle, point clouds generated from radar(s) / lidar(s), etc.), while a second ML / NN model may be trained / used to detect and track objects in a three-dimensional (3D) space (e.g., to perform 3D object detection (3D0D) tasks). Then, the outputs of different types of sensors (e.g., from the set of cameras 404 and the set of radars 406) may be processed and used by the ADAS or the autonomous driving system (e.g., for assisted / autonomous driving). A point cloud may refer to a discrete set of data points in space, where these points may represent a 3D shape or object. In some implementations, each point position may be associated with a set of Cartesian coordinates (X, Y, Z). Point clouds may be produced by radar(s) / lidar(s) by detecting multiple points on the external surfaces of objects.129025-2565WO01Qualcomm Ref. No. 2501028WO 26 / 53

[0074] For purposes of the present disclosure and in the context of assisted / autonomous driving, a vehicle that is capable of performing autonomous driving and / or certain amount of assisted driving may be referred to as an “ego vehicle” or simply “ego.” For example, an “ego lane” may refer to a lane in which an ego vehicle itself is currently driving. As such, the term “ego” in such context may refer to the vehicle itself, and the term “ego lane” may imply that it is the lane where the (ego) vehicle is actively maneuvering and making decisions. Depending on the context, the term “ego lane” may beused for differentiatingthe lane occupiedby the ego / autonomous vehicle from other lanes on the road.

[0075] Memory (or system memory) may be a fundamental component for most electronic devices (e.g., a UE, a mobile device, an ECU, etc.) responsible for (temporarily) storing and managing data that one or more processors specify to execute instructions efficiently. The memory hierarchy may be designed to balance speed, capacity, and cost, ensuring optimal performance in processing tasks. At its core, memory may be organized into different levels, each level configured to serve a specific role in data storage and retrieval. These levels may include registers, cache memory, main memory (RAM), and secondary storage, etc. with each layer varying in speed and accessibility.

[0076] In the context of memory, registers may refer to the smallest and fastest type of memory, typically located directly inside a central processing unit (CPU). Registers may store data and instructions that are immediately specified for execution. Cache memory (also simply as the “cache”), which is slightly larger but still limited in capacity, may serve as an intermediary between the CPU and the main memory, reducing access time by keeping frequently used data close to the processor. The main memory, or random access memory (RAM), may be configured to hold active programs and data, allowing for fast read and write operations. Unlike storage devices such as hard drives and solid state drives (SSDs), RAM is volatile, meaning its contents are lost when the system powers down.

[0077] Beyond the main memory, secondary storage may provide long-term data retention, ensuring that programs, files, and the operating system persist even when the computer is turned off. Modem systems may also incorporate virtual memory, which extends RAM capacity by using portions of the storage drive to simulate additional memory when physical RAM is insufficient. The efficiency of system memory129025-2565WO01Qualcomm Ref. No. 2501028WO 27 / 53structure may directly impact the performance of a device, as faster memory access is likely to result in smoother operation, reduced latency, and improved multitasking capabilities. For example, memory used in automotive systems / sub-systemsusually specifies high efficiency and / or performance as the memory may be specified to process time-critical tasks (e.g., autonomous / assisted driving related tasks).

[0078] In addition to the aforementioned types of memory, in some systems, there may also be vector tightly coupled memory (VTCM). VTCM may refer to memory that is designed to provide high-speed, low-latency access to important data and instructions, and VTCM may be configured to operate independently of the typical / traditional cache hierarchy. This means that data stored in VTCM may not typically be backed up in lower-level caches like level 2 (L2) or level 3 (L3) caches. Since VTCM does not have a backup storage, the state / system may not be recovered if there are uncorrectable errors in memory regions that have been modified. In addition, VTCM may not have the capability to correct and update memory region if correctable errors are detected by read instructions (such as vector loads). So, a sub-system may commonly be configured to mitigate the errors in two ways: system reset and / or system restart (discussed below).

[0079] FIG. 5 is a diagram 500 illustrating an example of types of errors in caches in accordance with various aspects of the present disclosure. As shown at 510, when cache memory is detected with a single bit error that is correctable, the single bit error may be handled with a single error correction, double error detection (SECDED) error correction code (ECC) correction. ECC may refer to a mechanism / algorithm used in memory to detect and correct errors that occur during data storage or transmission. It may work by adding extra bits to data to create a redundant code that allows errors to be identified and fixed automatically. ECC may be important in systems where data integrity is demanded, such as servers, databases, and mission-driven applications (e.g., vehicle applications). By using ECC, systems may be provided with higher reliability and protect again st data corruption caused by electrical interference, cosmic rays, or other sources of bit corruption, etc. For purposes of the present disclosure, an uncorrectable memory error (or simply as an uncorrectable error in the context of memory) may refer to a type of hardware memory error that cannot be automatically corrected by the ECC mechanism(s). When such an error occurs, it may lead to system crashes, application failures, or data corruption.129025-2565WO01Qualcomm Ref. No. 2501028WO 28 / 53

[0080] As shown at 512, when cache memory is detected with a multi -bit error that is correctable, the multi-bit error may be handled by fetching back-up data from lower- level caches, such as from L2 / L3 caches.

[0081] As shown at 514, if the cache memory is detected with an uncorrectable error, the uncorrectable error may be handled by triggering a system reset (which may also be referred to as and used interchangeably with the “system reboot”) or a subsystem restart. Depending on scenarios, as VTCM is often used for important, high-speed, and / or time sensitive operations, an uncorrectable error may severely impact the system stability and performance. Thus, a system reset may help to reset the state of the memory and clear the error, allowing the system to continue functioning. On the other hand, in some scenarios, a subsystem (e.g., an application digital signal processor (ADSP), a compute digital signal processor (CDSP), a modem processor subsystem (MPSS), and / or an always-on processor (AOP), etc.) restart may be sufficient to correct the error, especially if the error is isolated to a specific part of the system. However, a subsystem restart may depend on the system architecture and the ability to isolate and restart subsystems independently.

[0082] In subsystems with closely coupled memories like VTCM (static random-access memory (SRAM)), L2 caches, instruction translation lookaside buffer (ITLB), data translation lookaside buffer (DTLB), and / or j oint translation lookaside buffer (JTLB) (note JTLB may be a combined concept of both ITLB and DTLB), if an uncorrectable error is encountered, the subsystem may have no other option / way to recover except triggering a sub -system restart or a system reset. However, this may unfold the following problems: (1) bad user experience - an uncorrectable error may disrupt the normal operation of the system, and users may experience crashes, freezes, or unexpected behavior, (2) performance impact - the time spent from recovering the error may directly impact the overall system performance, where latency may increase, and important tasks may be delayed, and / or (3) power impact- the energy specified to bringthe sub system back to a consistent state may be substantial in some cases, and frequent uncorrectable errors triggering subsystem restarts may lead to increased power consumption, affecting battery life in mobile devices or energy efficiency in data centers.

[0083] Aspects presented herein may improve the overall performance and efficiency related to recovering a system / sub system from uncorrectable memory error(s), such as a129025-2565WO01Qualcomm Ref. No. 2501028WO 29 / 53system / sub system with resource constraints. In one aspect of the present disclosure, an interrupt-based mechanism is provided to handle uncorrectable memory errors) in a more efficient manner. When an uncorrectable error in memory, such as the VTCM, is encountered, the error instance is logged in the error logger region of the memory, and an interrupt is raised to the subsystem or a high-level operation system (HLOS) (e.g., an application processor) in order to take the specified action(s).

[0084] The interrupt-based mechanism described herein may include one or more of the following:(1) Victim process domain identification. For purposes of the present disclosure, a process domain (PD) may refer to an isolated execution environment within a system where a process is configured to operate with its own allocated resources, which may include memory, central processing unit (CPU) time, and / or system privileges, etc. Such configuration may provide better security, stability, and / or resource management, etc. In a system, a process domainmay ensure that different applications and system processes run independently without interference. In a subsystem, a process domain may be defined within a specialized component (e.g., modem, DSP, secure enclave, etc.) to manage tasks efficiently and securely. In one aspect, the interrupt-based mechanism proposed herein may be configured to identify the process domain identification / identifier (ID) based on software (SW) thread ID programmed in the hardware (HW) thread / core on which a fault / error (e.g., an ECC fault / error) is encountered. For illustrative purposes, a process domain that is detected / identified with the fault / error may be referred to as the “victim process domain.”(2) Victim resource identification. In another aspect, the interrupt-based mechanism proposed herein may be configured to include an additional set of registers (or use a set of reserved bits) in the ECC error logger register in memory / VTCM (of the subsystem), to keep the information of the faulty cache line / region (referringto the cache line / region that is associated with the fault / error). The faulty cache line / region information may be used by the software to identify the owner of this cache line / region or the victim PD. For illustrative purposes, resource(s) (e.g, memory) that are detected / identified with the fault / error (e.g., the faulty cache line / region) may collectively be referred to as the “victim resource(s).”129025-2565WO01Qualcomm Ref. No. 2501028WO 30 / 53(3) Interrupt to abort process. In another aspect, the interrupt-based mechanism proposed herein may be configured to interrupt a set of HW threads which are (already) executing the processes that belongs to the victim PD. This may be specified to make sure noHW thread is executing a process from the faulty cache line / region until the recovery is done.(4) Faulty cache line / region (e.g., victim resource(s)) restriction. In another aspect, the interrupt-based mechanism proposed herein may be configured to restrict or make the faulty cache line / region un-usable by other PD(s) for a specified period of time. For example, the interrupt-based mechanism may be configured to lock the faulty cache line / region (e.g., by settingthem with invalid bit(s)) for restricting the faulty cache line / region. Unless, the number of faults exceeds a threshold count or the subsystem is restarted, the faulty cache line / region is to be restricted from usage by other PD(s).(5) Interrupt for threshold uncorrectable errors. In another aspect, the interrupt-based mechanism proposed herein maybe configured to include an interrupt register in the error logger region to set the threshold for the number of uncorrectable errors and raise an interrupt to a HLOS / sub system (e.g., a safety island (SAIL) subsystem in automotive) to take further action(s). The further action(s) may be implementation specific, but a suitable / recommended approach may be to trigger a system restart.(6) Victim PD restart. In another aspect, the interrupt-based mechanism proposed herein (or an associated software) may be configured to trigger a PD restart based on the error logger information of the cache, to recover from the error. This may ensure that, when the PD is loaded to cache memory, the PD does not get the faulty line / region allocated (e.g., since this may be done at (3) above already).

[0085] FIG. 6 is a diagram 600 illustrating an example of an interrupt-based mechanism that is capable of handling uncorrectable memory error(s) in a more efficient manner in accordance with various aspects of the present disclosure. A device or at least one processor / memory of the device (collectively as the “device 602” hereafter) may be configured to monitor, detect, and / or identify one or more memories of the device 602, such as the VTCM, for uncorrectable error(s).

[0086] At shown at 610, when at least one uncorrectable error is detected / identified (e.g., in memory / cache of a subsystem), the device 602 may be configured to identify the129025-2565WO01Qualcomm Ref. No. 2501028WO 31 / 53PD(s) and the ID(s) of the PD(s) (referring to as the PD ID(s) hereafter) on which the at least one uncorrectable error is detected / identified. The PD ID(s) may be based on the SW thread ID programmed in the HW thread / core. In other word, the device may be configured identify the victim PD(s) after at least one uncorrectable error is detected / identified. For example, depending on implementations, anECC hardware (or an ECC error logger module) may be configured to include a set of registers (or use a set of reserved bits) to record / keep the information of the cache line / region (i.e., the victim resource(s)) that is associated with the at least one uncorrectable error. The faulty cache line / region information may be used by the software to identify the owner of the cache line / region or the victim PD.

[0087] As shown at 612, after the victim PD(s) and victim resource(s) are identified, the device 602 may be configured to interrupt a set of HW threads which are (already) executingthe processes thatbelongs to the victim PD(s) and / or the victim resource(s), such as by sending an interrupt (INTR) signal to the set of HW threads.

[0088] As shown at 614, the device 602 may also be configured to restrict or make the victim resource(s) un-usable by other PD(s) for a specified period of time. For example, the device 602 may be configured to lock the victim resource(s) (e.g., the faulty cache line / region) by setting them with invalid bit(s). Unless, the number of uncorrectable errors exceeds a threshold count (discussed below) or the subsystem is restarted, the victim resource(s) is to be restricted from usage by other PD(s). For example, the device 602 may be configuredto include an interrupt register in the error logger region to set the threshold for the number of uncorrectable errors and raise an interruption to a HLOS / sub system (or SAIL) (e.g., to take further specified action(s)).

[0089] As shown at 616, the device 602 may be configured to trigger a PD restart (i.e., a subsystem restart) based on the error logger information of the cache, to recover from the error. However, as shown at 618, if the number of uncorrectable errors exceeds a threshold count (discussed below), the device 602 may be configured to initiate a system reset.

[0090] FIG. 7 is a flow chart 700 illustrating an example of the interrupt-based mechanism with an uncorrectable error counter in accordance with various aspects of the present disclosure. As shown at 710, a set of processors of the device 602 may be configured to monitor, detect, and / or identify the cache of the device 602 for memory related error(s), such as described in connection with FIG. 5.129025-2565WO01Qualcomm Ref. No. 2501028WO 32 / 53

[0091] As shown at 712, when a memory error is detected, the device 602 may be configured to identify whether the detected memory error is a correctable error or an uncorrectable error. As discussed above, a correctable error may refer to an error that may be corrected automatically by the ECC, and an uncorrectable error may refer to an error that may not be corrected automatically by the ECC.

[0092] As shown at 714, if the detected memory error is a correctable error, the device 602 may be configured to correct the error, such as based on using the SECDED ECC correction and / or fetching backed-up data from lower-level caches as described in connection with 510 and 512 of FIG. 5. Then, as shown at 716, the device 602 may enable the system to continue to operate after the detected memory error is corrected.

[0093] On the other hand, as shown at 718, if the detected memory error is an uncorrectable error, the device 602 may save the save the faulty cache / memory information in an error logger. Then, as shown at 719, the device 602 may identify the victim PD(s) and the victim resource(s), such as described in connection with 610 of FIG. 6. Then, as shown at 720, the device 602 may interrupt (e.g., halt or abort) a set of HW threads which are (already) executing the processes that belongs to the victim PD(s) and / or the victim resource(s). The device 602 may restart or resume the set of HW threads after the system reset, the subsystem reset, the restart of a set of PDs, or the restart of a PD, etc.

[0094] As shown at 722, the device 602 may also include an uncorrectable error counter, where the uncorrectable error counter is configured to be incremented when uncorrectable error(s) are detected / identified. For example, the uncorrectable error counter may be configured to be incremented by one (1) when an uncorrectable error is detected / identified, and incremented by one (1) again when a subsequent uncorrectable error is detected / identified. In some configurations, if multiple uncorrectable error(s) are detected, the uncorrectable error counter may be configured to be incremented by the correspondingnumber of uncorrectable errors detected (e.g, incremented by two if two uncorrectable errors are detected, and incrementedby three if three uncorrectable errors are detected, etc.).

[0095] As shown at 724, after the uncorrectable error counter is incremented based on the detected uncorrectable error(s), the device 602 may determine whether the uncorrectable error counter exceeds a defined threshold (which may be referred to as a “counter threshold”).129025-2565WO01Qualcomm Ref. No. 2501028WO 33 / 53

[0096] As shown at 725, if the uncorrectable error counter does not exceed the defined threshold, the device 602 may restrict or make the victim resource(s) un-usable by other PD(s) for a specified period of time. In addition, as shown at 726, the device 602 may raise an interruption to a HLOS / sub system (e.g., SAIL) to take further action(s), such as restarting the victim PD or the subsystem.

[0097] On the other hand, as shown at 728, if the uncorrectable error counter exceeds the defined threshold, the device 602 may interrupt the subsystem to take defined action(s), and / or trigger a system reset.

[0098] FIG. 8 is a communication flow 800 illustrating an example system programming sequence forthe interrupt-basedmechanism in accordance with various aspects of the present disclosure. The system programming sequence described herein may occur between various modules / entities of the device 602.

[0099] At 810, aHW module 802 (e.g., a processor, a subsystem, etc.) may be configured to initialize an ECC module 804, such for providing ECC for a set of memories. As shown at 812 and 814, the HW module 802 and the ECC module 804 may also be configured to clear the error loggers, and enable an uncorrectable error counter (i.e., the uncorrectable error counter discussed in connection with 722 and 724 of FIG. 7).

[0100] At 816, after the ECC module 804 is initialized, the ECC module 804 may install a handler at an operating system (OS) 806, such as a high-level operating system (HLOS) and / or a real-time operating system (RTOS), etc. In the context of operating systems like HLOS and RTOS, a handler may refer to a routine or a function responsible for managing specific events or interrupts. Handlers may be important for handling tasks such as processing interrupts from hardware devices, managing exceptions, and / or responding to specific software requests, etc. In an RTOS, for instance, handlers may be configured to handle time-sensitive tasks with precise timing specifications, ensuring that interrupts are processed promptly and deterministically. Handlers may play a vital role in maintaining system stability and responsiveness, especially in environments where real-time performance is demanded, such as in automotive systems (e.g., as describedin connection with FIG.4).

[0101] At 818, the OS 806 may start a process domain (PD) 808 (which may be onePD ora plurality of PDs). At 820, after the PD 808 is started, the PD 808 may be configured to register a SECDED interrupt request (IRQ) with the OS 806. Registering a129025-2565WO01Qualcomm Ref. No. 2501028WO 34 / 53SECDED IRQ may refer to setting up an IRQ handler for SECDED error events in a system that implements ECC mechanisms. Registering a SECDED IRQ may enable the system to handle ECC-related faults by detecting memory corruption, correcting single-bit errors before they cause data corruption, and / or logging errors and potentially triggering fail-safe mechanisms in safety -critical applications, etc.

[0102] At 822 and 824, if uncorrectable error(s) occurred at the PD 808, the PD 808 may generate and send an IRQ to the OS 806.

[0103] At 826, based on the IRQ from the PD 808, the OS 806 may run the installed handler.Then, as shown at 828, the OS 806 may obtain the process identifier (PID) and cache line information from the ECC module 804. PID may be a unique number assigned by the operating system (e.g., the OS 806) to each running process. PID may be used to track and manage processes in multitasking environments.

[0104] At 830, the OS 806 may identify the PID from the error logger registers, and at 832, the OS 806 may restrict a set of victim resources, such as describedin connection with 614 of FIG. 6. Then, at 834, the OS 806 may initiate a restart for the PD 808, such as described in connection with 616 of FIG. 6 and / or 726 of FIG. 7.

[0105] Aspects presented herein may be suitable for advanced driver assistance systems (ADAS) (e.g., as described in connection with FIG. 4). For example, inconsistent power supply may cause transient errors in the cache. If an object detection algorithm fails, aPD restartmay recoverthis specific function while keeping mutually exclusive ADAS features operational which may be dependent on the same subsystem. Aspects presented herein may also be suitable for infotainment systems. For example, automotive environments may experience extreme temperatures, which may affect the stability and reliability of cache memory. If a navigation system encounters an error, a PD restart may reset just the navigation module without affecting the entire infotainment system, ensuring continuous music playback and other functions. Aspects presented herein may also be suitable foruncorrectable error recovery in low- power use scenarios / cases. For example, in case of low-power modes like idle mode power saving (IMPS) or island modes, there may be no way to recover from the uncorrectable errors and the system may crash. In these cases, restricting the memory resource may help the subsystem to recover from the errors by doing a PD restart.

[0106] FIG. 9 is a flowchart 900 of a method of recovering system from uncorrectable memory error(s). The method may be performed by a device (e.g., theUE 104, 402;129025-2565WO01Qualcomm Ref. No. 2501028WO 35 / 53the device 602; the apparatus 1104). The method may enable the device to use an interrupt-based mechanism to handle uncorrectable memory error(s) in a more efficient manner.

[0107] At 902, the device may detect an uncorrectable error in memory of a PD, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 610 of FIG. 6, a device or at least one processor / memory of the device (collectively as the “device 602” hereafter) may be configured to monitor, detect, and / or identify one or more memories of the device 602, such as the VTCM, for uncorrectable error(s). When at least one uncorrectable error is detected / identified (e.g., at memory / cache of a subsystem), the device 602 may be configured to identify the PD(s) and the ID(s) of the PD(s) (referring to as the PD ID(s) hereafter) on which the at least one uncorrectable error is detected / identified. The detection of the uncorrectable error in memory may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0108] In one example, to detect the uncorrectable error in the memory of thePD, the device may be configured to identify an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and obtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

[0109] At 910, the device may increment an uncorrectable error counter based on detection of the uncorrectable error in the memory, such as described in connection with FIGs.6 to 8. For example, as discussed in connection with 722 of FIG. 7, the device 602 may also include an uncorrectable error counter, where the uncorrectable error counter is configured to be incremented when uncorrectable error(s) are detected / identified. The incrementation of the uncorrectable error counter may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceivers) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0110] At 912, the device may perform, based on the uncorrectable error counter, at least one of : (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable129025-2565WO01Qualcomm Ref. No. 2501028WO 36 / 53error counter exceeds the counter threshold, or (3) issuinga restartforthe PD if the uncorrectable error counter does not exceed the counter threshold, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 724 of FIG. 7, after the uncorrectable error counter is incremented based on the detected uncorrectable error(s), the device 602 may determine whether the uncorrectable error counter exceeds a defined threshold (which may be referred to as a “counter threshold”). As shown at 725, if the uncorrectable error counter does not exceed the defined threshold, the device 602 may restrict or make the victim resource(s) unusable by otherPD(s) for a specified period of time. In addition, as shown at 726, the device 602 may raise an interruption to a HLOS / sub system (e.g., SAIL) to take further action(s), such as restarting the victim PD or the subsystem. On the other hand, as shown at 728, if the uncorrectable error counter exceeds the defined threshold, the device 602 may interrupt the subsystem to take defined action(s), and / or trigger a system reset. The triggering ofthe system reset or a subsystem resetand / orthe issuing of a restart for a set of PDs may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 ofthe apparatus 1104 in FIG. 11.

[0111] In one example, to issue the restart for the PD further, the device may be configured to interrupt a HLOS or a subsystem.

[0112] In another example, the system reset corresponds to a SoC system reset and the subsystem reset corresponds to a DSP subsystem reset or a micro-controller subsystem reset.

[0113] In another example, the memory corresponds to at least one SRAM or at least one VTCM.

[0114] In another example, the memory is protected with an ECC in the process domain, and the uncorrectable error corresponds to an ECC error.

[0115] In another example, the PD is associated with an ADAS system, an infotainment system, or a system operating below a power threshold or a capability threshold.

[0116] In another example, the device may identify an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and obtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable129025-2565WO01Qualcomm Ref. No. 2501028WO 37 / 53error in the PD, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 610 of FIG. 6, when at least one uncorrectable error is detected / identified (e.g., at memory / cache of a subsystem), the device 602 may be configured to identify the PD(s) and the ID(s) of the PD(s) (referring to as the PD ID(s) hereafter) on which the at least one uncorrectable error is detected / identified. The PD ID(s) may be based on the SW thread ID programmed in the HW thread / core. The identification of the ID of the PD may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0117] In another example, the device may halt or abort a set of hardware threads that is executing a set of processes related to the PD (and restart or resume the set of hardware threads after the system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD), such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 720 of FIG. 7, the device 602 may interrupt (e.g., halt or abort) a set of HW threads which are (already) executing the processes that belongs to the victim PD(s) and / or the victim resource(s). The device 602 may restart or resume the set of HW threads after the system reset, the subsystem reset, the restart of a set of PDs, or the restart of a PD, etc. The halting or aborting of the set of hardware threads may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0118] In another example, the device may mark a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non- usable b ased on the detection of the uncorrectable error in the memory, where memory lines or memory regions of the memory that are notmarked as restricted or non-usable are accessible and usable by the PD or other PDs, such as describedin connection with FIGs. 6 to 8. For example, as discussed in connection with 614 of FIG. 6, the device 602 may also be configured to restrict or make the victim resource(s) un-usable by other PD(s) fora specified period of time. For example, the device 602 may be configured to lockthe victim resource(s) (e.g., the faulty cache line / region) by setting them with invalid bit(s). The marking of the set of memory lines or the set of memory129025-2565WO01Qualcomm Ref. No. 2501028WO 38 / 53regions may be performed by, e.g., the memory error recover component 198, the one ormoreECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11. In some implementations, the device may output the set of memory lines or the set of memory regions to an error logger. In some implementations, the device may refrain from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, or prevent or disable the PD or the other PDs to access or use the set of memory lines or the set of memory regions.

[0119] FIG. 10 is a flowchart 1000 of a method of recovering system from uncorrectable memory error(s). The method may be performed by a device (e.g., the UE 104, 402; the device 602; the apparatus 1104). The method may enable the device to use an interrupt-based mechanism to handle uncorrectable memory error(s) in a more efficient manner.

[0120] At 1002, the device may detect an uncorrectable error in memory of a PD, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 610 of FIG. 6, a device or at least one processor / memory of the device (collectively as the “device 602” hereafter) may be configured to monitor, detect, and / or identify one or more memories of the device 602, such as the VTCM, for uncorrectable error(s). When at least one uncorrectable error is detected / identified (e.g., at memory / cache of a subsystem), the device 602 may be configured to identify the PD(s) and the ID(s) of the PD(s) (referring to as the PD ID(s) hereafter) on which the at least one uncorrectable error is detected / identified. The detection of the uncorrectable error in memory may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0121] In one example, to detect the uncorrectable error in the memory of thePD, the device may be configured to identify an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and obtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

[0122] At 1010, the device may increment an uncorrectable error counter based on detection of the uncorrectable error in the memory, such as described in connection with FIGs.129025-2565WO01Qualcomm Ref. No. 2501028WO 39 / 536 to 8. For example, as discussed in connection with 722 of FIG. 7, the device 602 may also include anuncorrectable error counter, where the uncorrectable error counter is configured to be incremented when uncorrectable error(s) are detected / identified. The incrementation of the uncorrectable error counter may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceivers) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0123] At 1012, the device may perform, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuing a restart for the PD if the uncorrectable error counter does not exceed the counter threshold, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 724 of FIG. 7, after the uncorrectable error counter is incremented based on the detected uncorrectable error(s), the device 602 may determine whether the uncorrectable error counter exceeds a defined threshold (which may be referred to as a “counter threshold”). As shown at 725, if the uncorrectable error counter does not exceed the defined threshold, the device 602 may restrict or make the victim resource(s) un-usable by other PD(s) for a specified period of time. In addition, as shown at 726, the device 602 may raise an interruption to a HLOS / sub system (e.g, SAIL) to take further action(s), such as restarting the victim PD or the sub system. On the other hand, as shown at728, if the uncorrectable error counter exceeds the defined threshold, the device 602 may interrupt the subsystem to take defined action(s), and / or trigger a system reset. The triggering of the system reset or a subsystem reset and / or the issuing of a restart fora set of PDs may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0124] In one example, to issue the restart for the PD further, the device may be configured to interrupt a HLOS or a subsystem.

[0125] In another example, the system reset corresponds to a SoC system reset and the subsystem reset corresponds to a DSP subsystem reset or a micro-controller subsystem reset.129025-2565WO01Qualcomm Ref. No. 2501028WO 40 / 53

[0126] In another example, the memory corresponds to at least one SRAM or at least one VTCM.

[0127] In another example, the memory is protected with anECC in the process domain, and the uncorrectable error corresponds to an ECC error.

[0128] In another example, the PD is associated with an ADAS system, an infotainment system, or a system operating below a power threshold or a capability threshold.

[0129] In another example, as shown at 1004, the device may identify an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and obtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 610 of FIG. 6, when at least one uncorrectable error is detected / identified (e.g., at memory / cache of a subsystem), the device 602 may be configured to identify the PD(s) and the ID(s) of the PD(s) (referring to as the PD ID(s) hereafter) on which the at least one uncorrectable error is detected / identified. The PD ID(s) may be based on the SW thread ID programmed in the HW thread / core. The identification of the ID of the PD may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11.

[0130] In another example, as shown at 1006, the device may halt or abort a set of hardware threads that is executing a set of processes related to the PD (and restart or resume the set of hardware threads after the system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD), such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 720 of FIG. 7, the device 602 may interrupt (e.g., halt or abort) a set of HW threads which are (already) executingthe processes that belongs to the victim PD(s) and / or the victim resource(s). The device 602 may restart or resume the set of HW threads after the system reset, the subsystem reset, the restart of a set of PDs, or the restart of a PD, etc. The halting or aborting of the set of hardware threads may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / orthe application processor(s) 1106 of the apparatus 1104 in FIG. 11.129025-2565WO01Qualcomm Ref. No. 2501028WO 41 / 53

[0131] In another example, as shown at 1008, the device may mark a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non-usable based on the detection of the uncorrectable error in the memory, where memory lines or memory regions of the memory that are not marked as restricted or non-usable are accessible and usable by the PD or other PDs, such as described in connection with FIGs. 6 to 8. For example, as discussed in connection with 614 of FIG. 6, the device 602 may also be configured to restrict or make the victim resource(s) un-usable by other PD(s) for a specified period of time. For example, the device 602 may be configured to lock the victim resource(s) (e.g., the faulty cache line / region) by setting them with invalid bit(s). The marking of the set of memory lines or the set of memory regions may be performed by, e.g., the memory error recover component 198, the one or more ECUs 1134, the transceiver(s) 1122, the cellular baseband processor(s) 1124, and / or the application processor(s) 1106 of the apparatus 1104 in FIG. 11. In some implementations, the device may output the set of memory lines or the set of memory regions to an error logger. In some implementations, the device may refrain from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, or prevent or disable the PD or the other PDs to access or use the set of memory lines or the set of memory regions.

[0132] FIG. 11 is a diagram 1100 illustrating an example of a hardware implementation for an apparatus 1104. The apparatus 1104 may be a UE, a component of a UE, or may implement UE functionality. In some aspects, the apparatus 1104 may include at least one cellular baseband processor 1124 (also referred to as a modem) coupled to one or more transceivers 1122 (e.g., cellular RF transceiver). The cellular baseband processor(s) 1124 may include at least one on-chip memory 1124'. In some aspects, the apparatus 1104 may further include one or more subscriber identity modules (SIM) cards 1120 and at least one application processor 1106 coupled to a secure digital (SD) card 1108 and a screen 1110. The application processor(s) 1106 may include on-chip memory 1106'. In some aspects, the apparatus 1104 may further include a Bluetooth module 1112, a WLAN module 1114, an ultrawide band (UWB) module 1138, an SPS module 1116 (e.g., GNSS module), one or more sensors 1118 (e.g., barometric pressure sensor / altimeter; motion sensor such as inertial measurement unit (IMU), gyroscope, and / or accelerometer(s); light detection and129025-2565WO01Qualcomm Ref. No. 2501028WO 42 / 53ranging (LIDAR), radio assisted detection and ranging (RADAR), sound navigation and ranging (SONAR), magnetometer, audio and / or other technologies used for positioning), additional memory modules 1126, a power supply 1130, a camera 1132, and / or one or more electronic control units (ECUs) 1134. The Bluetooth module 1112, the UWB module 1138, the WLAN module 1114, and the SPS module 1116 may include an on-chip transceiver (TRX) (or in some cases, just a receiver (RX)). The Bluetooth module 1112, the WLAN module 1114, and the SPS module 1116 may include their own dedicated antennas and / or utilize the antennas 1180 for communication. The cellular baseband processor(s) 1124 communicates through the transceiver(s) 1122 via one or more antennas 1180 with the UE 104 and / or with an RU associated with a network entity 1102. The cellular baseband processor(s) 1124 and the application processor(s) 1106 may each include a computer-readable medium / memory 1124', 1106', respectively. The additional memory modules 1126 may also be considered a computer-readable medium / memory. Each computer-readable medium / memory 1124', 1106', 1126 may be non-transitory. The cellular baseband processor(s) 1124 and the application processor(s) 1106 are each responsible for general processing, including the execution of software stored on the computer-readable medium / memory. The software, when executed by the cellular baseband processor(s) 1124 / application processor(s) 1106, causes the cellular baseband processor(s) 1124 / application processor(s) 1106 to perform the various functions described supra. The cellular baseband processor(s) 1124 and the application processor(s) 1106 are configured to perform the various functions described supra based at least in part of the information stored in the memory. That is, the cellular baseband processor(s) 1124 and the application processor(s) 1106 maybe configured to perform a first subset of the various functions described supra without information stored in the memory and may be configured to perform a second sub set of the various functions described supra based on the information stored in the memory. The computer-readable medium / memory may also be used for storing data that is manipulated by the cellular baseband processor(s) 1124 / application processors) 1106 when executing software. The cellular baseband processor(s) 1124 / application processor(s) 1106 may be a component of the UE 350 and may include the at least one memory 360 and / or at least one of the TX processor 368, the RX processor 356, and the controller / processor 359. In one configuration, the apparatus 1104 may be at129025-2565WO01Qualcomm Ref. No. 2501028WO 43 / 53least one processor chip (modem and / or application) and include just the cellular baseband processor(s) 1124 and / or the application processor(s) 1106, and in another configuration, the apparatus 1104 may be the entire UE (e.g., see UE 350 of FIG. 3) and include the additional modules of the apparatus 1104.

[0133] As discussed supra, the memory error recover component 198 may be configured to detect an uncorrectable error in memory of a PD. The memory error recover component 198 may also be configured to increment an uncorrectable error counter based on detection of the uncorrectable error in the memory. The memory error recover component 198 may also be configured to perform, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or(3)issuinga re start for the PD if the uncorrectable error counter does not exceed the counter threshold. The memory error recover component 198 may be within the cellular baseband processor(s) 1124, the application processor(s) 1106, or both the cellular baseband processor(s) 1124 and the application processor(s) 1106. The memory error recover component 198 may be one or more hardware components specifically configured to carry out the stated processes / algorithm, implemented by one or more processors configured to perform the stated processes / algorithm, stored within a computer-readable medium for implementation by one or more processors, or some combination thereof. When multiple processors are implemented, the multiple processors may perform the stated processes / algorithm individually or in combination. As shown, the apparatus 1104 may include a variety of components configured for various functions. In one configuration, the apparatus 1104, and in particular the cellular baseband processor(s) 1124 and / or the application processors) 1106, may include means for detecting an uncorrectable error in memory of a PD. The apparatus 1104 may further include means for incrementing an uncorrectable error counter based on detection of the uncorrectable error in the memory. The apparatus 1104 may further include means for performing, based on the uncorrectable error counter, atleastone of: (l)triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuing129025-2565WO01Qualcomm Ref. No. 2501028WO 44 / 53a restart for the PD if the uncorrectable error counter does not exceed the counter threshold.

[0134] In one configuration, the means for detecting the uncorrectable error in the memory of the PD may include configuring the apparatus 1104 to identify an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and obtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

[0135] In another configuration, the means for performing issuing the restart for the PD may further include configuring the apparatus 1104 to interrupt a HLOS or a subsystem.

[0136] In another configuration, the system reset corresponds to a SoC system reset and the subsystem reset corresponds to a DSP subsystem reset or a micro-controller subsystem reset.

[0137] In another configuration, the memory corresponds to at least one SRAM or at least one VTCM.

[0138] In another configuration, the memory is protected with an ECC in the process domain, and the uncorrectable error corresponds to an ECC error.

[0139] In another configuration, thePD is associated with an ADAS system, an infotainment system, or a system operating below a power threshold or a capability threshold.

[0140] In another configuration, the apparatus 1104 may further include means for identifying an ID of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected, and means for obtaining, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

[0141] In another configuration, the apparatus 1104 may further include means for halting or means f or aborting a set of hardware threads that is executing a set of processes related to thePD, and means for restarting or means for resuming the set of hardware threads after the system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD.

[0142] In another configuration, the apparatus 1104 may further include means for marking a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non-usable based on the detection of the uncorrectable error in the memory, where memory lines or memory regions of the129025-2565WO01Qualcomm Ref. No. 2501028WO 45 / 53memory that are not marked as restricted or non-usable are accessible and usable by the PD or other PDs. In some implementations, the apparatus 1104 may further include means for outputting the set of memory lines or the set of memory regions to an error logger. In some implementations, the apparatus 1104 may further include means for refraining from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, or means for preventing or means for disabling the PD or the other PDs to access or use the set of memory lines or the set of memory regions.

[0143] The means may be the memory error recover component 198 of the apparatus 1104 configured to perform the functions recited by the means. As described supra, the apparatus 1104 may include the TX processor 368, the RX processor 356, and the controller / processor 359. As such, in one configuration, the means may be the TX processor 368, the RX processor 356, and / or the controller / processor 359 configured to perform the functions recited by the means.

[0144] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of example approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts maybe rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not limited to the specific order or hierarchy presented.

[0145] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not limited to the aspects described herein, but are to be accorded the full scope consistent with the language claims. Reference to an element in the singular does not mean “one and only one” unless specifically so stated, but rather “one or more.” Terms such as “if,” “when,” and “while” do not imply an immediate temporal relationship or reaction. That is, these phrases, e.g., “when,” do notimply an immediate action in response to or during the occurrence of an action, but simply imply that if a condition is met then an action will occur, butwithoutrequiringa specific or immediate time constraintforthe action to occur. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not129025-2565WO01Qualcomm Ref. No. 2501028WO 46 / 53necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. Sets should be interpreted as a set of elements where the elements number one or more. Accordingly, for a set of X, X would include one or more elements. When at least one processor (i.e., a set of one or more processors P) is configured to perform a set of functions F, each processor of P may be configured to perform a subset S of F, where S £ F. Accordingly, each processor of the at least one processor may be configured to perform a particular subset of the set of functions, where the subset is the full set, a proper subset of the set, or an empty subset of the set. A processor may be referred to as processor circuitry. A memory / memory module may be referred to as memory circuitry. If a first apparatus receives datafrom ortransmits data to a second apparatus, the data may be received / transmitted directly between the first and second apparatuses, or indirectly between the first and second apparatuses through a set of apparatuses. A device configured to “output” data or “provide” data, such as a transmission, signal, or message, may transmit the data, for example with a transceiver, or may send the data to a device that transmits the data. A device configured to “obtain” data, such as a transmission, signal, or message, may receive, for example with a transceiver, or may obtain the data from a device that receives the data. Information stored in a memory includes instructions and / or data. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are encompassed by the claims. Moreover, nothing disclosed herein is dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The words “module,” “mechanism,” “element,” “device,” and the like may notbe a substitute for the word129025-2565WO01Qualcomm Ref. No. 2501028WO 47 / 53“means.” As such, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for.”

[0146] As used herein, the phrase “based on” shall not be construed as a reference to a closed set of information, one or more conditions, one or more factors, or the like. In other words, the phrase “based on A” (where “A” may be information, a condition, a factor, or the like) shall be construed as “based at least on A” unless specifically recited differently.

[0147] The following aspects are illustrative only and may be combined with other aspects or teachings described herein, without limitation.

[0148] Aspect 1 is a method at a device, comprising: detecting an uncorrectable error in memory of aprocess domain (PD); incrementing an uncorrectable error counter based on detection of the uncorrectable error in the memory; and performing, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or(3)issuinga re start for the PD if the uncorrectable error counter does not exceed the counter threshold.

[0149] Aspect 2 is the method of aspect 1, wherein detecting the uncorrectable error in the memory of the PD comprises: identifying an identification (ID) of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected; and obtaining, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

[0150] Aspect 3 is the method of aspect 1 or aspect 2, further comprising: halting or aborting a set of hardware threads that is executing a set of processes related to the PD; and restarting or resumingthe set of hardware threads afterthe system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD.

[0151] Aspect 4 is the method of any of aspects 1 to 3, further comprising: marking a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non-usable based on the detection of the uncorrectable error in the memory, wherein memory lines or memory regions of the memory that are not marked as restricted or non-usable are accessible and usable by the PD or other PDs.129025-2565WO01Qualcomm Ref. No. 2501028WO 48 / 53

[0152] Aspect 5 is the method of any of aspects 1 to 4, further comprising: outputting the set of memory lines or the set of memory regions to an error logger.

[0153] Aspect 6 is the method of any of aspects 1 to 5, further comprising: refraining from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, or preventing or disabling the PD or the other PDs to access or use the set of memory lines or the set of memory regions.

[0154] Aspect? is the method of any of aspects 1 to 6, wherein issuing the restart for the PD further comprises: interrupting a high-level op eration system (HLOS)ora subsystem.

[0155] Aspect 8 is the method of any of aspects 1 to 7, wherein the system reset corresponds to a system-on-chip (SoC) system reset and the subsystem reset corresponds to a digital signal processors (DSP) subsystem reset or a micro-controller subsy stem reset.

[0156] Aspect 9 is the method of any of aspects 1 to 8, wherein the memory corresponds to at least one static random-access memory (SRAM) or at least one vector tightly coupled memory (VTCM).

[0157] Aspect 10 is the method of any of aspects 1 to 9, wherein the memory is protected with an error correction code (ECC) in the process domain, and wherein the uncorrectable error corresponds to an ECC error.

[0158] Aspect 11 is the method of any of aspects 1 to 10, wherein the PD is associated with an advanced driver assistance systems (ADAS) system, an infotainment system, or a system operating below a power threshold or a capability threshold.

[0159] Aspect 12 is an apparatus at a device, including: at least one memory; and at least one processor coupledto the atleast one memory and, based atleastin part on information stored in the at least one memory, the at least one processor is configured to implement any of aspects 1 to 11.

[0160] Aspect 13 is the apparatus of aspect 12, further including at least one transceiver coupled to the at least one processor.

[0161] Aspect 14 is an apparatus at a device, including means for implementing any of aspects 1 to 11.

[0162] Aspect 15 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer executable code, where the code when executed by a processor causes the processor to implement any of aspects 1 to 11.129025-2565WO01

Claims

Qualcomm Ref. No. 2501028WO 49 / 53CLAIMS WHAT IS CLAIMED IS:

1. An apparatus at a device, comprising:at least one memory; andat least one processor coupled to the at least one memory, wherein the at least one processor is configured to:detect an uncorrectable error in memory of a process domain (PD); increment an uncorrectable error counter based on detection of the uncorrectable error in the memory; andperform, based on the uncorrectable error counter, at least one of: (1) triggering a system resetora subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuing a restartforthePD if the uncorrectable error counter does not exceedthe counter threshold.

2. The apparatus of claim 1 , wherein to detect the uncorrectable error in the memory of the PD, the at least one processor is configured to:identify an identification (ID) of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected; andobtain, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

3. The apparatus of claim 1 , wherein the at least one processor is further configured to:halt or abort a set of hardware threads that is executing a set of processes related to the PD; andrestart or resume the set of hardware threads after the system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD.

4. The apparatus of claim 1 , wherein the at least one processor is further configured to:129025-2565WO01Qualcomm Ref. No. 2501028WO 50 / 53mark a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non-usable based on the detection of the uncorrectable error in the memory, wherein memory lines or memory regions of the memory that are not marked as restricted or non-usable are accessible and usable by the PD or other PDs.

5. The apparatus of claim 4, wherein the at least one processor is further configured to:output the set of memory lines or the set of memory regions to an error logger.

6. The apparatus of claim 4, wherein the at least one processor is further configured to:refrain from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, orprevent or disable the PD or the other PDs to access or use the set of memory lines or the set of memory regions.

7. The apparatus of claim 1, wherein to issue the restart forthe PD, the at least one processor is further configured to:interrupt a high-level operation system (HLOS) or a subsystem.

8. The apparatus of claim 1, wherein the system reset corresponds to a system-on-chip(SoC) system reset and the subsy stem resetcorrespon ds toadigital signal processors (DSP) subsystem reset or a micro-controller subsystem reset.

9. The apparatus of claim 1, wherein the memory corresponds to at least one static random-access memory (SRAM) or at least one vector tightly coupled memory (VTCM).

10. The apparatus of claim 1, wherein the memory is protected with an error correction code (ECC) in the process domain, and wherein the uncorrectable error corresponds to an ECC error.129025-2565WO01Qualcomm Ref. No. 2501028WO 51 / 5311. The apparatus of claim 1, wherein the PD is associated with an advanced driver assistance systems (ADAS) system, an infotainment system, or a system operating below a power threshold or a capability threshold.

12. A method at a device, comprising:detecting an uncorrectable error in memory of a process domain (PD); incrementing an uncorrectable error counter based on detection of the uncorrectable error in the memory; andperforming, based on the uncorrectable error counter, atleastone of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuing a restart for the PD if the uncorrectable error counter does not exceed the counter threshold.

13. The method of claim 12, wherein detecting the uncorrectable error in the memory of the PD comprises:identifying an identification (ID) of the PD based on a software threshold ID programmed in a hardware thread or core on which the uncorrectable error is detected; andobtaining, based on the ID of the PD, information related to a set of memory lines or a set of memory regions that is associated with the uncorrectable error in the PD.

14. The method of claim 12, further comprising:halting or aborting a set of hardware threads that is executing a set of processes related to the PD; andrestarting or resuming the set of hardware threads after the system reset, the subsystem reset, the restart of the set of PDs, or the restart of the PD.

15. The method of claim 12, further comprising:marking a set of memory lines or a set of memory regions of the memory associated with the uncorrectable error as restricted or non-usable based on the detection of the uncorrectable error in the memory, wherein memory lines or memory regions of129025-2565WO01Qualcomm Ref. No. 2501028WO 52 / 53the memory that are not marked as restricted or non-usable are accessible and usable by the PD or other PDs.

16. The method of claim 15, further comprising:outputtingthe set of memory lines or the set of memory regions to an error logger.

17. The method of claim 15, further comprising:refraining from using the set of memory lines or the set of memory regions that are marked as restricted or non-usable, orpreventing or disablingthe PD or the other PDs to access or use the set of memory lines or the set of memory regions.

18. The method of claim 12, wherein issuing the restart for the PD further comprises:interrupting a high-level operation system (HLOS) or a subsystem.

19. The method of claim 12, wherein the memory corresponds to at least one static random-access memory (SRAM) or at least one vector tightly coupled memory (VTCM).

20. A computer-readable medium storing computer executable code at a device, the code when executed by at least one processor causes the at least one processor to:detect an uncorrectable error in memory of a process domain (PD); increment an uncorrectable error counter based on detection of the uncorrectable error in the memory; andperform, based on the uncorrectable error counter, at least one of: (1) triggering a system reset or a subsystem reset if the uncorrectable error counter exceeds a counter threshold, (2) issuing a restart for a set of PDs if the uncorrectable error counter exceeds the counter threshold, or (3) issuinga restartforthe PD if the uncorrectable error counter does not exceed the counter threshold.129025-2565WO01