Data Processing Method, Device, Storage Medium, and Electronic Device
By injecting the modified exception context into the kernel, the problem of inaccurate handling of UCE memory exceptions is solved, and higher processing accuracy and system stability are achieved.
Patent Information
- Application Number
- CN202510005757.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-02
AI Technical Summary
In hyperscale cloud computing data centers, the kernel has low accuracy in handling exceptions triggered by UCE memory by consumption, and the lack of exception context information leads to inaccurate processing.
The initial exception context of the target task is obtained through trusted firmware, modified and injected into the kernel, and the kernel performs exception processing based on the target exception context.
Improve the accuracy of exception handling, ensure that the kernel can effectively handle UCE memory exceptions and avoid system downtime.
Smart Images

Figure CN119396619B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a data processing method, an apparatus, a storage medium, and an electronic device. Background Art
[0002] In a hyperscale cloud computing data center, the server downtime rate has always been a key indicator for measuring RAS (Reliability, Availability, and Serviceability), and it is also the primary issue for meeting the SLA (Service Level Agreement) of cloud computing end users.
[0003] Unexpected system downtime not only affects the normal operation of tasks but also damages the enterprise's reputation. In computing clusters and data centers, the hardware and software density deployed on a single physical machine is getting higher and higher. Hundreds of VMs (Virtual Machines) and 2,500 container instances can be deployed on each physical machine. Although hardware failures rarely occur, any server downtime may result in huge cost losses. Memory failures are the main cause of server failures.
[0004] In the existing technology Firmware First (a firmware - first error - handling mode), a memory error interrupt of hardware is first responded to by firmware. The firmware collects the error information of the hardware, records the error in accordance with the APEI (ACPI Platform Error Interfaces, an industry - standard interface defined by the ACPI (Advance Configuration and Power Interface) specification), and notifies the kernel of the hardware error information. Then the kernel parses out the error information according to the APEI specification and processes different error types separately. However, when the kernel processes an exception, it lacks the cause of the exception and the state of the PE (Processing Element, an entity capable of executing a program) during the exception, and lacks a basis for judging whether the exception handling can be recovered. Therefore, there is a problem of relatively low accuracy in handling exceptions.
[0005] Regarding the problem of relatively low accuracy in exception handling when the kernel processes an exception triggered by the consumption of UCE memory in the above - mentioned related technologies, no effective solution has been proposed yet. Summary of the Invention
[0006] Embodiments of the present application provide a data processing method, apparatus, storage medium, and electronic device to at least solve the technical problem of low accuracy in exception handling when the kernel processes exceptions triggered by the consumption of UCE memory in related technologies.
[0007] According to one aspect of the embodiments of the present application, a data processing method is provided, including: triggering a synchronous exception when detecting that a target task is executed, where the target task is a task corresponding to the consumption of uncorrectable error memory; in response to the synchronous exception, obtaining an initial exception context of the target task through a trusted firmware; modifying the initial exception context through the trusted firmware to obtain a target exception context, and injecting the target exception context into the kernel; and performing exception handling by the kernel based on the target exception context.
[0008] Further, modifying the initial exception context through the trusted firmware to obtain a target exception context includes: modifying the exception type in the initial exception context to obtain a modified initial exception context; and determining the modified initial exception context as the target exception context.
[0009] Further, before performing exception handling by the kernel based on the target exception context, the method further includes: determining a target offset corresponding to an interrupt vector table in the kernel according to the exception level and stack pointer corresponding to the synchronous exception, and writing the target offset into an exception link register corresponding to the trusted firmware, where the interrupt vector table stores exception handling functions for performing exception handling; modifying the exception level and stack pointer in a save run state register corresponding to the trusted firmware to obtain a modified save run state register; and switching from the level where the trusted firmware is located to the level where the kernel is located and jumping to the interrupt vector table of the kernel based on the target offset in the exception link register and the modified save run state register through an exception return instruction, so as to perform exception handling by the kernel based on the target exception context.
[0010] Further, the method further includes: restoring data information of the target task in a stack corresponding to the trusted firmware to a general register, where after injecting the target exception context into the kernel, the kernel writes the data information of the target task in the general register into a stack corresponding to the kernel to perform exception handling by the kernel based on the target exception context.
[0011] Further, the kernel performing exception handling based on the target exception context includes: determining whether the synchronous exception is a repairable exception based on the target exception context; if the synchronous exception is a repairable exception, determining the exception level of the trigger source corresponding to the synchronous exception based on the target exception context; and performing exception handling according to the exception level.
[0012] Further, after triggering a synchronous exception when detecting that a target task is executed, the method further includes: performing an interruption process on the target task; performing exception handling according to the exception level includes: if the exception level is the level corresponding to the target task being a host user-mode task, sending a bus error signal to the host user-mode task through an exception handling function in the kernel; when the interruption of the host user-mode task is restored and the host user-mode task does not set a signal handling function for handling the bus error signal, ending the host user-mode task based on the bus error signal.
[0013] Further, performing exception handling according to the exception level includes: if the exception level is the level corresponding to the target task being a host kernel-mode task, determining the target instruction address of the target task from the target exception context; judging whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table according to the target instruction address; if there is no such exception repair instruction address, performing a crash handling.
[0014] Further, after judging whether there is an exception repair instruction address corresponding to the target instruction address in the preset exception table, the method further includes: if there is such an exception repair instruction address, determining the exception repair instruction address as the target instruction address of the target task in the target exception context to obtain a repaired target exception context; and restoring the host kernel-mode task based on the repaired target exception context.
[0015] Further, performing exception handling according to the exception level includes: if the exception level is the level corresponding to the target task being a virtual machine task, sending a bus error signal to a virtual machine simulator process through an exception handling function in the kernel, and the virtual machine task is implemented based on the virtual machine simulator process; when the interruption of the virtual machine task is restored, judging whether there is a target function for handling the bus error signal in the virtual machine simulator process; if there is no such target function, ending the virtual machine simulator process based on the bus error signal.
[0016] Further, after determining that there is a target function for processing the bus error signal in the virtual machine simulator process, the method further includes: if there is the target function, constructing a target exception type through the target function; injecting a virtual synchronous exception into the virtual machine kernel according to the target exception type; and performing exception handling through the virtual machine kernel.
[0017] According to another aspect of the embodiments of the present application, there is also provided a data processing device, including: a triggering unit, configured to trigger a synchronous exception when detecting that a target task is executed, where the target task is a task corresponding to the consumption of uncorrectable error memory; a first obtaining unit, configured to respond to the synchronous exception and obtain an initial exception context of the target task through a trusted firmware; a first modifying unit, configured to modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the kernel; and a first executing unit, configured to perform exception handling through the kernel based on the target exception context.
[0018] Further, the first modifying unit includes: a modifying subunit, configured to modify an exception type in the initial exception context to obtain a modified initial exception context; an obtaining subunit, configured to obtain an exception level and a stack pointer corresponding to the synchronous exception; and a processing subunit, configured to modify the modified initial exception context according to the exception level and the stack pointer corresponding to the synchronous exception to obtain the target exception context.
[0019] Further, the device further includes: a first determining unit, configured to determine a target offset corresponding to an interrupt vector table in the kernel according to the exception level and the stack pointer corresponding to the synchronous exception, and write the target offset into an exception link register corresponding to the trusted firmware before performing exception handling through the kernel based on the target exception context, where the interrupt vector table stores an exception handling function for performing exception handling; a second modifying unit, configured to modify the exception level and the stack pointer in a saved running state register corresponding to the trusted firmware to obtain a modified saved running state register; and a jumping unit, configured to switch from a level where the trusted firmware is located to a level where the kernel is located and jump to the interrupt vector table of the kernel based on the target offset in the exception link register and the modified saved running state register through an exception return instruction, so as to perform exception handling through the kernel based on the target exception context.
[0020] Further, the device further includes: a second acquisition unit, configured to restore data information of a target task in a stack corresponding to the trusted firmware to a general-purpose register, where after injecting the target exception context into the kernel, the kernel writes the data information of the target task in the general-purpose register into a stack corresponding to the kernel, so that the kernel performs exception handling based on the target exception context.
[0021] Further, the first execution unit includes: a judgment subunit, configured to judge whether the synchronous exception is a repairable exception based on the target exception context; a determination subunit, configured to, if the synchronous exception is a repairable exception, determine an exception level of a trigger source corresponding to the synchronous exception based on the target exception context; and an execution subunit, configured to perform exception handling according to the exception level.
[0022] Further, the device further includes: an interrupt unit, configured to perform an interrupt handling on the target task after triggering a synchronous exception when detecting that the target task is executed; the execution subunit includes: a first sending module, configured to, if the exception level is a level corresponding to the target task being a host user-mode task, send a bus error signal to the host user-mode task through an exception handling function in the kernel; and a first ending module, configured to end the host user-mode task based on the bus error signal when an interrupt of the host user-mode task is restored.
[0023] Further, the execution subunit includes: a determination module, configured to, if the exception level is a level corresponding to the target task being a host kernel-mode task, determine a target instruction address of the target task from the target exception context; a first judgment module, configured to judge whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table according to the target instruction address; and an execution module, configured to perform a crash handling if there is no such exception repair instruction address.
[0024] Further, the device further includes: a second determination unit, configured to, after judging whether there is an exception repair instruction address corresponding to the target instruction address in the preset exception table, if there is such an exception repair instruction address, determine the exception repair instruction address as the target instruction address of the target task in the target exception context to obtain a repaired target exception context; and a restoration unit, configured to restore the host kernel-mode task based on the repaired target exception context.
[0025] Further, the execution subunit includes: a second sending module, configured to, if the exception level is the level corresponding to the target task being a virtual machine task, send a bus error signal to the virtual machine simulator process through an exception handling function in the kernel, where the virtual machine task is implemented based on the virtual machine simulator process; a second determination module, configured to determine whether there is a target function for handling the bus error signal in the virtual machine simulator process when the interruption of the virtual machine task is restored and the host user-mode task does not set a signal handling function for handling the bus error signal; a second ending module, configured to, if there is no such target function, end the virtual machine simulator process based on the bus error signal.
[0026] Further, the apparatus further includes: a construction unit, configured to, after determining that there is a target function for handling the bus error signal in the virtual machine simulator process, if there is such a target function, construct a target exception type through the target function; an injection unit, configured to inject a virtual synchronization exception into the virtual machine kernel according to the target exception type; a second execution unit, configured to execute exception handling through the virtual machine kernel.
[0027] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the data processing method described in any one of the above.
[0028] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory storing an executable program; a processor, configured to run the program, where when the program runs, it executes the data processing method described in any one of the above.
[0029] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, where the computer program product includes a stored computer program, and when the computer program runs on a processor, it implements the data processing method described in any one of the above.
[0030] In the embodiments of the present application, the following steps are adopted: when it is detected that the target task is executed, a synchronization exception is triggered, where the target task is a task corresponding to the consumption of uncorrectable error memory; in response to the synchronization exception, the initial exception context of the target task is obtained through the trusted firmware; the initial exception context is modified through the trusted firmware to obtain the target exception context, and the target exception context is injected into the kernel; the kernel performs exception handling based on the target exception context, which solves the technical problem in the related art that when the kernel processes the exception triggered by the consumption of UCE memory, there is a lack of exception context, resulting in relatively low accuracy of exception handling. In this solution, by modifying the initial exception context to obtain the target exception context and injecting the target exception context into the kernel, the purpose that the kernel can perform exception handling in combination with the target exception context is achieved, thereby realizing the technical effect of improving the accuracy of exception handling. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0032] Figure 1 is the flowchart of the data processing method provided in Embodiment 1 of the present application Figure 1 ;
[0033] Figure 2 is the schematic diagram of the exception level provided in Embodiment 1 of the present application;
[0034] Figure 3 is the flowchart of the prior art exception handling;
[0035] Figure 4 is the flowchart of the data processing method provided in Embodiment 1 of the present application Figure 2 ;
[0036] Figure 5 is the flowchart of the data processing method provided in Embodiment 1 of the present application Figure 3 ;
[0037] Figure 6 is the schematic diagram of the data processing device provided in Embodiment 2 of the present application;
[0038] Figure 7 is the structural block diagram of the electronic device provided in Embodiment 3 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the scope of protection of this application.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0041] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations:
[0042] RAS: Reliability, Availability and Serviceability, the reliability, availability and serviceability of a computer system;
[0043] SLA: Service Level Agreement, a service level agreement;
[0044] UCE: Uncorrected Correctable Error, an uncorrectable error;
[0045] SEA: Synchronous External Abort, a synchronous exception in the ARM architecture, which refers to an exception triggered due to some error when the processor attempts to access an external memory location. This exception is related to the instruction currently executed by the processor;
[0046] ELx: Exception levels are referred to as EL <x>, with x as a number between 0 and 3, is a term used in the ARM architecture to define different privilege levels, and these exception levels define the logical separation of software permissions. X is 0, 1, 2, 3;
[0047] TF-A: Trusted Firmware-A, a set of secure and trusted software components.
[0048] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0049] Embodiment 1
[0050] According to an embodiment of the present application, a data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0051] The present application provides as Figure 1 the data processing method shown. Figure 1 The flow of the data processing method according to Embodiment 1 of the present application Figure 1 . The method includes:
[0052] Step S101, when detecting that the target task is executed, trigger a synchronous exception, where the target task is the task corresponding to the consumption of uncorrectable error memory.
[0053] Optionally, in ARMv8-A, the concept of privilege level is called exception level (EL), which is a logical division of the privilege level where the program runs. The exception levels are divided into the following four types. Note that the larger the number, the higher the privilege. To optimize the virtualization context switching overhead brought by multiple exception levels, the virtual Host Extension (VHE) feature was introduced in ARMv8.1. The Linux kernel and KVM both execute at EL2, which can reduce the number of mode switches between EL1 and EL2 and has higher performance.
[0054] Such as Figure 2 As shown in the figure, for the host user mode, that is, Host App(s), and the guest user mode, that is, Guest App(s), they run at EL0. EL0 is a non-privileged level with the lowest execution privilege. For the guest kernel mode, that is, Guest OS runs at EL1. EL1 is usually used to run the operating system kernel, which has privileges and can access more system resources and perform privileged operations. For the host kernel mode, that is, Host OS runs at EL2. EL2 provides support for virtualization and can be used to run the hypervisor (virtual machine monitor). For Secure Monitor / Firmware, it runs at EL3. EL3 provides the function of secure state switching and is usually used to run the security monitor program. The above-mentioned TF-A (Trusted Firmware) runs at EL3.
[0055] In the prior art, a solution for implementing SDEI (Software Delegated Exception Interface, a software proxy exception handling interface for the non-secure world to register processors with the firmware to receive system event notifications) notifications is implemented. For handling SEA, after the UCE error memory (i.e., the above-mentioned uncorrectable error memory) is accessed, a synchronous exception is triggered. The kernel first marks the memory page as "poisoned", terminates the process mapping the page, and avoids future use of the page, thereby isolating the UCE error. The process is as follows Figure 3 shown, and specifically includes the following steps:
[0056] Step 1, Task: A user-mode EL0 task initiates a read request to access the UCE error memory, that is, the user-mode task consumes poison;
[0057] Step 2, Hardware: Generate a synchronous exception (SEA);
[0058] Step 3, The task is preempted and trapped into EL3 for the TF-A (Trusted Firmware) to handle the exception;
[0059] Step 4, TF-A: Save the context of the interrupted task, denoted as Secure Context;
[0060] Step 5, TF-A: Respond to the synchronous exception, collect the physical address of the error, etc., and record the error information in the APEI format;
[0061] Step 6, TF-A: Notify the kernel of the error through the SDEI event;
[0062] Step 7, Kernel: SDEI processing function, parses the physical address of the UCE error from the APEI record, marks the corresponding memory page as poisoned and isolates it, and releases the address mapping of the user-mode task;
[0063] Step 8, kernel: complete SDEI event and return to EL3;
[0064] Step 9, TF-A: After restoring the interrupted task context using Secure Context, return to the preempted task;
[0065] Step 10: After the exception returns, the PC (Program Counter, pointing to the address of the instruction that caused the exception) is the current instruction, and the user-mode task re-accesses the unmapped page;
[0066] Step 11, triggering a page fault or page missing, and the kernel handles the page fault or page missing;
[0067] Step 12, kernel: The kernel sends a SIGBUS signal to the user-mode process to terminate the user-mode program.
[0068] The existing technology mainly has the following defects: The exception context lacks valid information: In the Firmware First mode, the task falls into EL3 after triggering SEA, so the preempted exception context scene is saved in the Firmware. The current SDEI implementation only provides a general register of the exception context scene. The kernel lacks the cause of the exception, the state of the PE at the time of the exception, etc., and lacks the basis for judging whether the exception handling can be restored. In addition, the above process can only use the technique of re-triggering page fault after the above exception to isolate and kill the user-mode task, and cannot handle the exception triggered by the kernel-mode task.
[0069] Lack of handling measures in the exception context: For the exception triggered when consuming poison, the current kernel lacks measures to handle kernel-mode consumption UCE errors. If the kernel-mode consumption UCE is simply treated as a recoverable error, that is, the mapping of the poison page is unmapped, then after the exception returns, the PC is the current instruction and the exception will be re-triggered. This cycle of triggering exceptions will eventually cause the system to become unavailable and crash.
[0070] In order to solve the above problems, the present application proposes a data processing method. When there is a process consumption poison, that is, when the uncorrectable error memory is accessed by the read request of the above target task, the hardware will trigger SEA (that is, the above-mentioned synchronization exception).
[0071] Step S102, in response to a synchronization exception, obtain the initial exception context of the target task through the trusted firmware.
[0072] Optionally, when SCR_EL3.EA = 1 (SCR_EL3 is a control register used to configure the security state and the behavior of exception handling, and SCR_EL3.EA = 1 means that external interrupts and error interrupts will be routed to EL3), after the software triggers SEA, the task is preempted and enters EL3. The TF-A saves the exception context. After error handling, the TF-A restores the exception context of the task, and the task continues to run. This is the so-called Firmware First mode. When SCR_EL3.EA = 0, after the software triggers SEA, the task is preempted and enters EL2. The kernel saves the exception context. After error handling, the kernel restores the exception context of the task, and the task continues to run. This is the Kernel First mode. Usually, the error register information of the hardware is set to the Secure attribute, and the kernel runs in non-secure. The kernel cannot obtain the error information. Therefore, the Firmware First mode is adopted in this application.
[0073] Therefore, at the same time as SEA is triggered, the task is preempted and enters EL3, where it is processed by the TF-A (i.e., the above-mentioned trusted firmware). That is, the trusted firmware responds to SEA and obtains the initial exception context of the interrupted task (i.e., the above-mentioned target task). It should be noted that the initial exception context will be saved in the system registers of EL3. The initial exception context includes information such as the exception type, the virtual address of the instruction when the exception occurred, and the PC when the exception occurred. The relevant EL3 system registers include:
[0074] ESR_EL3 (Exception Syndrome Register_EL3): exception type (EC, Exception Class, used to indicate the reason for the exception trigger), synchronous error type (SET, Synchronous Error Type, describing the state of the PE when a synchronous exception occurs), exception status code (DFSC, Data Fault Status Code, indicating the stage at which the exception occurred), etc.;
[0075] FAR_EL3 (Fault Address Register_EL3): the virtual address of the instruction when the exception occurred;
[0076] ELR_EL3 (Exception Link Register_EL3): the PC (Program Counter, pointing to the address of the instruction that caused the exception) when the exception occurred;
[0077] SPSR_EL3 (Saved Program Status Register_EL3): The state of the PE, PSTATE, when an exception is trapped (when an exception occurs, the processor state (PSTATE) is saved in SPSR_EL3).
[0078] It should be noted that when TF-A responds to the SEA, error record information will also be collected, such as the physical address of the above UCE error, etc., so as to perform subsequent work such as cause analysis of the above exceptions based on the error record information.
[0079] Step S103: Modify the initial exception context through the trusted firmware to obtain the target exception context, and inject the target exception context into the kernel.
[0080] Optionally, since the exception interrupt is routed to EL3 and the corresponding initial exception context conforms to EL3, in order to be able to inject the exception context into EL2 for processing, it is necessary to modify the initial exception context through TF-A to obtain the target exception context. It should be noted that the target exception context can be obtained by modifying the exception type of the initial exception context. Finally, inject the target exception context into EL2. By modifying the initial exception context and injecting the target exception context into EL2, the kernel in EL2 can process the exception context normally.
[0081] Step S104: Perform exception handling by the kernel based on the target exception context.
[0082] Optionally, the Kernel in EL2 (i.e., the above-mentioned kernel) performs exception handling according to the injected target exception context. It should be noted that after the kernel performs exception handling based on the target exception context, directly restore the interrupted task context in EL2, that is, return to the preempted task through the ERET instruction, and there is no need to jump to EL3 to restore the interrupted task context again.
[0083] For example, determine the exception level according to the injected target exception context, and then perform corresponding processing for different exception levels. For example, if the exception level is EL0, then unmap the poisoned page in the user state, and then restore the interrupted task in EL2 according to the target exception context. The user-state program accesses the unmapped page again, triggering a page fault. The kernel processes the fault, and the kernel sends a SIGBUS signal to the user-state process to kill the user-state program.
[0084] In summary, the following steps are adopted: when it is detected that a target task is executed, a synchronization exception is triggered, where the target task is a task corresponding to the consumption of uncorrectable error memory; in response to the synchronization exception, the initial exception context of the target task is obtained through a trusted firmware; the initial exception context is modified through the trusted firmware to obtain a target exception context, and the target exception context is injected into the kernel; the kernel executes exception handling based on the target exception context, solving the technical problem in the related art that when the kernel processes an exception triggered by the consumption of UCE memory, there is a lack of an exception context, resulting in relatively low accuracy of exception handling. In this solution, by modifying the initial exception context to obtain a target exception context and injecting the target exception context into the kernel, the purpose that the kernel can execute exception handling in combination with the target exception context is achieved, thereby realizing the technical effect of improving the accuracy of exception handling.
[0085] In order to enable the Kernel in EL2 to process the exception context, in the data processing method provided in the first embodiment of this application, modifying the initial exception context through the trusted firmware to obtain the target exception context includes: modifying the exception type in the initial exception context to obtain the modified initial exception context; and determining the modified initial exception context as the target exception context.
[0086] Optionally, for ESR_EL3, it is one of the most important registers for exception cause analysis in the ARM64 architecture. For example, after EL0 consumes poison, it will trigger SEA, and ESR_EL3 is 0x92000410. The parsing of the important fields of ESR_EL3 is as follows:
[0087] Exception Class [31:26]: 0x24 0b100100 Data Abort from a lowerException level
[0088] Synchronous Error Type [12:11]: 0x0 0b00 Recoverable state
[0089] Data Fault Status Code [5:0]: 0x10 0b010000 Synchronous Externalabort, not on translation table walk or hardware update of translation table。
[0090] Since TF-A is at EL3, software in the Guest user mode, Guest kernel, Host kernel, and Host user mode all come from lower-level exceptions from the perspective of TF-A after consuming the poison to trigger SEA. Therefore, after triggering SEA from EL0, EL1, and EL2, the Exception Class of ESR_EL3 will indicate that the exception comes from a lower-level exception. For example:
[0091] EC == 0b100000: Instruction Abort from a lower Exception level
[0092] EC == 0b100100: Data Abort exception from a lower Exception level.
[0093] Exceptions from EL0 and EL1 are still from lower-level exceptions for EL2 and can remain unchanged. However, exceptions from EL2 are from the current exception for EL2, so the EC needs to be modified, that is, modify the exception type (EC) in the initial exception context as described above to obtain the modified initial exception context.
[0094] For example, the above two ECs are modified to:
[0095] EC == 0b100001: Instruction Abort taken without a change in Exception level
[0096] EC == 0b100101: Data Abort exception taken without a change in Exception level
[0097] That is, ESR_EL2 = ESR_EL3 | (1 << ESR_EC_SHIFT).
[0098] By modifying the exception type in the initial exception context, the target exception context that needs to be injected into the EL2 kernel can be obtained.
[0099] It should be noted that the modification of the initial exception context can be understood as follows: in TF-A of EL3, after the constructed poison is consumed, the SEA exception is triggered and falls into the target exception context of EL2. Specifically, the EC in ESR_EL3 (Exception Syndrome Register_EL3) is modified to obtain the modified information, which is written into ESR_EL2 (Exception Syndrome Register_EL2). FAR_EL3 (Fault Address Register_EL3), ELR_EL3 (Exception Link Register_EL3), and SPSR_EL3 (Program Status Controller_EL3) are copied and written into FAR_EL2 (Fault Address Register_EL2), ELR_EL2 (Exception Link Register_EL2), and SPSR_EL2 (Program Status Controller_EL2).
[0100] Through the above modification of the exception context, from the perspective of the kernel, the exception handling of TF-A is transparent. After the software triggers SEA, the task exception context obtained by the kernel is exactly the same as the task exception context after SEA is triggered in Kernel First mode. Therefore, the kernel can obtain complete exception context information, improving the accuracy of subsequent exception handling.
[0101] After constructing the target exception context, in order to be able to jump from EL3 to the exception handling function of EL2, before the kernel executes exception handling based on the target exception context, this method further includes: obtaining the exception level and stack pointer corresponding to the synchronous exception; determining the target offset corresponding to the interrupt vector table in the kernel based on the exception level and stack pointer corresponding to the synchronous exception, and writing the target offset into the exception link register corresponding to the trusted firmware, where the interrupt vector table stores exception handling functions for executing exception handling; modifying the exception level and stack pointer in the saved run state register corresponding to the trusted firmware to obtain the modified saved run state register; and switching from the level where the trusted firmware is located to the level where the kernel is located and jumping to the interrupt vector table of the kernel through the exception return instruction based on the target offset in the exception link register and the modified saved run state register, so as to execute exception handling by the kernel based on the target exception context.
[0102] Optionally, it is necessary to modify ELR_EL3 so that the PC jumps to the interrupt vector table of the kernel. Specifically, it includes: determining the offset of the exception vector table according to the exception level of the SEA source and the used stack pointer, that is, using SPSR_EL3.M[3:0], and writing it into ELR_EL3 (i.e., the above-mentioned exception link register corresponding to the trusted firmware). It should be noted that the interrupt vector table stores exception handling functions for executing exception handling.
[0103] After a synchronous exception is triggered and the system enters EL3, the corresponding exception level information for the triggered synchronous exception is stored in SPSR_EL3 (i.e., the Saved Program Status Register). Usually, it is reflected in SPSR_EL3.M[3:2] and SPSR_EL3.M[0]. In this solution, in order to be able to switch from EL3 to EL2 when executing the exception return instruction (ERET instruction), it is also necessary to modify the exception level (i.e., SPSR_EL3.M[3:2]) and the stack pointer (i.e., SPSR_EL3.M[0]) in SPSR_EL3 (i.e., the Saved Program Status Register) to obtain the modified Saved Program Status Register. That is, it is necessary to modify SPSR_EL3.M[3:2] to EL2 and modify SPSR_EL3.M[0] to SP_ELX (the value of SP_ELX is 1).
[0104] Finally, based on the target offset in the exception link register and the modified Saved Program Status Register, the exception return instruction (ERET instruction) switches from EL3 to EL2, and the PC jumps to the kernel's interrupt vector table so that the kernel can subsequently perform exception handling based on the target exception context.
[0105] It should be noted that the target offset corresponding to the interrupt vector table to be written in ELR_EL3 is determined according to SPSR_EL3.M[3:0], as shown in Table 1:[[]]END]]
[0106] Table 1
[0107] Data Processing Method, Device, Storage Medium, and Electronic Device
[0108] The above-mentioned VBAR_EL2 (Vector Base Address Resister_EL3) is a system register used to store the base address of the interrupt vector table of the EL2 kernel, and this interrupt vector table includes exception handling functions (or exception handlers) for handling exceptions.
[0109] By modifying ELR_EL3 and SPSR_EL3, it is ensured that EL3 switches to EL2, the PC jumps to the kernel's interrupt vector table, and the SEA exception can be subsequently processed in EL2_Kernel.
[0110] In order to make the task context of the exception context equivalent to directly entering EL2_Kernel after the task is interrupted after the exception is injected, in the data processing method provided in Embodiment 1 of this application, the method further includes: restoring the data information of the target task in the stack corresponding to the trusted firmware to the general registers so that after injecting the target exception context into the kernel, the kernel writes the data information of the target task in the general registers into the stack corresponding to the kernel.
[0111] Optionally, when a synchronous exception triggers a trap to EL3, the data information of the target task in the general-purpose registers (i.e., 32 64-bit registers X0-X31 in the ARM64 platform) is written into the stack of EL3 (i.e., the stack corresponding to the above-mentioned trusted firmware). In order to make the task context of the exception context equivalent to directly trapping to EL2_Kernel after the task is interrupted after the exception is injected, first, the data information of the target task in the stack of EL3 needs to be restored to the general-purpose registers. After injecting the target exception context into the kernel and jumping to the kernel for processing through ERET, the kernel will write the data information of the target task in the general-purpose registers into the stack of EL2 (i.e., the stack corresponding to the above-mentioned kernel), so that the kernel can perform exception handling based on the target exception context. Through the above processing process, the kernel can obtain the complete exception context when the target task triggers an exception, ensuring the normal processing process of subsequent exceptions.
[0112] How to perform exception handling is crucial. Therefore, in the data processing method provided in the first embodiment of the present application, the kernel performing exception handling based on the target exception context includes: judging whether the synchronous exception is a repairable exception based on the target exception context; if the synchronous exception is a repairable exception, determining the exception level of the trigger source corresponding to the synchronous exception based on the target exception context; and performing exception handling according to the exception level.
[0113] Optionally, TF-A injects the complete exception context of the interrupted task into the kernel. The exception switches from EL3 to EL2, and the PC jumps to the corresponding exception handling function in the kernel's exception vector table, such as el0t_64_sync_handler, which is responsible for handling synchronous exceptions from EL0, and el1h_64_sync_handler is responsible for handling synchronous exceptions from EL2. Note that el10_64_sync_handler has no practical meaning and will immediately end the operation.
[0114] Moreover, since the kernel can directly obtain the exception context of the task by accessing system registers, just like in the kernel first mode. For example:
[0115] ESR_EL2: Exception type (EC), PE's error status (SET), exception status code (DFSC), etc.;
[0116] FAR_EL2: The virtual address of the instruction when the exception occurs;
[0117] ELR_EL2: The PC when the exception occurs;
[0118] SPSR_EL2: The state PSTATE of PE when the exception traps.
[0119] First, the kernel determines whether the SEA is recoverable based on ESR_EL2. In an optional embodiment, if the following three conditions are all met, it is recoverable:
[0120] The Exception Class is Data abort or Instruction abort
[0121] SET is UER (i.e., Synchronous Error Type [12:11] is 0x0 0b00 Recoverablestate)
[0122] DFSC is SEA (i.e., Data Fault Status Code [5:0] is 0x10 0b010000 SynchronousExternal abort, not on translation table walk or hardware update oftranslation table).
[0123] If the SEA does not meet the above three conditions, it is determined to be an unrecoverable error and immediate downtime processing is required.
[0124] If the synchronous exception is a repairable exception, the exception level of the error trigger source of the SEA is determined according to SPSR_EL2.M[3:2]:
[0125] If SPSR_EL2.M[3:2] == 0b10, the exception level of the error trigger source of the SEA is EL2
[0126] If SPSR_EL2.M[3:2] == 0b01, the exception level of the error trigger source of the SEA is EL1
[0127] If SPSR_EL2.M[3:2] == 0b00, the exception level of the error trigger source of the SEA is EL0
[0128] Process them separately.
[0129] The kernel can more accurately determine whether the exception handling can be recovered based on the injected target exception information, and can more accurately handle SEAs of different exception levels.
[0130] To improve the accuracy of SEA processing for different exception levels, in the data processing method provided in the first embodiment of this application, after detecting that a target task is executed and triggering a synchronous exception, the method further includes: performing an interruption process on the target task; performing exception handling according to the exception level, including: if the exception level is the level corresponding to the target task being a host user-mode task, sending a bus error signal to the host user-mode task through an exception handling function in the kernel; when the interruption of the host user-mode task is restored, ending the host user-mode task based on the bus error signal.
[0131] Optionally, when triggering a synchronous exception, the currently executing target task will be interrupted, and then according to the exception level, performing exception handling includes the following steps: If the exception level is the level corresponding to the target task being a host user-mode task (Host user-mode task), then through the exception handling function, it can be clarified as a synchronous exception error. When it is a memory_failure, specify the parameter as MF_ACTION_REQUIRED, and directly send a SIGBUS signal (i.e., the above-mentioned bus error signal) to the host user-mode task, with the error code being BUS_MCEERR_AR 4. When resuming the abnormal context of the task and continuing to run (i.e., when the interruption of the host user-mode task is restored), if the user-mode task has registered a signal handling function for SIGBUS, then execute the corresponding signal handling function. It should be noted that how the signal handling function is processed can be set according to user requirements.
[0132] If there is no above-mentioned signal handling function for processing SIGBUS, then end the host user-mode task according to the SIGBUS signal. It should be noted that the level corresponding to the Host user mode is EL0.
[0133] By directly sending a SIGBUS signal to the host user-mode task, it is avoided to trigger a page fault again to end the exception when resuming the abnormal context of the task and continuing to run. Therefore, it can effectively improve the processing efficiency of the host user-mode task.
[0134] To improve the accuracy of SEA processing for different exception levels, in the data processing method provided in the first embodiment of this application, performing exception handling according to the exception level further includes: if the exception level is the level corresponding to the target task being a host kernel-mode task, determining the target instruction address of the target task from the target abnormal context; based on the target instruction address, determining whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table; if there is no exception repair instruction address, performing a crash handling.
[0135] If there is an abnormal repair instruction address, the abnormal repair instruction address is determined as the target instruction address of the target task in the target abnormal context, and the repaired target abnormal context is obtained; based on the repaired target abnormal context, the host kernel task is restored.
[0136] Optionally, when the exception level is the level corresponding to the target task being a host kernel task (Host kernel task), isolation cannot be performed by directly sending a SIGBUS signal. However, when the kernel implements memory copy, repair jumps can be added to memory access instructions such as LDTR / LDTRB / LDTRH, STTR / STTRB / STTRH, etc., and recorded in the extable (i.e., the above-mentioned preset exception table).
[0137] In an optional embodiment, the exception table entry format of the extable is a quadruple (insn, fixup, type, data): insn: the PC of the current instruction (i.e., the above-mentioned target instruction address); fixup: the PC of the abnormal repair instruction (i.e., the above-mentioned abnormal repair instruction address); type: the type of abnormal repair; data: the data for abnormal repair (referring to the data returned when executing the abnormal repair instruction address). It should be noted that the level corresponding to the host kernel task (Host kernel task) is EL2.
[0138] Therefore, when the kernel processes the SEA triggered by the host kernel task (Host kernel task), it uses the interrupted task PC saved in the ELR_EL2 to check whether there is an exception table entry in the extable, that is, it uses the target instruction address to determine whether there is an abnormal repair instruction address corresponding to the target instruction address in the preset exception table. If there is no corresponding fixup, it directly crashes.
[0139] If there is an abnormal repair instruction address corresponding to the target instruction address in the exception table, the PC in the target abnormal context is modified to the fixup specified in the entry, that is, the abnormal repair instruction address is determined as the target instruction address of the target task in the target abnormal context to obtain the repaired target abnormal context. Finally, based on the repaired target abnormal context, the host kernel task is restored, and the task to be re-executed will jump to the modified new PC, thereby achieving the skipping of the exception.
[0140] Compared with the SDEI scheme in the prior art, for kernel tasks, the abnormal context of the task cannot be effectively modified. After the abnormal return, it returns to the current instruction, triggering the exception in a loop, resulting in the system being unavailable and finally crashing. Through the above steps, when processing the SEA of kernel tasks, the kernel can effectively modify the abnormal context of the task, achieve the error recovery of the task, and improve the accuracy of exception handling.
[0141] To improve the accuracy of SEA processing for different exception levels, in the data processing method provided in the first embodiment of this application, according to the exception level, performing exception handling includes: if the exception level is the level corresponding to the target task being a virtual machine task, send a bus error signal to the virtual machine simulator process through the exception handling function in the kernel, and the virtual machine task is implemented based on the virtual machine simulator process; when the interruption of the virtual machine task is restored, determine whether there is a target function in the virtual machine simulator process for handling the bus error signal; if there is no target function, end the virtual machine simulator process based on the bus error signal.
[0142] Optionally, for the Host, the virtual machine VM is a QEMU user-space process (i.e., the above-mentioned virtual machine simulator process). Therefore, if the exception level is the level corresponding to the target task being a virtual machine task (Guest task), first use the current processing method for the Host user space to send a SIGBUS signal (i.e., the above-mentioned bus error signal) with an error code of BUS_MCEERR_AR_4 to the QEMU process through the exception handling function in the kernel. When the exception of the virtual machine task is restored, QEMU checks whether there is a target function for handling the SIGBUS signal (for example, the target function is kvm_arch_on_sigbus_vcpu). If there is no corresponding target function, end the virtual machine simulator process based on the bus error signal.
[0143] By sending the SIGBUS signal, it is possible to avoid triggering a page fault again and improve the efficiency of exception handling.
[0144] If there is a target function, in the data processing method provided in the first embodiment of this application, after determining that there is a target function in the virtual machine simulator process for handling the bus error signal, the method further includes: constructing a target exception type through the target function; injecting a virtual synchronous exception into the virtual machine kernel according to the target exception type; and performing exception handling through the virtual machine kernel.
[0145] Optionally, if there is kvm_arch_on_sigbus_vcpu, when kvm_arch_on_sigbus_vcpu determines that the error code of SIGBUS is BUS_MCEERR_AR_4, it constructs an ESR value (i.e., the above-mentioned target exception type), and then injects a virtual synchronous exception (vSEA) into the virtual machine kernel according to the target exception type, and performs exception handling through the virtual machine kernel.
[0146] For the Guest Kernel (i.e., the virtual machine kernel mentioned above) to execute the exception handling process, it is the same as receiving a hardware-triggered SEA. Errors of user-mode tasks and kernel-mode tasks are processed separately. If it is a user-mode SEA error, a SIGBUS signal is sent to the user-mode process, and the user-mode process within the Guest is killed. If it is a kernel-mode SEA error, the PC of the interrupted task saved in ELR_EL1 is used to check if there is an exception entry in the extable. If not, the system crashes immediately; if there is, the PC of the exception context is modified to the fixp specified in the exception entry. After the kernel restores the exception context of the task, it will jump to the modified new PC.
[0147] In the data processing method provided by Embodiment 1 of this application, the kernel can directly obtain the exception context of a task by accessing system registers, just like in the kernel first mode, such as the trigger reason of the SEA exception, the instruction PC that triggers the SEA, the running task, etc. When the kernel processes the SEA of a user-mode task, the kernel can directly obtain that the exception reason of the task is a synchronous exception and directly send a SIGBUS signal to avoid triggering a page fault again. When the kernel processes the SEA of a kernel-mode task, the kernel can effectively modify the exception context of the task to achieve error recovery of the task. In the existing SDEI solution, for kernel-mode tasks, the exception context of the task cannot be effectively modified, and in this solution, the advantages of the Firmware First mode can also be utilized simultaneously to obtain hardware error record information so that the operation and maintenance personnel can accurately locate the faulty unit of the hardware.
[0148] In an optional embodiment, the exception handling can be implemented through a flowchart as Figure 4 shown. Specifically, it includes: Step 1, software task: consume poison (i.e., consume the uncorrectable error); Step 2, hardware: generate a synchronous exception SEA; Step 3, the task is preempted and enters EL3 for the TF-A to handle the exception; Step 4, TF-A: save the context of the interrupted task, denoted as the EL3_STATE context; Step 5, TF-A: respond to the SEA, collect the physical address of the error, etc., and record the error information in the APEI format so that the operation and maintenance personnel can accurately locate the faulty unit of the hardware; Step 6, TF-A: modify the EL3_STATE context; Step 7, TF-A: inject the modified EL3_STATE context into the kernel; Step 8, kernel: the exception handling function repairs the recoverable exception context to complete the error handling process; Step 9, kernel: after restoring the context of the interrupted task, return to the preempted task through the ERET instruction.
[0149] It should be noted that the EL3_STATE context includes system registers related to exceptions (i.e., the above-mentioned target exception context) and general-purpose registers used by interrupted tasks. For injecting the modified EL3_STATE context into the kernel, it includes: writing the data of the system registers in the modified EL3_STATE to ESR_EL2, FAR_EL2, ELR_EL2, and SPSR_EL2, and then, according to the Procedure Call Standard of the ARM64 architecture, restoring the general-purpose registers of the interrupted task from the stack in EL3; using SPSR_EL3.M[3:0] to determine the offset of the exception vector table, writing it to ELR_EL3, modifying SPSR_EL3.M[3:2] to EL2, and modifying SPSR_EL3.M[0] to SP_ELX (the value of SP_ELX is 1). Finally, based on the modified ELR_EL3 and SPSR_EL3, the ERET instruction will complete the switch from EL3 to EL2 and jump the PC to the kernel's exception handling function.
[0150] In an optional embodiment, exception injection can be implemented through a flowchart as Figure 5 shown: Step 1, trigger SEA; Step 2, enter EL3; Step 2.1, according to the Procedure Call Standard of the ARM64 architecture, write the values of the general-purpose registers used by the target task to the stack in EL3; Step 2.2, modify the initial exception context to obtain the modified target exception context, and write the modified target exception context to EL2; Step 2.3, in response to SEA, collect the incorrect physical address, etc., and record the error information in the APEI format; Step 2.4, modify ELR_EL3 and SPSR_EL3 to jump to the interrupt vector table of EL2-kernel, that is, ELR_EL3 = VBAR_EL2 + offset, SPSR_EL3.M[3:2] = EL2, SPSR_EL3.M[0] = SP_ELX (the value of SP_ELX is 1); Step 2.5, restore the values of the task general-purpose registers saved in the stack in EL3 to the general-purpose registers; Step 3, complete the exception switch and PC jump through the ERET instruction; Step 3.1, according to the Procedure Call Standard of the ARM64 architecture, write the values of the general-purpose registers used by the target task to the stack in EL3, Step 3.3, execute exception handling in EL2-kernel, Step 3.4, restore the values of the task general-purpose registers saved in the stack in EL3 to the general-purpose registers; Step 4, resume the task from EL2 through the ERET instruction and continue running.
[0151] In the data processing method provided in the first embodiment of this application, the following steps are adopted: when it is detected that a target task is executed, a synchronization exception is triggered, where the target task is a task corresponding to the consumption of uncorrectable error memory; in response to the synchronization exception, the initial exception context of the target task is obtained through a trusted firmware; the initial exception context is modified through the trusted firmware to obtain a target exception context, and the target exception context is injected into the kernel; the kernel executes exception handling based on the target exception context, which solves the technical problem in the related art that when the kernel processes the exception triggered by the consumption of UCE memory, there is a lack of exception context, resulting in relatively low accuracy of exception handling. In this solution, by modifying the initial exception context to obtain a target exception context and injecting the target exception context into the kernel, the purpose that the kernel can execute exception handling in combination with the target exception context is achieved, thereby realizing the technical effect of improving the accuracy of exception handling.
[0152] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of this application.
[0154] Embodiment 2
[0155] According to an embodiment of this application, there is also provided a data processing device for implementing the above data processing method, as Figure 6 shown. The device includes: a trigger unit 601, a first acquisition unit 602, a first modification unit 603, and a first execution unit 604.
[0156] The trigger unit 601 is configured to trigger a synchronization exception when it is detected that a target task is executed, where the target task is a task corresponding to the consumption of uncorrectable error memory;
[0157] A first acquisition unit 602, configured to respond to a synchronization exception and acquire an initial exception context of a target task through a trusted firmware;
[0158] A first modification unit 603, configured to modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the kernel;
[0159] A first execution unit 604, configured to perform exception handling based on the target exception context through the kernel.
[0160] In the data processing apparatus provided in the second embodiment of the present application, a trigger unit 601 triggers a synchronization exception when detecting that a target task is executed, where the target task is a task corresponding to the consumption of uncorrectable error memory; the first acquisition unit 602 responds to the synchronization exception and acquires an initial exception context of the target task through the trusted firmware; the first modification unit 603 modifies the initial exception context through the trusted firmware to obtain a target exception context, and injects the target exception context into the kernel; the first execution unit 604 performs exception handling based on the target exception context through the kernel, thereby solving the technical problem in the related art that when the kernel processes an exception triggered by the consumption of UCE memory, there is a lack of an exception context, resulting in relatively low accuracy of exception handling.
[0161] In this solution, by modifying the initial exception context to obtain a target exception context and injecting the target exception context into the kernel, the purpose that the kernel can perform exception handling in combination with the target exception context is achieved, thereby realizing the technical effect of improving the accuracy of exception handling.
[0162] Optionally, in the data processing apparatus provided in the second embodiment of the present application, the first modification unit includes: a modification subunit, configured to modify the exception type in the initial exception context to obtain a modified initial exception context; a processing subunit, configured to determine the modified initial exception context as the target exception context.
[0163] Optionally, in the data processing device provided in the second embodiment of the present application, the device further includes: a first determination unit, configured to determine a target offset corresponding to an interrupt vector table in the kernel according to an exception level and a stack pointer corresponding to a synchronous exception before performing exception handling by the kernel based on a target exception context, and write the target offset into an exception link register corresponding to a trusted firmware, where the interrupt vector table stores exception handling functions for performing exception handling; a second modification unit, configured to modify the exception level and the stack pointer in a saved running state register corresponding to the trusted firmware to obtain a modified saved running state register; a jump unit, configured to switch from a level where the trusted firmware is located to a level where the kernel is located based on the target offset in the exception link register and the modified saved running state register through an exception return instruction, and jump to an exception vector table of the kernel, so as to perform exception handling by the kernel based on the target exception context.
[0164] Optionally, in the data processing device provided in the second embodiment of the present application, the device further includes: a second acquisition unit, configured to restore data information of a target task in a stack corresponding to the trusted firmware to a general-purpose register, where after injecting the target exception context into the kernel, write the data information of the target task in the general-purpose register into a stack corresponding to the kernel, so as to perform exception handling by the kernel based on the target exception context.
[0165] Optionally, in the data processing device provided in the second embodiment of the present application, the first execution unit includes: a judgment subunit, configured to judge whether a synchronous exception is a repairable exception based on a target exception context; a determination subunit, configured to, if the synchronous exception is a repairable exception, determine an exception level of a trigger source corresponding to the synchronous exception based on the target exception context; an execution subunit, configured to perform exception handling according to the exception level.
[0166] Optionally, in the data processing device provided in the second embodiment of the present application, the device further includes: an interrupt unit, configured to perform interrupt handling on a target task after triggering a synchronous exception when detecting that the target task is executed; the execution subunit includes: a first sending module, configured to, if the exception level is a level corresponding to a target task being a host user-mode task, send a bus error signal to the host user-mode task through an exception handling function in the kernel; a first ending module, configured to end the host user-mode task based on the bus error signal when an interrupt of the host user-mode task is restored and the host user-mode task does not set a signal handling function for handling the bus error signal.
[0167] Optionally, in the data processing device provided in the second embodiment of the present application, the execution subunit includes: a determination module, configured to, if the exception level is the level corresponding to the target task being a host kernel task, determine the target instruction address of the target task from the target exception context; a first judgment module, configured to judge whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table according to the target instruction address; an execution module, configured to perform a crash handling if there is no exception repair instruction address.
[0168] Optionally, in the data processing device provided in the second embodiment of the present application, the device further includes: a second determination unit, configured to, after judging whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table, if there is an exception repair instruction address, determine the exception repair instruction address as the target instruction address of the target task in the target exception context, so as to obtain a repaired target exception context; a recovery unit, configured to recover the host kernel task based on the repaired target exception context.
[0169] Optionally, in the data processing device provided in the second embodiment of the present application, the execution subunit includes: a second sending module, configured to, if the exception level is the level corresponding to the target task being a virtual machine task, send a bus error signal to a virtual machine simulator process through an exception handling function in the kernel, and the virtual machine task is implemented based on the virtual machine simulator process; a second judgment module, configured to judge whether there is a target function for processing the bus error signal in the virtual machine simulator process when an interruption of the virtual machine task is recovered; a second ending module, configured to end the virtual machine simulator process based on the bus error signal if there is no target function.
[0170] Optionally, in the data processing device provided in the second embodiment of the present application, the device further includes: a construction unit, configured to, after judging whether there is a target function for processing the bus error signal in the virtual machine simulator process, if there is a target function, construct a target exception type through the target function; an injection unit, configured to inject a virtual synchronization exception into the virtual machine kernel according to the target exception type; a second execution unit, configured to perform exception handling through the virtual machine kernel.
[0171] It should be noted here that the above trigger unit 601, first acquisition unit 602, first modification unit 603 and first execution unit 604 correspond to steps S101 to S104 in the first embodiment. The functions of the four units are the same as those of the corresponding steps in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the first embodiment above.
[0172] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios and implementation processes provided in the first embodiment, but are not limited to the schemes provided in the first embodiment.
[0173] Embodiment 3
[0174] An embodiment of the present application may provide an electronic device, and the electronic device may be any one of the electronic device terminals in the electronic device terminal group. Optionally, in this embodiment, the above-mentioned electronic device may also be replaced with a terminal device such as a mobile terminal.
[0175] Optionally, in this embodiment, the above-mentioned electronic device may be located in at least one of the multiple network devices of the computer network.
[0176] In this embodiment, the above-mentioned electronic device may execute the program code of the following steps in the data processing method: when detecting that the target task is executed, trigger a synchronization exception, where the target task is a task corresponding to the consumption of non-correctable error memory; in response to the synchronization exception, obtain the initial exception context of the target task through the trusted firmware; modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the kernel; and execute exception handling based on the target exception context through the kernel.
[0177] The above-mentioned electronic device may execute the program code of the following steps in the data processing method: modifying the initial exception context through the trusted firmware to obtain a target exception context includes: modifying the exception type in the initial exception context to obtain a modified initial exception context; obtaining the exception level and stack pointer corresponding to the synchronization exception; and determining the modified initial exception context as the target exception context.
[0178] The above-mentioned electronic device may execute the program code of the following steps in the data processing method: before executing exception handling based on the target exception context through the kernel, the method further includes: determining a target offset corresponding to the interrupt vector table in the kernel according to the exception level and stack pointer corresponding to the synchronization exception, and writing the target offset into the exception link register corresponding to the trusted firmware, where the interrupt vector table stores an exception handling function for executing exception handling; modifying the exception level and stack pointer in the save run state register corresponding to the trusted firmware to obtain a modified save run state register; and switching from the level where the trusted firmware is located to the level where the kernel is located based on the target offset in the exception link register and the modified save run state register through an exception return instruction, and jumping to the exception vector table of the kernel to execute exception handling based on the target exception context through the kernel.
[0179] The above electronic device can execute the program code for the following steps in the data processing method: restoring the data information of the target task in the stack corresponding to the trusted firmware to the general-purpose register, where after injecting the target exception context into the kernel, the data information of the target task in the general-purpose register is written into the stack corresponding to the kernel, so that the kernel executes exception handling based on the target exception context.
[0180] The above electronic device can execute the program code for the following steps in the data processing method: the kernel executing exception handling based on the target exception context includes: determining whether the synchronous exception is a repairable exception based on the target exception context; if the synchronous exception is a repairable exception, determining the exception level of the trigger source corresponding to the synchronous exception based on the target exception context; and executing exception handling according to the exception level.
[0181] The above electronic device can execute the program code for the following steps in the data processing method: after detecting that the target task is executed and triggering a synchronous exception, the method further includes: performing interrupt processing on the target task; executing exception handling according to the exception level includes: if the exception level is the level corresponding to the target task being a host user-mode task, sending a bus error signal to the host user-mode task through an exception handling function in the kernel; and ending the host user-mode task based on the bus error signal when the interrupt of the host user-mode task is restored and the host user-mode task does not set a signal handling function for processing the bus error signal.
[0182] The above electronic device can execute the program code for the following steps in the data processing method: executing exception handling according to the exception level includes: if the exception level is the level corresponding to the target task being a host kernel-mode task, determining the target instruction address of the target task from the target exception context; judging whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table according to the target instruction address; if there is no exception repair instruction address, performing a crash handling.
[0183] The above electronic device can execute the program code for the following steps in the data processing method: after judging whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table, the method further includes: if there is an exception repair instruction address, determining the exception repair instruction address as the target instruction address of the target task in the target exception context to obtain a repaired target exception context; and restoring the host kernel-mode task based on the repaired target exception context.
[0184] The above-mentioned electronic device can execute the program code for the following steps in the data processing method: According to the exception level, the exception handling is performed, including: If the exception level is the level corresponding to the virtual machine task as the target task, then a bus error signal is sent to the virtual machine simulator process through the exception handling function in the kernel, and the virtual machine task is implemented based on the virtual machine simulator process; When the interruption of the virtual machine task is restored, it is judged whether there is a target function for processing the bus error signal in the virtual machine simulator process; If there is no target function, the virtual machine simulator process is ended based on the bus error signal.
[0185] The above-mentioned electronic device can execute the program code for the following steps in the data processing method: After judging that there is a target function for processing the bus error signal in the virtual machine simulator process, the method further includes: If there is a target function, a target exception type is constructed through the target function; According to the target exception type, a virtual synchronous exception is injected into the virtual machine kernel; Through the virtual machine kernel, the exception handling is executed.
[0186] Optionally, Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 7 shown, the electronic device 70 may include: one or more ( Figure 7 only one is shown in the figure) processors 702, a memory 704. The electronic device 70 may further include a storage controller for controlling and managing the memory 704 through the storage controller; The electronic device 70 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc. through the peripheral interface.
[0187] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the above-mentioned data processing method is realized. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories may be connected to the electronic device 70 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0188] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: when detecting that a target task is executed, trigger a synchronous exception, where the target task is a task corresponding to the consumption of uncorrectable error memory; in response to the synchronous exception, obtain the initial exception context of the target task through a trusted firmware; modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the kernel; and execute exception handling based on the target exception context through the kernel.
[0189] Optionally, the above-mentioned processor can also execute the program code of the following steps: modifying the initial exception context through the trusted firmware to obtain a target exception context includes: modifying the exception type in the initial exception context to obtain a modified initial exception context; and determining the modified initial exception context as the target exception context.
[0190] Optionally, the above-mentioned processor can also execute the program code of the following steps: before executing exception handling based on the target exception context through the kernel, the method further includes: determining a target offset corresponding to an interrupt vector table in the kernel according to the exception level and stack pointer corresponding to the synchronous exception, and writing the target offset into an exception link register corresponding to the trusted firmware, where the interrupt vector table stores exception handling functions for executing exception handling; modifying the exception level and stack pointer in a saved run state register corresponding to the trusted firmware to obtain a modified saved run state register; and switching from the level where the trusted firmware is located to the level where the kernel is located based on the target offset in the exception link register and the modified saved run state register through an exception return instruction, and jumping to the interrupt vector table of the kernel to execute exception handling based on the target exception context through the kernel.
[0191] Optionally, the above-mentioned processor can also execute the program code of the following steps: restoring the data information of the target task in the stack corresponding to the trusted firmware to a general register, where after injecting the target exception context into the kernel, writing the data information of the target task in the general register into the stack corresponding to the kernel to execute exception handling based on the target exception context through the kernel.
[0192] Optionally, the above-mentioned processor can also execute the program code of the following steps: executing exception handling based on the target exception context through the kernel includes: judging whether the synchronous exception is a repairable exception based on the target exception context; if the synchronous exception is a repairable exception, determining the exception level of the trigger source corresponding to the synchronous exception based on the target exception context; and executing exception handling according to the exception level.
[0193] Optionally, the above-mentioned processor may also execute the program code of the following steps: When detecting that the target task is executed and triggering a synchronous exception, the method further includes: performing interruption processing on the target task; according to the exception level, performing exception handling includes: if the exception level is the level corresponding to the target task being a host user-mode task, sending a bus error signal to the host user-mode task through an exception handling function in the kernel; when the interruption of the host user-mode task is restored and the host user-mode task does not set a signal handling function for handling the bus error signal, ending the host user-mode task based on the bus error signal.
[0194] Optionally, the above-mentioned processor may also execute the program code of the following steps: According to the exception level, performing exception handling includes: if the exception level is the level corresponding to the target task being a host kernel-mode task, determining the target instruction address of the target task from the target exception context; according to the target instruction address, determining whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table; if there is no exception repair instruction address, performing a crash handling.
[0195] Optionally, after determining whether there is an exception repair instruction address corresponding to the target instruction address in the preset exception table, the method further includes: if there is an exception repair instruction address, determining the exception repair instruction address as the target instruction address of the target task in the target exception context to obtain a repaired target exception context; based on the repaired target exception context, restoring the host kernel-mode task.
[0196] Optionally, according to the exception level, performing exception handling includes: if the exception level is the level corresponding to the target task being a virtual machine task, sending a bus error signal to the virtual machine simulator process through an exception handling function in the kernel, and the virtual machine task is implemented based on the virtual machine simulator process; when the interruption of the virtual machine task is restored, determining whether there is a target function for handling the bus error signal in the virtual machine simulator process; if there is no target function, ending the virtual machine simulator process based on the bus error signal.
[0197] Optionally, after determining that there is a target function for handling the bus error signal in the virtual machine simulator process, the method further includes: if there is a target function, constructing a target exception type through the target function; according to the target exception type, injecting a virtual synchronous exception into the virtual machine kernel; through the virtual machine kernel, performing exception handling.
[0198] Those of ordinary skill in the art can understand that Figure 7 The structure shown is only schematic. The electronic device 70 can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a personal digital assistant, and terminal devices such as Mobile Internet Devices (MID) and PAD. Figure 7 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device 70 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 7 in the figure, or have a different configuration from that shown Figure 7 in the figure.
[0199] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0200] Embodiment 4
[0201] The embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above computer program product can be used to store the program code executed by the data processing method provided in the first embodiment above.
[0202] Embodiment 5
[0203] The embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product can be used to store the program code executed by the data processing method provided in the first embodiment above.
[0204] Optionally, in this embodiment, the above computer program product can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0205] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0206] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0207] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0208] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0209] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0210] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0211] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.< / x>
Claims
1. A data processing method, characterized in that, Including: When it is detected that the target task is executed, a synchronization exception is triggered, where the target task is a task corresponding to the consumption of uncorrectable error memory; In response to the synchronization exception, obtain the initial exception context of the target task through the trusted firmware; Modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the EL2 kernel; Execute exception handling through the EL2 kernel based on the target exception context; Among them, modifying the initial exception context through the trusted firmware to obtain a target exception context includes: modifying the exception type in the initial exception context according to the exception level corresponding to the synchronization exception to obtain the target exception context.
2. The method according to claim 1, wherein, Before executing exception handling through the EL2 kernel based on the target exception context, the method further includes: Determine the target offset corresponding to the interrupt vector table in the EL2 kernel according to the exception level and stack pointer corresponding to the synchronization exception, and write the target offset into the exception link register corresponding to the trusted firmware, where the interrupt vector table stores an exception handling function for executing exception handling; Modify the exception level and stack pointer in the saved run state register corresponding to the trusted firmware to obtain a modified saved run state register; Based on the target offset in the exception link register and the modified saved run state register through an exception return instruction, switch from the layer where the trusted firmware is located to the layer where the EL2 kernel is located, and jump to the interrupt vector table of the EL2 kernel to execute exception handling through the EL2 kernel based on the target exception context.
3. The method according to claim 1, wherein The method further includes: Restore the data information of the target task in the stack corresponding to the trusted firmware to the general register. After injecting the target exception context into the EL2 kernel, the EL2 kernel writes the data information of the target task in the general register into the stack corresponding to the EL2 kernel to execute exception handling through the EL2 kernel based on the target exception context.
4. The method according to any one of claims 1 to 3, characterized in that Executing exception handling through the EL2 kernel based on the target exception context includes: Judging whether the synchronization exception is a repairable exception based on the target exception context; If the synchronization exception is a repairable exception, determine the exception level of the trigger source corresponding to the synchronization exception based on the target exception context; Execute exception handling according to the exception level.
5. The method according to claim 4, wherein: After triggering a synchronization exception when it is detected that the target task is executed, the method further includes: Performing interrupt handling on the target task; Executing exception handling according to the exception level includes: If the exception level is the level corresponding to when the target task is a host user state task, send a bus error signal to the host user state task through the exception handling function in the EL2 kernel; When the interruption of the host user-mode task is restored and no signal processing function for handling the bus error signal is set in the host user-mode task, end the host user-mode task based on the bus error signal.
6. The method according to claim 4, wherein Performing exception handling according to the exception level includes: If the exception level is the level corresponding to the case where the target task is a host EL2 kernel-mode task, determine the target instruction address of the target task from the target exception context; According to the target instruction address, determine whether there is an exception repair instruction address corresponding to the target instruction address in a preset exception table; If the exception repair instruction address does not exist, perform a crash handling.
7. The method according to claim 6, characterized in that, After determining whether there is an exception repair instruction address corresponding to the target instruction address in the preset exception table, the method further includes: If the exception repair instruction address exists, determine the exception repair instruction address as the target instruction address of the target task in the target exception context, and obtain a repaired target exception context; Based on the repaired target exception context, resume the host EL2 kernel-mode task.
8. The method according to claim 4, characterized in that Performing exception handling according to the exception level includes: If the exception level is the level corresponding to the case where the target task is a virtual machine task, send a bus error signal to a virtual machine simulator process through an exception handling function in the EL2 kernel, and the virtual machine task is implemented based on the virtual machine simulator process; When the interruption of the virtual machine task is restored, determine whether there is a target function for handling the bus error signal in the virtual machine simulator process; If the target function does not exist, end the virtual machine simulator process based on the bus error signal.
9. The method according to claim 8, wherein After determining whether there is a target function for handling the bus error signal in the virtual machine simulator process, the method further includes: If the target function exists, construct a target exception type through the target function; According to the target exception type, inject a virtual synchronous exception into the virtual machine EL2 kernel; Through the virtual machine EL2 kernel, perform exception handling.
10. A data processing device, characterized in that, Including: A trigger unit, configured to trigger a synchronous exception when detecting that a target task is executed, where the target task is a task corresponding to the consumption of uncorrectable error memory; A first obtaining unit, configured to respond to the synchronous exception and obtain an initial exception context of the target task through a trusted firmware; A first modification unit, configured to modify the initial exception context through the trusted firmware to obtain a target exception context, and inject the target exception context into the EL2 kernel; A first execution unit, configured to perform exception handling through the EL2 kernel based on the target exception context; Wherein, modifying the initial exception context through the trusted firmware to obtain a target exception context includes: modifying the exception type in the initial exception context according to the exception level corresponding to the synchronous exception to obtain the target exception context.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the data processing method according to any one of claims 1 to 9.
12. An electronic device, characterized in that, Comprising: a memory storing an executable program; a processor for running the program, wherein when the program runs, it executes the data processing method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a stored computer program, and when the computer program is run by a processor, it implements the data processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Abnormality repair method and device and storage medium
CN115495278A