An error message processing method, device, and storage medium

After the memory error triggers an interrupt, collect and judge the memory area errors to avoid writing log information in the error area, the system downtime problem is solved and the stable operation of the system is achieved.

CN113536320BActive Publication Date: 2025-07-22LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110772948.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-08
Publication Date
2025-07-22
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

During the startup of the system boot program, when there is a serious memory error, it causes constant interruption, and ultimately leads to system downtime.

Method used

After the memory error triggers an interrupt, collect error information, obtain an error-free memory area for writing log information, and determine whether the area contains an error area when it is necessary to write the log information. If so, skip the writing step to avoid triggering the interrupt again.

Benefits of technology

Avoid system downtime caused by log information being written to the wrong memory area and ensures stable system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536320B_ABST
    Figure CN113536320B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus and storage medium for error information processing. After a memory error triggers a system interrupt (such as an MCE interrupt), the method collects error information of the memory error, including the memory area where the memory error occurs; then when log information needs to be written, the following operations are added to avoid triggering the system interrupt again due to the same memory error: obtaining the memory area for writing log information, and determining whether the memory area includes the memory area where the memory error occurs. If so, the step of writing the log information into the memory area is skipped. In this way, the system interrupt will not be triggered again when writing the log information into the memory area where the memory error occurs, thus avoiding system downtime caused by continuous generation of interrupts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information processing, and particularly to a method, apparatus, and storage medium for error information processing. Background Art

[0002] During the startup process of the system bootloader, the computer hardware is detected. When a serious error (fatal error) occurs in the memory of certain areas, subsequent operations will still access the memory where the fatal error has occurred, thus continuously generating interrupts and ultimately causing the system to crash. Summary of the Invention

[0003] The applicant of the present application creatively provides a method, apparatus, and storage medium for error information processing.

[0004] According to a first aspect of an embodiment of the present application, a method for error information processing is provided. The method includes the following operations when a memory error triggers a first interrupt: collecting error information of the memory error, where the error information includes a first memory area where the memory error occurs; obtaining a second memory area for writing log information, and determining whether the second memory area includes the first memory area. If so, the step of writing the log information into the second memory area is skipped.

[0005] According to an embodiment of the present application, before collecting the error information of the memory error, the method further includes: when a memory error triggers a first interrupt, recording the error information of the memory error.

[0006] According to an embodiment of the present application, obtaining a second memory area for writing log information includes: obtaining a second memory area for writing log information from the system memory defined in the error record serialization table.

[0007] According to an embodiment of the present application, after skipping the step of writing the log information into the second memory area, the method further includes: writing the log information into a backup memory area of the second memory area.

[0008] According to an embodiment of the present application, before writing the log information into the backup memory area of the first memory area, the method further includes: during system initialization, reserving a third memory area as the backup memory area of the second memory area.

[0009] According to an embodiment of the present application, the method further includes: obtaining a fourth memory area required for a first handler for processing the first interrupt, and determining whether the fourth memory area includes the first memory area. If so, marking the first memory area as unavailable.

[0010] According to an embodiment of the present application, after marking the first memory area as unavailable, the method further includes: obtaining a fifth memory area that does not include the first memory area; and allocating the fifth memory area for use by the first handler.

[0011] According to an embodiment of the present application, the first handler package crashes the kernel program. Correspondingly, marking the first memory area as unavailable includes: modifying the system memory mapping table provided by the kernel crash dump tool to mark the first memory area in the system memory mapping table as unavailable.

[0012] According to a second aspect of the embodiments of the present application, there is provided an error information processing device, including: an error information collection module, configured to collect error information of a memory error, where the error information includes a first memory area where the memory error occurs; and a log information writing module, configured to obtain a second memory area for writing log information, and determine whether the second memory area includes the first memory area. If so, skip the step of writing the log information to the second memory area.

[0013] According to a third aspect of the embodiments of the present application, there is provided a computer-readable storage medium, including a set of computer-executable instructions that, when executed, are used to execute the error information processing method in any one of the above.

[0014] The embodiments of the present application provide an error information processing method, device, and storage medium. After a memory error triggers a system interrupt (such as an MCE interrupt), the method collects error information of the memory error, including the memory area where the memory error occurs. Then, when log information needs to be written, the following operations are added to avoid triggering the system interrupt again due to the same memory error: obtaining the memory area for writing log information, and determining whether the memory area includes the memory area where the memory error occurs. If so, skip the step of writing the log information to the memory area. In this way, the system interrupt will not be triggered again when writing the log information to the memory area where the memory error occurs, thereby avoiding system downtime caused by continuous interrupts.

[0015] It should be understood that the implementation of the present application does not necessarily achieve all the above beneficial effects. Instead, specific technical solutions can achieve specific technical effects, and other embodiments of the present application can also achieve the beneficial effects not mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understood. In the drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, where:

[0017] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.

[0018] Figure 1 FIG. is a schematic flow chart of the implementation of an embodiment of the error information processing method of the present application;

[0019] Figure 2 FIG. is a schematic flow chart of the implementation of another embodiment of the error information processing method of the present application;

[0020] Figure 3 FIG. is a schematic structural diagram of the composition of an embodiment of the error information processing device of the present application. Detailed implementation manners

[0021] In order to make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0022] In the description of this specification, the descriptions referring to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0023] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0024] Figure 1 Shows the implementation process of an embodiment of the error information processing method of the present application. Refer to Figure 1, the method includes performing the following operations when a memory error triggers a first interrupt: Operation S110, collecting error information of the memory error, where the error information includes a first memory area where the memory error occurs; Operation S120, obtaining a second memory area for writing log information, and determining whether the second memory area contains the first memory area. If so, skip the step of writing the log information into the second memory area.

[0025] Among them, the first interrupt mainly refers to an operating system interrupt triggered by a memory error, and the error information processing method of the present application is also mainly executed by the operating system.

[0026] Furthermore, the error information processing method of the present application is usually executed by an interrupt handler corresponding to the first interrupt in the operating system. For example, assuming the first interrupt is a Machine Check Exception (MCE) interrupt, then the error information processing method of the present application is executed by an MCE interrupt handler corresponding to the MCE interrupt.

[0027] In Operation S110, the error information of the memory error is usually obtained by firmware through tools such as an Interrupt Handler from relevant CPU registers and placed in a register corresponding to the first interrupt when the first interrupt is triggered. For example, in an MCE register (MCE bank). Among them, common firmware includes: Unified Extensible Firmware Interface (UEFI), or Basic Input Output System (BIOS), etc.

[0028] In this way, when the operating system processes the first interrupt through the first interrupt handler, it can collect the error information of the memory error from the register corresponding to the first interrupt, including the first memory area where the memory error occurs.

[0029] The first memory area where the memory error occurs is usually a memory address range or a specific memory page number.

[0030] Once the first memory area where the memory error occurs is obtained, it can be compared with the memory area to be operated on subsequently. If the memory area to be operated on subsequently contains the first memory area, certain measures can be taken to avoid accessing the first memory area again in subsequent operations, thereby avoiding triggering the first interrupt again.

[0031] In operation S120, the second memory area for writing log information is obtained because when the operating system processes the first interrupt, it is usually necessary to write relevant log information, especially the error information that causes the interrupt, in the specified second memory area so that maintenance personnel can analyze the cause of the interrupt and fix related problems. If the second memory area contains the first memory area where a memory error occurs, the first interrupt will be triggered again when writing the log information.

[0032] Therefore, in the error information processing method provided in the embodiment of the present application, before writing the log information, the second memory area for writing the log information will be obtained first and compared with the first memory area where a memory error occurs, and the following judgment will be added: if the second memory area contains the first memory area, the step of writing the log information into the second memory area can be skipped, so as to avoid accessing the first memory area again and thus avoid triggering the first interrupt again.

[0033] It is not difficult to see that in the error information processing method of this embodiment, after a memory error triggers a system interrupt (such as an MCE interrupt), the first memory area where the memory error occurs is obtained through operation S110; then, before writing the log information is required, the second memory area for writing the log information is obtained first, and the processing logic for judging whether the second memory area contains the first memory area where the memory error occurs is added. Once it is found that the second memory area contains the first memory area where the memory error occurs, the step of writing the log information into this memory area is skipped.

[0034] In this way, the problem of system downtime caused by continuously triggering the first interrupt due to the log information writing operation during the processing of the first interrupt can be avoided.

[0035] It should be noted that Figure 1 The shown embodiment of the present application is only the most basic embodiment of an error information processing method of the present application, and implementers can further refine and expand it on this basis.

[0036] According to an embodiment of the present application, before collecting the error information of the memory error, the method further includes: when the memory error triggers the first interrupt, recording the error information of the memory error.

[0037] To ensure that the operating system can obtain the first memory area where a memory error occurs when processing the first interrupt, it is necessary to record the error information of the memory error before collecting the error information of the memory error.

[0038] As mentioned above, usually when a memory error triggers the first interrupt, the firmware will obtain the error information from the relevant CPU register through tools such as the interrupt handler and put it into the register corresponding to the first interrupt. Of course, different operating systems or firmware may also differ or change. The implementer can also use other methods, such as modifying the interrupt handler in the firmware, or adding a custom interrupt processing step, to record the error information of the memory error (including the first memory area where the memory error occurs) in the register or other data storage medium that can be accessed by the interrupt handler corresponding to the first interrupt in the operating system.

[0039] According to an embodiment of the present application, obtaining a second memory area for writing log information includes: obtaining the second memory area for writing log information from a system memory defined in an Error Record Serialization Table (ERST).

[0040] When the operating system and the firmware interact, they usually use the system memory defined in the ERST table to transmit information, that is, the second memory area for writing log information. Therefore, in this embodiment, the system memory defined in the ERST is used to obtain the second memory area for writing log information. In this way, the second memory area for writing log information can be obtained conveniently and quickly.

[0041] According to an embodiment of the present application, after skipping the step of writing the log information into the second memory area, the method further includes: writing the log information into a spare memory area of the second memory area.

[0042] Usually, after skipping the step of writing the log information into the second memory area, the log information related to the first interruption may no longer be obtained, which is very disadvantageous for repairing the related problems causing the first interruption.

[0043] In the implementation of the present application, after skipping the step of writing the log information into the second memory area, a spare memory area is obtained and the log information is written into the memory area, thereby ensuring that the first interrupt will not be triggered again and the log information can be retained to analyze the cause of the first interrupt and fix the problem.

[0044] The spare memory area can be determined dynamically and randomly, but it is necessary to ensure that the randomly assigned spare memory area is passed to the subsequent processing process so that log information can be obtained from the spare memory area; the spare memory area can also be pre-set, so there is no need to pass the spare memory area, and the pre-set spare memory area can be used directly to obtain log information.

[0045] According to an embodiment of the present application, before writing the log information into the backup memory area of the first memory area, the method further includes: reserving a third memory area as a backup memory area for the second memory area when the system is initialized.

[0046] In this embodiment, the backup memory area of the second memory area is a memory area reserved during system initialization, and the memory area can be set as the backup memory area of the second memory area during initialization.

[0047] In this way, when the second memory area includes the first memory area where an error occurs, the log information can be written into the spare memory area.

[0048] According to an embodiment of the present application, the method further includes: obtaining a fourth memory area required to be used by a first handler for processing the first interrupt, determining whether the fourth memory area includes the first memory area, and if so, marking the first memory area as unavailable.

[0049] Other handlers may also be called in the process of handling the first interrupt. For example, if the first interrupt is caused by an error with a higher error level that triggers a system restart (panic), in addition to recording log information, it may be necessary to perform a system restart. The system restart program is the other handler that needs to be called in this case, that is, the first handler.

[0050] If the first processing program also needs to access a memory area, and the memory area happens to include the first memory area where the memory error occurs, the first memory area will be accessed again, thereby triggering the first interrupt again.

[0051] Therefore, in this embodiment, in addition to detecting the second memory area where the log information is written, the memory area that the first processing program needs to use, that is, the fourth memory area, is also obtained. Afterwards, logic is added to determine whether the fourth memory area contains the first memory area where the memory error occurs. Once it is found that the fourth memory area contains the first memory area where the memory error occurs, corresponding measures are taken so that the first processing program does not use the first memory area where the memory error occurs, for example, marking the first memory area as unavailable.

[0052] In this way, the problem of system downtime caused by repeatedly accessing the first memory area where the error occurs and continuously triggering the first interrupt can be avoided due to calling the first processing program when processing the first interrupt.

[0053] According to an embodiment of the present application, after marking the first memory area as unavailable, the method further includes: acquiring a fifth memory area that does not include the first memory area; and allocating the fifth memory area to the first processing program for use.

[0054] As described above, once the first memory area is marked as unavailable, the fourth memory area that may contain the first memory area cannot meet the data storage needs of the first handler, resulting in the failure of the first handler to execute successfully.

[0055] In the embodiments of the present application, in this case, a fifth memory area that does not contain the first memory area is obtained; the fifth memory area is allocated for use by the first handler to ensure the smooth execution of the first handler.

[0056] According to an embodiment of the present application, the first handler package crashes the kernel program. Correspondingly, marking the first memory area as unavailable includes: modifying the system memory mapping table provided by the kernel crash dump tool to the crashed kernel program, and marking the first memory area in the system memory mapping table as unavailable.

[0057] Among them, the kernel crash dump tool refers to a system tool that can timely save the system state and important information at the time of the crash when the system crashes, such as Kdump.

[0058] The crashed kernel program refers to a standby kernel program that can be used when the main kernel program has crashed. For example, the crash kernel in the Linux operating system.

[0059] The system memory mapping table refers to a table that maps virtual memory or memory allocated for a specific purpose to the actual physical memory address. For example, the E820table provided by Kdump to the crash kernel.

[0060] It should be noted that the above embodiments of the error information processing method of the present application are only exemplary descriptions refined and extended on the basis of the Figure 1 shown basic embodiment. Implementers can also flexibly combine the above embodiments according to specific implementation requirements and implementation conditions to form new embodiments.

[0061] Figure 2 Another specific embodiment of the error information processing method of the present application is shown. This embodiment is applied to the scenario where the operating system exception MCE is triggered after a memory error is found during the hardware detection of the computer.

[0062] Such as Figure 2As shown in the figure, assume that the hardware management platform 21 (Hardware Platform), for example, the memory controller, detects the memory page frames 20 (Pageframe) and finds that there are two memory pages with errors: memory page 201 (address 0x2A3ED018) and memory page 202 (0xC400). Then it sends an MCE interrupt (INT18MCE) with interrupt number 18 to the processor 22 (Processor); the processor 22 generates an MSMI interrupt by pulling down the MSMI pin and sends it to the firmware 23 (Firmware); the firmware 23 is equipped with a system management interrupt (System Management Interrupt, SMI) processor, for example, SMIhandler231. SMIhandler231 will collect error information such as the memory address where the error occurred and store the collected error information in the MCE register; then, the firmware 23 sends an interrupt request (Interrupt ReQuest, IRQ) representing the MCE to the operating system kernel 24 (OS Kernel); after receiving this MCE IRQ, the operating system kernel 24 looks up the interrupt vector (X86_TRAP_MC) corresponding to the MCE in the interrupt vector table 25 (IDTTable) and obtains the entry address of the MCE interrupt handler 26 (MCEinterrupt flow); then, it calls the MCE interrupt handler 26 to execute the corresponding processing procedure. In terms of the program call relationship and execution order, the MCE interrupt handler mainly executes: machine_check->do_mce->machine_check_vector->do_machine_check. When it executes to do_machine_check, the MCE interrupt handler 26 will call the interrupt processor 27 (interrupt handler) to perform the following operations:

[0063] Operation S210: Read the memory address where the error occurred from the MCE register;

[0064] Among them, the memory address where the error occurred is the one that SMIhandler231 put into the MCE register when collecting the error information.

[0065] Operation S220: Determine whether the error level will trigger a system panic. If it will trigger, continue with operation S230. If it will not trigger, perform other corresponding processing;

[0066] Among them, the error level mainly refers to the error level that triggers the MCE interrupt, and different systems may have different definitions. But usually the system will define which error levels will trigger a system panic.

[0067] If the error level is relatively low and can be automatically corrected, the system can still continue to run. At this time, when using the corresponding memory area, it is very likely that there is no longer the error that occurred before, and the corresponding interrupt will not be triggered again. Therefore, no additional processing is required, and only waiting for the error to be self-repaired can reduce the processing steps and maximize the availability of the memory area.

[0068] If the error level is relatively high and causes the system to be unable to continue running, this error level is usually defined as the error level that triggers a system panic. At this time, once an error of this level is detected, an emergency plan will be triggered, such as the mce_panic program executed by operation S230, for rescue.

[0069] In the mce_panic program executed by operation S230, there are two operations: 1) Call the apei write interface to write log information in memory page 201; 2) Copy the error information to memory.

[0070] To avoid the above two operations from triggering an infinite loop due to an exception where memory pages 201 and 202 have errors again, the following detection and judgment logic is added in this embodiment before executing these two operations:

[0071] Operation S240, determine whether the memory address to be operated is the memory address with an error. If not, continue with operation S250; if so, continue with operation S260;

[0072] Among them, for the "apei write" operation, the memory address to be operated can be obtained through the system memory defined in ERST for handling MCE; for the operation of "copying error information", the memory address to be operated can be obtained through the E820 table provided by the kdump service to the crash kernel.

[0073] Operation S250, call the apei write interface to write log information;

[0074] Operation S260, skip writing log information, modify the E820 table, and set the memory page with a serious error to unavailable;

[0075] After that, operation S270 can be executed to start the Kdump system and call the crash kernel to copy the error information.

[0076] In this way, when the memory area used by the apei write interface to write log information contains the memory page where an error occurs, log information is no longer written to this memory page, and when the Kdump system is subsequently started and the crash kernel is called to copy the error information, the memory page with a serious error will no longer be used, thus avoiding the MCE exception of interrupt number 18 from being triggered again.

[0077] It should be noted that Figure 2 The illustrated embodiments are only exemplary descriptions of the error information processing method of the present application, rather than limitations on the implementation manners and application scenarios of the error information processing method of the present application. Implementers can adopt any applicable implementation manner according to specific implementation conditions and apply it to any applicable application scenario.

[0078] Furthermore, an embodiment of the present application further provides an error information processing device. As Figure 3 shown, the device 30 includes: an error information collection module 301, configured to collect error information of a memory error, where the error information includes a first memory area where a memory error occurs; a log information writing module 302, configured to obtain a second memory area for writing log information, and determine whether the second memory area includes the first memory area. If so, the step of writing log information to the second memory area is skipped.

[0079] According to an embodiment of the present application, the device 30 further includes: an error information recording module, configured to record error information of a memory error when the memory error triggers a first interrupt.

[0080] According to an embodiment of the present application, the log information writing module 302 is specifically configured to obtain a second memory area for writing log information from the system memory defined in the error record serialization table.

[0081] According to an embodiment of the present application, the log information writing module 302 is further configured to write log information to a backup memory area of the second memory area.

[0082] According to an embodiment of the present application, the device 30 further includes: a backup memory area reservation module, configured to reserve a third memory area as a backup memory area of the second memory area during system initialization.

[0083] According to an embodiment of the present application, the device 30 further includes: a memory area processing module, configured to obtain a fourth memory area required for a first processing program for processing the first interrupt, and determine whether the fourth memory area includes the first memory area. If so, the first memory area is marked as unavailable.

[0084] According to an embodiment of the present application, the device 30 further includes: a memory area acquisition module, configured to acquire a fifth memory area that does not include the first memory area; and a memory area allocation module, configured to allocate the fifth memory area for use by the first handler.

[0085] According to an embodiment of the present application, the first handler package crashes the kernel program. Correspondingly, the memory area processing module is specifically configured to modify the system memory mapping table provided by the kernel crash dump tool to the crashed kernel program, and mark the first memory area in the system memory mapping table as unavailable.

[0086] According to a third aspect of the embodiments of the present application, there is provided a computer-readable storage medium, where the storage medium includes a set of computer-executable instructions, which are used to execute any of the above error information processing methods when the instructions are executed.

[0087] It should be noted here that: the above descriptions of the embodiments of the error information processing device and the above descriptions of the embodiments of the computer-readable storage medium are similar to the descriptions of the foregoing method embodiments, and have beneficial effects similar to those of the foregoing method embodiments, so details will not be repeated. For the technical details not disclosed in the descriptions of the embodiments of the error information processing device and the embodiments of the computer-readable storage medium of the present application, please refer to the descriptions of the foregoing method embodiments of the present application for understanding. For the sake of saving space, details will not be repeated here.

[0088] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0089] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another device, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.

[0090] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0091] In addition, each functional unit in the embodiments of the present application may be fully integrated into one processing unit, or each unit may be separately used as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0092] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage media, read-only memory (ROM), magnetic disks or optical disks and other various media that can store program codes.

[0093] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the embodiments of the present application. And the foregoing storage medium includes: removable storage media, ROM, magnetic disks or optical disks and other various media that can store program codes.

[0094] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for processing error information, the method includes performing the following operations when a memory error triggers a first interrupt: Collect error information of the memory error, the error information including a first memory area where the memory error occurs; Obtain a second memory area for writing log information, determine whether the second memory area contains the first memory area, and if so, skip the step of writing the log information into the second memory area; Obtain a fourth memory area required by a first handler for processing the first interrupt, determine whether the fourth memory area contains the first memory area, and if so, mark the first memory area as unavailable.

2. The method according to claim 1, before collecting the error information of the memory error, the method further includes: When a memory error triggers a first interrupt, record the error information of the memory error.

3. The method according to claim 1, the obtaining the second memory area for writing log information includes: Determine the corresponding defined system memory based on an error record serialization table; Obtain the second memory area for writing log information from the system memory.

4. The method according to claim 1, after skipping the step of writing the log information into the second memory area, the method further includes: Write the log information into a backup memory area of the second memory area.

5. The method according to claim 4, before writing the log information into the backup memory area of the second memory area, the method further includes: Reserve a third memory area as the backup memory area of the second memory area during system initialization.

6. The method according to claim 1, after marking the first memory area as unavailable, the method further includes: Obtain a fifth memory area that does not contain the first memory area; Allocate the fifth memory area for use by the first handler.

7. The method according to claim 1, the first handler includes a crash kernel program, and correspondingly, marking the first memory area as unavailable includes: Modify the system memory mapping table provided by the kernel crash dump tool to the crash kernel program, and mark the first memory area in the system memory mapping table as unavailable.

8. An error information processing device, the device includes: An error information collection module, configured to collect error information of a memory error, the error information including a first memory area where the memory error occurs; A log information writing module, configured to obtain a second memory area for writing log information, determine whether the second memory area contains the first memory area, and if so, skip the step of writing the log information into the second memory area; A memory area processing module, configured to obtain a fourth memory area required by a first handler for processing the first interrupt, determine whether the fourth memory area contains the first memory area, and if so, mark the first memory area as unavailable.

9. A computer-readable storage medium, characterized in that, The storage medium includes a set of computer-executable instructions that, when executed, are used to perform the error message processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for reducing write latency for database logging utilizing multiple storage devices

    CN103443773A

  • Memory error processing method and device and server

    CN111625387A