Abnormal information processing method and device, electronic equipment and storage medium
By re-establishing the communication connection between the processing cores using non-maskable interrupts in the server, obtaining the abnormal information of the main operating system and generating diagnostic files, the problem that the server cannot obtain abnormal information in the non-debug phase is solved, and the fault location efficiency is improved.
Patent Information
- Application Number
- CN202510237595.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The server cannot obtain abnormal information of the operating system during the non-debug phase, resulting in the inability to directly locate the fault point and cause of the fault.
By re-establishing the communication connection between the first processing core and the second processing core in a server with an isomorphic multi-core CPU, the abnormal information of the main operating system in the second processing core is obtained, and a diagnostic file is generated.
It realizes the acquisition of operating system abnormal information in the non-debug phase, reduces manpower consumption and improves the efficiency of fault location.
Smart Images

Figure CN120179460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method, apparatus, electronic device, and storage medium for processing abnormal information. Background Art
[0002] With the rapid development of electronic technologies, the number of cores of the central processing unit (CPU) of a server is increasing. For example, dual-core, quad-core, and even 12-core CPUs are quite common. The operating system deployed on the CPU core (such as the Linux system), as an open-source operating system, is widely used in various fields such as computers, servers, embedded systems, mobile devices, and the Internet of Things. As a software system, when the Linux system is applied to products, various problems will inevitably occur, especially when encountering very serious problems, resulting in system crashes or failures.
[0003] In related technologies, after the Linux system crashes or fails, generally only the system is restarted. Although there is a debug serial port on the server during the debugging phase to print out abnormal information to locate the fault point and the cause of the fault, this debug serial port is not visible during the non-debugging phase of the server (such as at the customer site), so the fault point and the cause of the fault cannot be directly located. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for processing abnormal information, so as to at least solve the problem in related technologies that the server cannot obtain the abnormal information of the operating system during the non-debugging phase, and thus cannot directly locate the fault point and the cause of the fault.
[0005] This application provides a method for processing abnormal information, which is applied to a first processing core. The first processing core is connected to a second processing core. The first processing core and the second processing core are isomorphic, and different operating systems are deployed inside the cores. The method includes: when an exception occurs in the main operating system in the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, re-establishing the communication connection between the first processing core and the second processing core by setting a non-maskable interrupt; sending a first command to the second processing core, where the first command is used to instruct the second processing core to find the abnormal information recorded by the main operating system and feedback the abnormal information to the first processing core; receiving the abnormal information fed back by the second processing core and generating a diagnostic file for the main operating system based on the abnormal information; and when the main operating system in the second processing core restarts, sending the diagnostic file to the second processing core.
[0006] The present application also provides a processing device for abnormal information, which is applied to a first processing core. The first processing core is connected to a second processing core. The first processing core and the second processing core are isomorphic, and the operating systems deployed inside the cores are different. The processing device for abnormal information includes: a processing module, configured to, when an exception occurs in the main operating system in the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, re - establish the communication connection between the first processing core and the second processing core by setting a non - maskable interrupt; a transceiver module, configured to send a first command to the second processing core, where the first command is used to instruct the second processing core to find the abnormal information recorded by the main operating system and feedback the abnormal information to the first processing core; the processing module is further configured to receive the abnormal information fed back by the second processing core and generate a diagnostic file for the main operating system based on the abnormal information; the transceiver module is further configured to, when the main operating system in the second processing core restarts, send the diagnostic file to the second processing core.
[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above - mentioned processing methods for abnormal information when executing the computer program.
[0008] The present application also provides a computer - readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above - mentioned processing methods for abnormal information are implemented.
[0009] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above - mentioned processing methods for abnormal information are implemented.
[0010] Through the present application, when an exception occurs in the main operating system in the second processing core, and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, and the operating system in the first processing core is different from the main operating system, that is, when the main operating system in the second processing core crashes or has an exception, the operating system in the first processing core can still run normally. Since the first processing core and the second processing core are isomorphic, that is, the operating system in the first processing core and the main operating system in the second processing core access the same resources, therefore, the communication connection between the first processing core and the second processing core can be re - established through a non - maskable interrupt, and the abnormal information of the main operating system can be obtained through inter - core communication. Furthermore, a diagnostic file for the main operating system is generated based on the obtained abnormal information; when the main operating system in the second processing core restarts, the diagnostic file is sent to the second processing core so that the second processing core outputs the diagnostic file, solving the problem in the related art that the server cannot obtain the abnormal information of the operating system during the non - debugging stage, and thus cannot directly locate the fault point and the cause of the fault, greatly reducing the labor consumption and improving the efficiency of fault location. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0012] Figure 1 It is a flowchart for manually diagnosing faults in the operating system provided by the embodiments of the present application;
[0013] Figure 2 It is a topology diagram of a processing system for abnormal information provided by the embodiments of the present application;
[0014] Figure 3 It is an operation block diagram of a homogeneous dual-system provided according to the embodiments of the present application;
[0015] Figure 4 It is a flowchart of a method for processing abnormal information provided by the embodiments of the present application;
[0016] Figure 5 It is a schematic diagram of another method for processing abnormal information provided by the embodiments of the present application;
[0017] Figure 6 It is a schematic diagram of a method for remotely processing abnormal information provided by the embodiments of the present application;
[0018] Figure 7 It is a structural block diagram of a device for processing abnormal information provided by the embodiments of the present application;
[0019] Figure 8 It is a schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0021] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0022] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the drawings and specific embodiments.
[0023] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the abnormal information processing method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0024] The embodiments of this application are applied to a server with a homogeneous multi-core Central Processing Unit (CPU). Specifically, it is applied to the scenario where the main operating system of a homogeneous multi-core central processing unit crashes and abnormal information of the main operating system cannot be obtained in the non-debugging stage.
[0025] Homogeneous multi-core CPU: From a hardware perspective, all cores of the CPU or the CPU have the same architecture.
[0026] Homogeneous multi-core CPUs usually have two operating modes. One is the asymmetric multi-processing (AMP) mode; the other is the symmetric multi-processing (SMP) mode.
[0027] In the AMP mode, multiple cores of the CPU run different tasks relatively independently, and each core may run a different operating system or bare-metal program, or different versions of the operating system.
[0028] In the SMP mode, all cores of the CPU run one operating system.
[0029] In the related art, the existing fault diagnosis of the operating system of the processing core mainly relies on manual analysis. As Figure 1 shown, Figure 1 is the flowchart of the manual fault diagnosis of the operating system provided by the embodiments of this application. In Figure 1In the case where the CPU has problems during the non - debugging phase, generally, it is necessary to go through processes such as querying the scenario, setting up the problem machine environment, burning the problem version code, reproducing the problem phenomenon, analyzing the exception stack, deducing disassembly parameters, source code analysis, speculating on the cause of the fault, reproducing the fault through fault injection, and testing and verifying the solution. However, manual diagnosis of operating system faults often faces problems such as a large number and diverse types of faults. Especially when dealing with occasional problems or unclear problems, which may not occur for several weeks, it greatly increases the difficulty of solving problems. It requires strong manual professionalism, long time consumption, low efficiency, and also faces the dilemmas of repeated analysis and passive follow - up.
[0030] In addition, during the non - debugging phase of the CPU or server, it is impossible to obtain the exception information of the operating system, and it is also impossible to directly locate the fault point and the cause of the fault of the operating system.
[0031] To solve the above - mentioned technical problems, the embodiments of the present application provide a method for processing exception information. When an exception occurs in the main operating system within the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, a non - maskable interrupt is set to re - establish the communication connection between the first processing core and the second processing core. Then, based on this communication connection, the exception information recorded by the main operating system of the second processing core is obtained, and a diagnostic file of the main operating system is generated based on the exception information. When the main operating system within the second processing core restarts, the diagnostic file is sent to the second processing core so that the second processing core outputs the diagnostic file, which shortens processes such as "querying the scenario, setting up the problem machine environment, burning the problem version code, reproducing the problem phenomenon, analyzing the exception stack, and deducing disassembly parameters". It is possible to directly perform source code analysis based on the diagnostic file, solve the problems of difficult problem reproduction and difficult location, greatly reduce the human consumption, and enable users to quickly locate the fault point and the cause of the fault of the operating system based on the diagnostic file.
[0032] Next, taking Figure 2 the exception information processing system 200 shown as an example, the method provided by the embodiments of the present application will be described. Figure 2 It is only a schematic diagram and does not constitute a limitation on the applicable scenarios of the technical solutions provided by the present application.
[0033] As Figure 2 shown, Figure 2 is a topology diagram of an exception information processing system provided by the embodiments of the present application. Figure 2 In it, the exception information processing system 200 may include a first processing core 201, a second processing core 202, and a memory card 203. Optionally, the exception information processing system 200 may further include a remote debugging server 204.
[0034] In the embodiments of the present application, the first processing core 201 and the second processing core 202 can be any two processing cores in the same central processing unit. The first processing core and the second processing core are isomorphic, and the operating systems deployed inside the cores are different.
[0035] A real-time operating system is deployed inside the first processing core 201 as the secondary operating system of the server. When an external event or data is generated, this real-time operating system (RTOS) can receive and process it at a sufficiently fast speed, and the processed result can control the production process or make a quick response to the processing system within a specified time, schedule all available resources to complete real-time tasks, and control all real-time tasks to run in coordination. It is an operating system with the characteristics of timely response and high reliability. The RTOS system detects the running state of the main operating system (Linux system) as the secondary operating system. Among them, the RTOS system can also be a freeftos system.
[0036] A main operating system is deployed inside the second processing core 202. For example, the main operating system can be a Linux system. The main operating system is a Unix-like operating system and is also a multi-user, multi-tasking, multi-threaded, and multi-CPU operating system based on POSIX. The main operating system runs business normally.
[0037] In one example, as Figure 3 shown, Figure 3 is a block diagram of the operation of an isomorphic dual-system provided according to the embodiments of the present application. When the server starts to power on, bootrom boots Uboot; Uboot wakes up the first processing core in the CPU of the server to load the RTOS system, wakes up the second processing core to load the Linux system, the Linux system loads the remaining processing cores to run the Linux system, and the Linux system runs business after it starts up. The RTOS system monitors the running state of the Linux system through inter-core communication.
[0038] The memory card 203 in the embodiments of the present application can be any kind of memory card. The memory card is used to store diagnostic files and disassembly files.
[0039] It can be understood that the first processing core 201, the second processing core 202, and the memory card 203 belong to the same server.
[0040] The remote debugging server 204 in the embodiments of the present application can be any device with computing and communication functions. For example, the remote debugging server 204 can be a server, a cloud server, or a virtual machine. There is a communication connection between the remote debugging server 204 and the second processing core 202.
[0041] Figure 2The processing system 200 for the abnormal information shown is only for illustration and is not used to limit the technical solution of the present application. Those skilled in the art should understand that in the specific implementation process, the processing system 200 for the abnormal information may further include other processing cores in addition to the first processing core 201 and the second processing core 202, which is not limited.
[0042] According to an embodiment of the present application, an embodiment of a method for processing abnormal information is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0043] In this embodiment, a method for processing abnormal information is provided, which can be used in the above-mentioned first processing core. Figure 4 The flowchart of a method for processing abnormal information provided by an embodiment of the present application is shown in Figure 4 As shown, the process includes the following steps:
[0044] S401: When an exception occurs in the main operating system in the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, re-establish the communication connection between the first processing core and the second processing core by setting a non-maskable interrupt.
[0045] Among them, the occurrence of an exception in the main operating system can also be referred to as the main operating system crashing or an inevitable fatal exception occurring.
[0046] A non-maskable interrupt (NonMaskable Interrupt, NMI) is an unmaskable emergency interrupt with the highest priority, which is used to handle serious system exceptions such as CPU errors and memory failures.
[0047] It can be understood that since the first processing core and the second processing core default to use a shared interrupt connection, the disconnection of the connection between the first processing core and the second processing core means that the shared interrupt used by the first processing core and the second processing core cannot communicate. Therefore, it is necessary to set the NMI interrupt to re-establish the communication connection between the first processing core and the second processing core.
[0048] When an inevitable fatal exception occurs in the main operating system, there will be two phenomena. The first phenomenon is that the main operating system hangs without any response and can only wait for the watchdog to restart. The second phenomenon is that the fatal exception of the main operating system is detected by the main operating system itself, triggering the panic mechanism and restarting the main operating system.
[0049] Regarding the first phenomenon:
[0050] In some alternative embodiments, the first processing core sends a heartbeat packet to the second processing core to detect whether the second processing core returns response information; if no response information is received within a preset time period, it is determined that the main operating system has crashed.
[0051] Among them, the crash of the main operating system causes the connection between the first processing core and the second processing core to be disconnected.
[0052] The preset time period can be a time period without response due to heartbeat timeout set according to actual needs. For example, the preset time period can be 5 seconds.
[0053] In one example, the first processing core sends a heartbeat packet to the second processing core for the first time; if the second processing core does not return response information, it continues to send a heartbeat packet to the second processing core for the second time after a first time period; if the second processing core does not return response information, it continues to send a heartbeat packet to the second processing core for the third time after the first time period; if the second processing core still does not return response information, it is determined that no response information is received within the preset time period. Among them, the first time period can be set according to actual needs. For example, the first time period can be 1 second.
[0054] It can be understood that when the second processing core does not return heartbeat response information three times, it means that the main operating system has crashed, that is, the communication connection established between the first processing core and the second processing core is disconnected, and the inter-core communication interruption is switched or set to an NMI interruption.
[0055] Regarding Phenomenon 2:
[0056] In some alternative embodiments, if the first processing core receives response information within a preset time period, it receives the exception information of the main operating system sent by the second processing core, and the exception information is obtained by triggering an error handling mechanism after an exception occurs in the main operating system.
[0057] Among them, the error handling mechanism is the panic mechanism.
[0058] It can be understood that a fatal exception of the main operating system is detected by the main operating system itself, the panic mechanism is triggered, and the exception information is fed back to the first processing core.
[0059] S402: Send a first command to the second processing core.
[0060] Among them, the first command is used to instruct the second processing core to find the exception information recorded by the main operating system and feed back the exception information to the first processing core.
[0061] It can be understood that after the first processing core re-establishes the communication connection between the first processing core and the second processing core by setting a non-maskable interrupt, the first command can be sent through inter-core communication.
[0062] S403: Receive the exception information fed back by the second processing core, and generate a diagnostic file for the main operating system based on the exception information.
[0063] Among them, the exception information includes register information and stack information. The register information includes program counter information, link register information, and general register information.
[0064] The diagnostic file is also called a panic_code file.
[0065] In some alternative embodiments, the first processing core loads the disassembly file of the main operating system; and generates a diagnostic file for the main operating system according to the register information, stack information, and disassembly file.
[0066] In some alternative embodiments, the first processing core locates the fault code position of the currently running main operating system in the disassembly file according to the program counter information and the first corresponding relationship; locates the upper-level code position of the fault code in the disassembly file according to the link register information and the second corresponding relationship; determines the call function code position of the call function associated with the fault code and the variable values of the call function according to the stack information and general register information; and generates a diagnostic file with the fault code position, upper-level code position, call function code position, and variable values.
[0067] Among them, the first corresponding relationship is the corresponding relationship between the program counter information and the fault code position. The second corresponding relationship is the corresponding relationship between the link register information and the upper-level code position.
[0068] In one example, the first processing core saves the diagnostic file in the memory card.
[0069] It can be understood that the first processing core obtains the register information and stack information of the second processing core through inter-core communication, accesses the stack memory according to the register information and stack information, and then traces back the code running process according to the disassembly file of the main operating system in the read-only memory of the CPU, and generates a diagnostic file including the fault code position, upper-level code position, call function code position, and variable value information. That is, the diagnostic file may include fault information, which includes the fault point, variable values at each fault point, stack information, register information, and approximate fault cause, etc.
[0070] S404: When the main operating system in the second processing core restarts, send the diagnostic file to the second processing core.
[0071] Understandably, after the main operating system in the second processing core starts up, it runs a self-diagnosis process to detect whether a diagnostic file (panic_code file) is stored in the memory card; if so, it outputs the diagnostic file through the web interface. For example, developers or maintenance personnel can obtain the diagnostic file through the web interface. Based on the diagnostic file, developers can infer the cause of the main operating system failure, inject errors, and modify and test the corresponding solutions.
[0072] In some alternative embodiments, before the main operating system in the second processing core restarts, the first processing core sends a second command to the second processing core.
[0073] Wherein, the second command is used to instruct the second processing core to restart the main operating system and output a diagnostic file.
[0074] Through the method described above Figure 4 When an exception occurs in the main operating system in the second processing core, and the first processing core determines that the connection between the first processing core and the second processing core is disconnected due to the exception, and the operating system in the first processing core is different from the main operating system, that is, when the main operating system in the second processing core crashes or has an exception, the operating system in the first processing core can still run normally. Since the first processing core and the second processing core are isomorphic, that is, the operating system in the first processing core and the main operating system in the second processing core access the same resources, therefore, the communication connection between the first processing core and the second processing core can be re-established through a non-maskable interrupt, and the exception information of the main operating system can be obtained through inter-core communication. Furthermore, a diagnostic file of the main operating system is generated based on the obtained exception information; when the main operating system in the second processing core restarts, the diagnostic file is sent to the second processing core so that the second processing core outputs the diagnostic file, solving the problem in the related art that the server cannot obtain the exception information of the operating system during the non-debugging stage, and thus cannot directly locate the fault point and the cause of the fault, greatly reducing the manpower consumption and improving the efficiency of fault location.
[0075] In this embodiment, another method for processing exception information is provided. Figure 5 FIG. is a schematic diagram of another method for processing exception information provided by an embodiment of the present application, as Figure 5 shown. An RTOS system runs in the first processing core, and a Linux system runs in the second processing core. In this schematic diagram:
[0076] S501: The Linux system crashes or freezes abnormally.
[0077] S502: The Linux system detects whether it returns a response message.
[0078] S503: If so, the Linux system triggers a panic and returns exception information.
[0079] S504: If not, the Linux system hangs.
[0080] S505: The RTOS system executes the inter-core communication management task. The inter-core communication management task means that when the shared interrupt cannot communicate, the NMI interrupt is used.
[0081] S506: The RTOS system sends a heartbeat packet to detect the running state of the Linux system in real time.
[0082] S507: The RTOS system detects a timeout, and the Linux system does not return response information within a preset time period, determining that the Linux system hangs.
[0083] S508: The RTOS system switches the inter-core communication interrupt to the NMI interrupt and sends the first command to the Linux system.
[0084] S509: The Linux system receives the first command and obtains the exception information.
[0085] S510: The Linux system returns the exception information to the RTOS system.
[0086] S511: The RTOS system receives the exception information.
[0087] S512: The RTOS system loads the disassembly file of the Linux system.
[0088] S513: The RTOS system retrieves the fault point according to the program counter, link register and disassembly file in the exception information.
[0089] S514: The RTOS system performs function call tracing and function temporary variable assignment according to the stack information in the exception information, and converts the corresponding machine language into the assembly language file during the execution process.
[0090] S515: The RTOS system generates a diagnostic file from the fault point and the assembly language file, and resets and restarts the Linux system through the watchdog or hardware signal control system.
[0091] S516: The Linux system restarts.
[0092] S517: The Linux system accesses the diagnostic file and outputs the diagnostic file through the web and other human-computer interaction maintenance interfaces for maintenance personnel to view.
[0093] S518: The Linux system speculates on the cause of the fault based on the diagnostic file.
[0094] S519: The Linux system performs fault injection and reproduction.
[0095] S520: The Linux system conducts solution testing and verification.
[0096] S521: The Linux system resolves fault problems.
[0097] For the specific implementation manners of the above steps S501 - S521, reference can be made to the steps S401 - S404 of the foregoing embodiments, which will not be elaborated herein.
[0098] Furthermore, in a scenario where it is not easy to speculate on the cause of a fault, and even if there is a fault point, it is difficult to quickly find the cause of the fault and remote debugging is required to find the cause of the fault, the first processing core can also receive a network packet instruction forwarded by the second processing core; in response to the network packet instruction, establish a remote communication connection between the first processing core and the remote debugging server; when the main operating system has an exception, initialize the network interface and take over the network; generate and send a third command to the remote debugging server.
[0099] Among them, the network packet instruction is sent by the remote debugging server to the second processing core, and the network packet instruction is used to instruct to pause and restart when the main operating system has an exception.
[0100] The third command is used to notify the remote debugging server that the main operating system has an exception and instruct the remote debugging server to send a debugging instruction to the first processing core. The debugging instruction is used to instruct the first processing core to obtain exception information.
[0101] It can be understood that the method for the first processing core to obtain exception information has been described in the above S401 - S403, and will not be elaborated herein.
[0102] Specifically, as Figure 6 shown, Figure 6 is a schematic diagram of the remote exception information processing method provided by the embodiment of the present application. In Figure 6 , the RTOS system in the first processing core monitors the running state of the Linux system in the second processing core in real time; when it detects that the Linux system hangs, re-initialize the network interface and take over the network; notify the remote debugging server that the Linux system hangs.
[0103] The Linux system in the second processing core enables the remote debugging function; the Linux system hangs; receives the debugging instruction sent by the first processing core, obtains the exception information of the main operating system, and sends the exception information to the RTOS system through inter-core communication. The RTOS system then sends the exception information to the remote debugging server through network packets.
[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0105] An embodiment of the present application further provides a processing device for abnormal information, which is applied to a first processing core, as Figure 7 shown, Figure 7 is a structural block diagram of a processing device for abnormal information provided by an embodiment of the present application; the device includes: a processing module 701, configured to, when an exception occurs in the main operating system in the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, re-establish a communication connection between the first processing core and the second processing core by setting a non-maskable interrupt.
[0106] A transceiver module 702 is further configured to send a first command to the second processing core, and the first command is used to instruct the second processing core to find abnormal information recorded by the main operating system and feedback the abnormal information to the first processing core.
[0107] The processing module 701 is further configured to receive the abnormal information fed back by the second processing core and generate a diagnostic file of the main operating system based on the abnormal information.
[0108] The transceiver module 702 is further configured to send the diagnostic file to the second processing core after the main operating system in the second processing core restarts.
[0109] In some alternative embodiments, the abnormal information includes register information and stack information; the processing module 701 is specifically configured to load the disassembly file of the main operating system; generate a diagnostic file of the main operating system according to the register information, stack information, and disassembly file.
[0110] In some alternative embodiments, the register information includes program counter information, link register information, and general register information; the processing module 701 is specifically configured to find the fault code position of the currently running main operating system from the disassembly file according to the program counter information and a first correspondence relationship, where the first correspondence relationship is the correspondence relationship between the program counter information and the fault code position; find the upper-level code position of the fault code from the disassembly file according to the link register information and a second correspondence relationship, where the second correspondence relationship is the correspondence relationship between the link register information and the upper-level code position; determine the call function code position and variable values of the call function associated with the fault code according to the stack information and general register information; generate a diagnostic file from the fault code position, upper-level code position, call function code position, and variable values.
[0111] In some alternative embodiments, the processing module 701 is further specifically configured to send a heartbeat packet to the second processing core to detect whether the second processing core returns response information; if no response information is received within a preset time period, it is determined that the main operating system has crashed, and the crash of the main operating system causes the connection between the first processing core and the second processing core to be disconnected.
[0112] In some alternative embodiments, the transceiver module 702 is further configured to, if response information is received within a preset time period, receive the exception information of the main operating system sent by the second processing core, where the exception information is obtained by triggering an error handling mechanism after an exception occurs in the main operating system.
[0113] In some alternative embodiments, before the main operating system in the second processing core restarts, the transceiver module 702 is further configured to send a second command to the second processing core, and the second command is used to instruct the second processing core to restart the main operating system and output a diagnostic file.
[0114] In some alternative embodiments, the transceiver module 702 is further configured to receive a network packet instruction forwarded by the second processing core, where the network packet instruction is sent by a remote debugging server to the second processing core, and the network packet instruction is used to instruct to pause the restart when an exception occurs in the main operating system; the processing module 701 is further configured to, in response to the network packet instruction, establish a remote communication connection between the first processing core and the remote debugging server; the processing module 701 is further configured to initialize the network interface and take over the network when an exception occurs in the main operating system; the transceiver module 702 is further configured to generate and send a third command to the remote debugging server, and the third command is used to instruct the remote debugging server to send a debugging instruction to the first processing core, and the debugging instruction is used to instruct the first processing core to obtain exception information.
[0115] For the description of the features in the corresponding embodiments of the apparatus for processing exception information, reference may be made to the relevant description in the corresponding embodiments of the method for processing exception information, which will not be elaborated here one by one.
[0116] An embodiment of the present application further provides an electronic device, such as Figure 8 shown Figure 8 is a schematic hardware structure diagram of the electronic device provided in the embodiment of the present application. The electronic device includes a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above embodiments of the method for processing exception information.
[0117] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any of the above embodiments of the method for processing exception information when running.
[0118] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0119] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the method for processing abnormal information.
[0120] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the method for processing abnormal information.
[0121] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0122] The above has introduced in detail a method, device, electronic device, and storage medium for processing abnormal information provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for processing abnormal information, characterized in that: Applied to a first processing core, the first processing core is connected to a second processing core, the first processing core and the second processing core are isomorphic, and different operating systems are deployed in the cores; the method includes: When an exception occurs in the main operating system in the second processing core, and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, reestablishing the communication connection between the first processing core and the second processing core by setting a non-maskable interrupt; Sending a first command to the second processing core, where the first command is used to instruct the second processing core to search for exception information recorded by the main operating system and feed the exception information back to the first processing core; receiving the abnormal information fed back by the second processing core, and generating a diagnostic file of the main operating system based on the abnormal information; When the main operating system in the second processing core is restarted, the diagnostic file is sent to the second processing core.
2. The method for processing abnormal information according to claim 1, characterized in that: The abnormal information includes register information and stack information; The generating of the diagnostic file of the main operating system based on the abnormal information includes: Loading the disassembly file of the main operating system; The diagnostic file of the main operating system is generated according to the register information, the stack information and the disassembly file.
3. The method for processing abnormal information according to claim 2, characterized in that: The register information includes program counter information, link register information and general register information; The step of generating the diagnostic file of the main operating system according to the register information, the stack information and the disassembly file comprises: searching, from the disassembled file, for a fault code position of a fault code of the currently running main operating system according to the program counter information and a first corresponding relationship, wherein the first corresponding relationship is a corresponding relationship between the program counter information and the fault code position; searching the previous level code position of the fault code from the disassembled file according to the link register information and the second corresponding relationship, wherein the second corresponding relationship is the corresponding relationship between the link register information and the previous level code position; Determine, according to the stack information and the general register information, a calling function code position of a calling function associated with the fault code and a variable value of the calling function; The fault code position, the upper level code position, the calling function code position and the variable value are used to generate the diagnostic file.
4. The method for processing abnormal information according to claim 1, characterized in that: An exception occurs in the main operating system in the second processing core, and determining that the connection between the first processing core and the second processing core is disconnected due to the exception includes: Sending a heartbeat packet to the second processing core to detect whether the second processing core returns response information; If the response information is not received within a preset time period, it is determined that the main operating system is hung, and the main operating system hangs up, resulting in a disconnection between the first processing core and the second processing core.
5. The method for processing abnormal information according to claim 4, characterized in that: The method further comprises: If the response information is received within the preset time period, the exception information of the main operating system sent from the second processing core is received, where the exception information is obtained by triggering an error handling mechanism after an exception occurs in the main operating system.
6. The method for processing abnormal information according to claim 5, characterized in that: Before the main operating system in the second processing core is restarted, the method further includes: A second command is sent to the second processing core, where the second command is used to instruct the second processing core to restart the main operating system and output the diagnostic file.
7. The method for processing abnormal information according to claim 6, characterized in that: The method further comprises: receiving a network message instruction forwarded from the second processing core, the network message instruction being sent to the second processing core by a remote debugging server, and the network message instruction being used to instruct the main operating system to suspend and restart when an exception occurs; In response to the network message instruction, establishing a remote communication connection between the first processing core and the remote debugging server; When an abnormality occurs in the main operating system, the network port is initialized and the network is taken over; A third command is generated and sent to the remote debugging server, where the third command is used to instruct the remote debugging server to send a debugging instruction to the first processing core, where the debugging instruction is used to instruct the first processing core to obtain the exception information.
8. A device for processing abnormal information, characterized in that: Applied to a first processing core, the first processing core is connected to a second processing core, the first processing core and the second processing core are isomorphic, and different operating systems are deployed in the cores; The abnormal information processing device comprises: a processing module, configured to, when an exception occurs in a main operating system in the second processing core and it is determined that the connection between the first processing core and the second processing core is disconnected due to the exception, reestablish the communication connection between the first processing core and the second processing core by setting a non-maskable interrupt; a transceiver module, configured to send a first command to the second processing core, wherein the first command is configured to instruct the second processing core to search for abnormal information recorded by the main operating system and to feed back the abnormal information to the first processing core; The processing module is further used to receive the abnormal information fed back by the second processing core, and generate a diagnostic file of the main operating system based on the abnormal information; The transceiver module is further configured to send the diagnostic file to the second processing core after the main operating system in the second processing core is restarted.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for processing abnormal information as claimed in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for processing abnormal information according to any one of claims 1 to 7.