Virtual machine crashing event processing method and device and storage medium
Through the first virtual machine assisting the second virtual machine to restore resources and restart, the problem of abnormal machines caused by virtual machine crashes is solved, reducing user perception and improving user experience.
Patent Information
- Application Number
- CN202311479735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-13
AI Technical Summary
The crash of virtual machines causes abnormalities in the entire electronic device. The existing technology is usually solved by force restarting the entire machine, but this method affects the user experience.
The first virtual machine that did not have a crash event assisted the second virtual machine that had a crash event to restore normality, release and redistribute resources, and avoid restarting the entire machine.
Reduces user perception, improves user experience, and avoids the black screen and startup delay caused by the entire machine restart.
Smart Images

Figure CN119987934A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtualization technology, and in particular to a method, device and storage medium for processing a virtual machine crash event. Background Art
[0002] As virtualization technology is increasingly used, the system on chip (SOC) architecture used in current electronic devices makes extensive use of virtual machine architecture, which not only isolates business scenarios but also facilitates decoupling. However, the crash of the virtual machine may also cause the entire electronic device to malfunction.
[0003] At present, the processing of virtual machine crashes is usually to force the entire machine to restart so that the electronic device can resume normal use. However, this forced restart of the entire machine will cause the electronic device to have a black screen, and then the startup screen will flash again before entering the operating system and displaying the desktop. This not only slows down the processing, but also affects the user experience. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a method, device and storage medium for processing virtual machine crash events, aiming that when a crash event occurs in a hardware layer virtual machine, a virtual machine that has not experienced a crash event and hosts an operating system assists the virtual machine that has experienced a crash event to recover to normal without triggering a restart of the entire machine, thereby reducing user perception and improving user experience.
[0005] In the first aspect, the present application provides a method for handling a virtual machine crash event, which is applied to an electronic device. A first virtual machine and a second virtual machine are running on the electronic device, the first virtual machine is a virtual machine that carries the operating system of the electronic device, and the second virtual machine is a virtual machine for the hardware layer of the electronic device. The method includes: in response to a crash event of the second virtual machine, releasing the first resource occupied by the second virtual machine, and the first virtual machine taking over the first resource; reallocating memory resources for the second virtual machine and restarting the second virtual machine; after the second virtual machine is successfully restarted, releasing the first resource taken over by the first virtual machine; after the second virtual machine obtains the first resource, rebuilding the service link between the second virtual machine and the service that the second virtual machine is responsible for, and restoring the processing capability of the service.
[0006] Taking the operating system of the electronic device as an Android system as an example, the first virtual machine is a virtual machine that carries the Android system. In some implementations, it can be called a primary virtual machine (PVM), which can be responsible for most of the business processing logic of the Android system.
[0007] Among them, the second virtual machine facing the hardware layer can be described as a hardware layer virtual machine in some implementations, such as the hardware layer virtual machine A listed in this application, which is responsible for fingerprint, face, gesture and other recognition and processing services; for example, the hardware layer virtual machine B listed in this application, which is responsible for electronic identity identification, mobile POS (Point of sales) and other services; for example, the hardware layer virtual machine C listed in this application, which is responsible for the system reliability of the central processing unit (CPU) of the electronic device.
[0008] The crash event of the second virtual machine is specifically a common crash problem in Android.
[0009] Among them, sensing the crash of the second virtual machine, notifying the second virtual machine to release the first resource, the first virtual machine taking over the first resource, and reallocating memory resources to the second virtual machine to trigger the restart of the second virtual machine can be implemented by the virtual machine management service (Hyper-V).
[0010] The virtual machine management service is a tool that can be used to create virtual machines, shut down virtual machines, start virtual machines, and migrate virtual machines.
[0011] Therefore, when the second virtual machine crashes, the first virtual machine that has not crashed and carries the operating system temporarily takes over the resources occupied by the second virtual machine, restarts the second virtual machine alone, and after the second virtual machine restarts, the resources of the second virtual machine taken over by the first virtual machine are returned to the second virtual machine. In this way, only the second virtual machine that crashes is self-recovered, and the electronic device is not triggered to restart the entire machine, thereby reducing user perception and improving user experience.
[0012] According to the first aspect, releasing the first resource occupied by the second virtual machine includes: closing the input and output IO resources and interrupt resources of the first driver, where the first driver is the driver to be called by the second virtual machine when processing business; and clearing the contents of the queues corresponding to the IO resources and interrupt resources.
[0013] The first driver is related to the second virtual machine where the crash occurs.
[0014] Exemplarily, when the second virtual machine that crashes is a virtual machine responsible for fingerprint, face, gesture and other recognition and processing services, the first driver may include a touch driver, a display driver, a camera driver and the like.
[0015] Therefore, when the second virtual machine crashes, the IO resources and interrupt resources of each driver called by the second virtual machine are closed, and the contents in the queues corresponding to the IO resources and interrupt resources are cleared, thereby releasing the first resources occupied by the second virtual machine, avoiding memory leaks, and further causing performance degradation or even crash of other virtual machines that have not crashed.
[0016] According to the first aspect, or any implementation method of the first aspect above, clearing the contents in the queues corresponding to the IO resources and interrupt resources includes: clearing the events in the first-in-first-out queues corresponding to the IO resources and interrupt resources, and clearing the messages in the message queues corresponding to the IO resources and interrupt resources.
[0017] Thereby, various resources occupied by the second virtual machine are released.
[0018] According to the first aspect, or any implementation of the first aspect above, after clearing the contents in the queues corresponding to the IO resources and the interrupt resources, the method further includes: reclaiming the memory resources occupied by the second virtual machine.
[0019] Therefore, by reclaiming the memory resources occupied by the second virtual machine, the memory pressure is relieved, thereby ensuring the normal operation of other virtual machines and programs and the normal processing of services.
[0020] According to the first aspect, or any implementation of the first aspect above, the method further includes: taking over the reclaimed memory resources by the first virtual machine.
[0021] According to the first aspect, or any implementation method of the first aspect above, reallocate memory resources for the second virtual machine and restart the second virtual machine, including: reallocate memory resources for the second virtual machine and re-establish the mapping between the second virtual machine and the memory resources; call the second driver to load the image of the second virtual machine; restart the operating system running on the second virtual machine.
[0022] Thus, the second virtual machine that crashed is restarted.
[0023] According to the first aspect, or any implementation of the first aspect above, in response to a crash event of the second virtual machine, the method further includes: obtaining data in a memory resource corresponding to the second virtual machine, and dumping the data to a storage medium of the electronic device.
[0024] When the second virtual machine crashes, the data in the memory resources corresponding to the second virtual machine obtained are dynamic or volatile data, that is, data that will be lost after use or when an exception occurs.
[0025] Dump refers to dumping dynamic data in memory resources into a static form, such as dumping into a file (hereinafter referred to as a dump file), and then storing it in a storage medium of an electronic device.
[0026] The storage medium may be an internal memory of the electronic device or a connected external memory.
[0027] Therefore, when the second virtual machine crashes, by collecting and dumping data related to the second virtual machine, it is convenient to locate the problem later according to the dump file in the storage medium.
[0028] According to the first aspect, or any implementation of the first aspect above, the method also includes: in the process of creating the first virtual machine and the second virtual machine, configuring monitoring for the first virtual machine and the second virtual machine; using monitoring to collect system event information of the first virtual machine and the second virtual machine, the system event information includes an event type; when the event type indicates that the system event currently occurring in the second virtual machine is a crash event, executing steps in response to the crash event occurring in the second virtual machine.
[0029] The creation of the first virtual machine and the second virtual machine is achieved, for example, through a virtual machine management service (Hyper-V).
[0030] The configuration of monitoring, interaction with the created first virtual machine and the second virtual machine, and obtaining the system event information collected by monitoring to determine whether a crash event occurs are all implemented by the virtual machine management service.
[0031] According to the first aspect, or any implementation of the first aspect above, the method further includes: when the event type indicates that the system event currently occurring in the first virtual machine is a crash event, in response to the crash event occurring in the first virtual machine, restarting the entire electronic device.
[0032] The whole machine restart specifically refers to the restart at the operating system level, that is, the restart method of directly restarting the operating system or reloading the system kernel without re-powering on. Since the operating system needs to be restarted or the system kernel needs to be reloaded, a black screen will appear, and then the restart animation will be displayed, and then the process of entering the operating system will occur.
[0033] According to the first aspect, or any implementation method of the first aspect above, when the event type indicates that the current system event occurring in the first virtual machine is a crash event, the entire electronic device is restarted in response to the crash event occurring in the first virtual machine, including: when the crash event is a panic exception, the entire electronic device is restarted in response to the crash event occurring in the first virtual machine.
[0034] Panic is a way of handling errors. It is used to indicate that an error that cannot be handled has occurred, causing the program to be unable to continue. The occurrence of panic will force the program to stop running, and developers need to handle it effectively. Otherwise, the program may crash and affect the normal business process.
[0035] Since the first virtual machine carries most of the business logic of the Android system, when the first virtual machine panics, in order to ensure the normal use of the Android system, the entire machine needs to be restarted to restore to normal. For other types of abnormalities, such as abnormalities in Bluetooth and cellular network functions, they can be restored by restarting the function without restarting the entire machine.
[0036] Therefore, when a panic occurs in the first virtual machine, the electronic device is controlled to restart the entire machine, which greatly reduces the probability of restarting the entire machine.
[0037] In a second aspect, the present application provides an electronic device. The electronic device includes: a memory and a processor, the memory and the processor are coupled; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes instructions of the method in the first aspect or any possible implementation of the first aspect.
[0038] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.
[0039] In a third aspect, the present application provides a computer-readable medium for storing a computer program, wherein the computer program includes instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0040] The third aspect and any implementation of the third aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the third aspect and any implementation of the third aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.
[0041] In a fourth aspect, the present application provides a computer program, comprising instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0042] The fourth aspect and any implementation of the fourth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.
[0043] In a fifth aspect, the present application provides a chip, the chip comprising a processing circuit and a transceiver pin, wherein the transceiver pin and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method in the first aspect or any possible implementation of the first aspect to control the receiving pin to receive a signal and control the sending pin to send a signal.
[0044] The fifth aspect and any implementation of the fifth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a schematic diagram of the hardware structure of an electronic device shown as an example;
[0046] Figure 2 A schematic diagram of the software structure of an electronic device is shown as an example;
[0047] Figure 3 A schematic diagram showing, by way of example, virtual machines running on a system on a chip in an electronic device, and services that each virtual machine is responsible for, and operations after a crash event occurs;
[0048] Figure 4A and Figure 4B This is a schematic diagram of interface changes for example when a virtual machine crashes and the entire machine is restarted;
[0049] Figure 5 A schematic diagram showing the operation of virtual machines running on a system on a chip in an electronic device and services responsible for each virtual machine after a crash event occurs after the self-recovery logic of the method for handling a virtual machine crash event provided in an embodiment of the present application is deployed;
[0050] Figure 6 A flowchart of a method for processing a virtual machine crash event provided by an embodiment of the present application is exemplified;
[0051] Figure 7 The present invention is a flowchart of another method for processing a virtual machine crash event provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0053] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0054] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects. For example, a first target object and a second target object are used to distinguish different target objects rather than to describe a specific order of target objects.
[0055] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0056] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more than two. For example, multiple processing units refer to two or more processing units; multiple systems refer to two or more systems.
[0057] In order to better understand the technical solution provided by the embodiment of the present application, before describing the technical solution of the embodiment of the present application, the hardware structure of the electronic device (such as a mobile phone, a tablet computer, a smart wearable device, etc.) to which the embodiment of the present application is applicable is first described in conjunction with the accompanying drawings. Figure 1 Let's take a mobile phone as an example.
[0058] See also Figure 1The mobile phone 100 may include: a system on chip (SOC) 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0059] Understandably, SOC 110 can also be called a system-on-chip, which is a miniature system. Specifically in practical applications, SOC 110 may include one or more processing units. For example, an application processor (AP), a modem processor (Modem), a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc., which are not listed here one by one, and this application does not limit this.
[0060] Regarding the SOC 110 including the above-mentioned processing units, in some implementations, different processing units may be independent devices. That is, each processing unit may be regarded as a processor. In other implementations, different processing units may also be integrated into one or more processors.
[0061] It should be understood that the above description is merely an example listed for a better understanding of the technical solution of this embodiment, and is not intended to be the sole limitation to this embodiment.
[0062] In addition, the SOC 110 may further include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc., which are not listed here one by one, and the present application does not limit this.
[0063] In addition, a memory may be provided in the SOC 110 for storing instructions and data. In some implementations, the memory in the SOC 110 is a cache memory. The memory may store instructions and data that the SOC 110 has just used or is cyclically using. In this way, if the SOC 110 needs to use these instructions or data again, it can be directly called from the memory. Thus, repeated access from a storage medium, such as the internal memory 121 or an external memory card, is avoided, the waiting time of the SOC 110 is reduced, and the efficiency of the system is improved.
[0064] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the SOC 110 through the external memory interface 120 to implement a data storage function.
[0065] Specifically, in the technical solution provided in the embodiment of the present application, when a crash occurs in a virtual machine at the hardware layer of an electronic device, the dump file collected and dumped by the virtual machine management service can be stored in an external memory card connected to the external memory interface 120.
[0066] The internal memory 121 can be used to store computer executable program codes, which include instructions. The SOC 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system (such as an Android system), and at least one application required for a function (such as programs corresponding to fingerprint recognition, gesture recognition, face recognition, etc.). The data storage area can store data created during the use of the mobile phone 100, etc.
[0067] Specifically, in the technical solution provided in the embodiment of the present application, when a crash occurs in a virtual machine at the hardware layer of an electronic device, the dump file collected and dumped by the virtual machine management service can be stored in the storage data area of the internal memory 121.
[0068] In addition, the internal memory 121 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0069] The charging management module 140 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging implementations, the charging management module 140 may receive charging input from a wired charger through the USB interface 130. In some wireless charging implementations, the charging management module 140 may receive wireless charging input through a wireless charging coil of the mobile phone 100. While the charging management module 140 is charging the battery 142, it may also power the electronic device through the power management module 141.
[0070] The power management module 141 is used to connect the battery 142, the charging management module 140 and the SOC 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the SOC 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some other implementations, the power management module 141 can also be set in the SOC 110. In some other implementations, the power management module 141 and the charging management module 140 can also be set in the same device.
[0071] The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0072] It should be noted that antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other implementations, the antenna can be used in combination with a tuning switch.
[0073] The mobile communication module 150 can provide wireless communication solutions including 2G / 3G / 4G / 5G etc. applied on the mobile phone 100 .
[0074] The wireless communication module 160 can provide wireless communication solutions for application in the mobile phone 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.
[0075] The audio module 170 may include a speaker 170A, a receiver 170B, a microphone 170C, an earphone jack 170D, and the like.
[0076] The sensor module 180 may include a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc., which are not listed here one by one and the present application does not impose any limitation on this.
[0077] The key 190 includes a power key, a volume key, etc. The key 190 can be a mechanical key or a touch key. The mobile phone 100 can receive key input and generate key signal input related to the user settings and function control of the mobile phone 100.
[0078] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback.
[0079] Indicator 192 may be an indicator light, which may be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0080] The camera 193 is used to capture still images or videos. In some implementations, the mobile phone 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0081] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some implementations, the mobile phone 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0082] The hardware structure of the mobile phone 100 is introduced here. It should be understood that: Figure 1 The illustrated mobile phone 100 is merely an example. In a specific implementation, the mobile phone 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have a different component configuration. Figure 1 The various components shown in the drawings may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0083] With the increasing application of virtualization technology, a large number of virtual machines are running in the SOCs used in current electronic devices, such as the virtual machine that carries the operating system of the electronic device (referred to as the first virtual machine for ease of distinction) and the virtual machine oriented to the hardware layer (referred to as the second virtual machine for ease of distinction).
[0084] In order to better understand the software structure of electronic devices with this structure, we still take mobile phones as an example. Figure 2 Provide explanation.
[0085] Before describing the software structure of the mobile phone 100 , the architecture that can be adopted by the software system of the mobile phone 100 is first described.
[0086] Specifically, in actual applications, the software system of the mobile phone 100 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture.
[0087] In addition, it is understandable that the software systems used by current mainstream electronic devices include but are not limited to Windows systems, Android systems, and iOS systems. For ease of explanation, the present application embodiment takes the layered Android system as an example to exemplify the software structure of the mobile phone 100.
[0088] In addition, the subsequent method for handling virtual machine crash events provided in the embodiments of the present application is also applicable to other systems in specific implementations.
[0089] See also Figure 2 , which is a software structure block diagram of the mobile phone 100 according to an embodiment of the present application.
[0090] like Figure 2 As shown, the layered architecture of the mobile phone 100 divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some implementations, the Android system is divided into five layers, from top to bottom, namely, the application layer (Application Layer, APP layer, also known as: application layer), the application framework layer (Application Framework Layer, FWK layer, also known as: framework layer), Android runtime (AndroidRuntime) and system library, hardware abstraction layer (HAL layer) and kernel layer (LinuxKernel Layer, Linux kernel layer, also known as: Linux kernel).
[0091] The application layer can include a series of application packages. Figure 2 As shown, the application package may include camera, game, video, live broadcast and other applications, which are not listed here one by one and are not limited in this application.
[0092] The framework layer provides application programming interfaces (APIs) and programming frameworks for the application layer's applications. In some implementations, these programming interfaces and programming frameworks can be described as functions (or services or functional modules). Figure 2 As shown, the framework layer may include functions such as content provider, window manager, resource manager, phone manager, etc., which are not listed here one by one and are not limited in this application.
[0093] Among them, Android Runtime includes core libraries and virtual machines. Android Runtime is responsible for the scheduling and management of the Android system.
[0094] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.
[0095] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.
[0096] in, Figure 2 The virtual machines shown may be a first virtual machine and a second virtual machine running in a SOC.
[0097] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0098] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.
[0099] The media library supports the playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0100] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0101] It can be understood that the 2D graphics engine mentioned above is a drawing engine for 2D drawing.
[0102] Among them, the hardware abstraction layer is used to encapsulate the hardware driver and then provide a unified interface to the upper-level framework.
[0103] Among them, the Linux kernel layer is the layer between hardware and software, including various processes, threads, power management, and various hardware drivers, such as sensor drivers, display drivers, microphone drivers, WLAN drivers, camera drivers, etc.
[0104] This concludes the introduction to the software structure of the mobile phone 100. It can be understood that: Figure 2 The layers in the software structure shown and the components included in each layer do not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer layers than shown, and each layer may include more or fewer components, which is not limited in the present application.
[0105] Each virtual machine running in the SOC is configured with its own set of virtual hardware resources. Therefore, the functions / logics / businesses that each virtual machine is responsible for are separated from each other.
[0106] That is to say, by running virtual machines (a first virtual machine and a second virtual machine) in the SOC, it is possible to isolate business scenarios and facilitate decoupling.
[0107] It should be noted that, usually, the first virtual machine that carries the operating system of the electronic device, such as most of the business logic of the Android system, is one. In a possible implementation, the first virtual machine is, for example, a primary virtual machine (PVM), such as Figure 3 shown.
[0108] Exemplarily, PVM mainly runs the Android system, and therefore can be responsible for most of the logic of the Android system. Exemplarily, these logics include the logic of the underlying Linux kernel to the upper application layer.
[0109] It should be noted that for most of the logic of the Android system that PVM is responsible for, the abnormal problems of a certain function can be solved by restarting the program of that function separately. For example, when the Bluetooth function is abnormal, you can turn off the Bluetooth function and then turn it back on to restore the normal use of the Bluetooth function. In other words, the problem can be solved without restarting the entire electronic device. Only when a panic problem occurs, it is necessary to restart the entire device.
[0110] Understandably, panic is a way of handling errors. It is used to indicate that an error that cannot be handled has occurred, causing the program to be unable to continue. For example, when we try to access a null pointer, the program will throw a panic error. The occurrence of a panic error will force the program to stop running. If it is not handled, the program may crash and affect the normal business process.
[0111] For the second virtual machine of the hardware layer of the electronic device, it can be divided into multiple virtual machines according to the functions to be implemented and the services to be responsible. For the sake of convenience, take the example of three second virtual machines included in the SOC. These three second virtual machines can be described as hardware layer virtual machine A, hardware layer virtual machine B and hardware layer virtual machine C.
[0112] See also Figure 3 For example, hardware layer virtual machine A may be a virtual machine responsible for fingerprint, face, gesture and other recognition services. Hardware layer virtual machine B may be responsible for electronic IDentity (eID), mobile POS and other services. Hardware layer virtual machine C may be responsible for CPU system reliability services.
[0113] Therefore, by running PVM, hardware layer virtual machine A, hardware layer virtual machine B and hardware layer virtual machine C in SOC, not only can the business scenarios be isolated, but also decoupling can be facilitated.
[0114] However, due to problems such as misoperation, insufficient resources, and malicious access, the virtual machine may crash, which in turn may cause the entire electronic device to malfunction and panic.
[0115] Currently, for chip platforms that provide SOC, the processing of virtual machine crash events is usually to force the entire machine to restart so that the electronic device can resume normal use. Figure 3 As shown, when the crash event of the PVM is a panic exception / problem, the mobile phone will restart the entire machine; when any hardware layer virtual machine among hardware layer virtual machine A, hardware layer virtual machine B and hardware layer virtual machine C crashes, the mobile phone will restart the entire machine.
[0116] See also Figure 4A In (1), for example, in one possible implementation, when a user clicks on the icon of the gallery application displayed in the interface 10a, the mobile phone responds to the operation and can jump to Figure 4A Interface 10b shown in (2).
[0117] For example, when the user browses the pictures displayed in the interface 10b, the mobile phone may crash the first virtual machine or the second virtual machine due to insufficient resources or user's misoperation. Based on the existing virtual machine crash event processing strategy, the mobile phone will directly restart the entire machine.
[0118] It should be noted that the whole machine restart mentioned in each embodiment of the present application is specifically a restart at the operating system level of the electronic device, that is, a restart method in which the operating system is directly restarted or the system kernel is reloaded without re-powering on. Therefore, after a virtual machine crashes, the mobile phone directly restarts the whole machine. During the process of restarting the operating system or the system kernel, the interface of the mobile phone will be changed from Figure 4A The interface 10b shown in (2) becomes Figure 4B In one possible implementation, the interface 10c shown in (1) is a black screen. In addition, as time goes by, the corresponding restart screen may also be displayed in the process. Figure 4B The interface 10c shown in (1) becomes Figure 4B After the mobile phone operating system is successfully started, the mobile phone desktop will be displayed, for example Figure 4A Interface 10a shown in (1).
[0119] That is to say, in the process of restoring the electronic device to normal use by forcibly restarting the entire device, the interface of the electronic device will change from Figure 4A The interface 10b shown in (2) to Figure 4B (1) shown in the interface 10c, and then Figure 4B The interface 10d shown in (2) finally reaches Figure 4A The changes of the interface 10a shown in (1).
[0120] It should be understood that the above description is merely an example listed for a better understanding of the technical solution of this embodiment, and is not intended to be the sole limitation to this embodiment.
[0121] Although this method of forcing the entire machine to restart can help a crashed virtual machine recover, it will cause the phone to display a black screen, re-display the startup screen, and then enter the operating system and display the desktop. Therefore, not only is the processing speed slow, but it also affects the user experience.
[0122] In view of this, an embodiment of the present application provides a method for processing a virtual machine crash event, aiming to deploy self-recovery logic for a second virtual machine, such as Figure 5 As shown. When the second virtual machine crashes, the first virtual machine that has not crashed and is hosting the operating system assists the second virtual machine that has crashed to recover to normal based on the self-recovery logic. In this way, without affecting most of the services of the electronic device, only the virtual machine responsible for part of the services is restarted, and the entire machine is not triggered to restart, thereby reducing user perception and improving user experience.
[0123] The self-recovery logic mentioned above can be divided into electronic devices for user use (hereinafter referred to as: commercial machines) and electronic devices used in the research and development stage (hereinafter referred to as: internal test machines).
[0124] Specifically, considering user privacy issues, the self-recovery logic of commercial machines may not involve the collection and storage of log information when a crash occurs; while internal test machines need to be designed to collect and store log information so that R&D personnel can locate and solve problems based on the data obtained from the memory when a crash occurs.
[0125] The following describes the process of a method for handling a virtual machine crash event based on the self-recovery logic deployed in these two types of electronic devices in conjunction with the accompanying drawings.
[0126] See also Figure 6 , exemplarily illustrates a method for processing a virtual machine crash event applied to a commercial machine, specifically comprising:
[0127] 101. The virtual machine management service notifies the hardware layer virtual machine where the crash occurs to process current resources.
[0128] Specifically, the virtual machine management service in this embodiment is Hyper-V, which can be used to create a virtual machine, shut down a virtual machine, start a virtual machine, and perform migration operations on a virtual machine.
[0129] That is, during the startup of the electronic device, the virtual machine management service will create each virtual machine that can run in the SOC. Specifically, in the embodiment of the present application, it is responsible for hosting the operating system, such as the first virtual machine of the Android system, and each second virtual machine facing the hardware layer.
[0130] For the sake of convenience, this embodiment still takes the first virtual machine as Figure 3 and Figure 5 The primary virtual machine (PVM) shown in FIG. 1 is physical, and the second virtual machine includes Figure 3 and Figure 5 Take the hardware layer virtual machine A, hardware layer virtual machine B and hardware layer virtual machine C shown in as an example.
[0131] For example, for a SOC of this architecture, the virtual machine service first creates a PVM, a hardware layer virtual machine A, a hardware layer virtual machine B, and a hardware layer virtual machine C, and configures monitoring for the created virtual machines. Thus, during the operation of these virtual machines, the virtual machine management service can collect system event information of the PVM, the hardware layer virtual machine A, the hardware layer virtual machine B, and the hardware layer virtual machine C by using the configured monitoring.
[0132] It should be noted that, in this embodiment, the system event information includes the type of event currently occurring, wherein the event type can identify what type of event the currently occurring system event is.
[0133] In addition, the system event information can also carry the source, that is, which virtual machine it comes from. In this way, according to the event type in the system event information, it can be determined which virtual machine has crashed.
[0134] In addition, the virtual machine management service can act as a manager of the virtual machine and be responsible for interacting with each virtual machine. Therefore, when the virtual machine management service uses the system event information collected by monitoring to determine that a hardware layer virtual machine, such as hardware layer virtual machine A, has a crash event, it can interact with the hardware layer virtual machine A that has crashed, specifically by sending instructions to the hardware layer virtual machine A, and then notifying the hardware layer virtual machine A to release the currently occupied resources, that is, executing step 101.
[0135] It should be noted that when the hardware layer virtual machine A is started, part of the IO resources and interrupt resources of the electronic device are taken over by the hardware layer virtual machine A. Therefore, when the hardware layer virtual machine A crashes, it is necessary to release these resources first.
[0136] See also Figure 6 For example, after receiving the instruction sent by the virtual machine management service to release the occupied resources, the hardware layer virtual machine A will close the IO and interrupt of the driver to be called when processing the business it is responsible for, and clean up the events and messages in the queues corresponding to these driver IO and interrupts.
[0137] pass Figure 3 and Figure 5 It can be seen that the services that the hardware layer virtual machine A is responsible for may include fingerprint, face, gesture and other recognition and processing services. For ease of description, this embodiment takes the fingerprint recognition service and gesture recognition service as examples of the services that the hardware layer virtual machine A is responsible for.
[0138] Understandably, when implementing fingerprint recognition services and gesture recognition services, it is necessary to call the touch driver and display driver of the electronic device. Therefore, in this embodiment, the hardware layer virtual machine A needs to turn off the IO and interrupt of the touch driver (execute step 1011), clean up the events and messages in the queue corresponding to the IO and interrupt of the touch driver (execute step 1012), and turn off the IO and interrupt of the display driver (execute step 1011'), clean up the events and messages in the queue corresponding to the IO and interrupt of the display driver (execute step 1012').
[0139] For example, in a possible implementation, the queues corresponding to the IO and interrupt of the touch driver may include a first input first output (FIFO) queue and a message queue (MSGQ). The queues corresponding to the IO and interrupt of the display driver may also include a FIFO queue and a MSGQ.
[0140] Each FIFO queue can cache events that are not processed by the hardware layer virtual machine A, and each MSGQ can cache messages that are not processed by the hardware layer virtual machine A. Therefore, the processing that the hardware layer virtual machine A needs to perform can include clearing events in each FIFO queue and clearing messages in each MSGQ.
[0141] It should be understood that the above description is merely an example listed for a better understanding of the technical solution of this embodiment, and is not intended to be the sole limitation to this embodiment.
[0142] Continue to see Figure 6 For example, after the hardware layer virtual machine A that crashes completes the release of occupied resources, the virtual machine management service will reclaim the memory resources occupied by the hardware layer virtual machine A and hand over the memory resources to the PVM.
[0143] 102. The virtual machine management service notifies the primary virtual machine to take over the resources of the hardware layer virtual machine where the crash event occurs.
[0144] Since the virtual machine management service can act as a virtual machine manager and is responsible for interacting with each virtual machine, when a crash occurs in the hardware layer virtual machine A but no crash occurs in the PVM (specifically, no panic exception / problem occurs), the virtual machine management service can send instructions to the PVM to notify the PVM to take over the IO resources, interrupt resources, memory resources, etc. released by the hardware layer virtual machine A.
[0145] In this way, the PVM can temporarily store the corresponding IO resources and interrupt resources of the hardware layer virtual machine A, and the recovered memory resources can also be used by the PVM to avoid the PVM crashing due to insufficient resources.
[0146] 103, the primary virtual machine takes over the IO resources, interrupt resources, and recovered memory resources of the hardware layer virtual machine where the crash event occurs.
[0147] 104. After the primary virtual machine takes over the resources of the hardware virtual machine where the crash occurs, the virtual machine management service coordinates the memory resources of the hardware layer virtual machine, re-establishes the mapping, and calls the kernel driver to reload the image of the hardware layer virtual machine where the crash occurs, and restarts the hardware layer virtual machine where the crash occurs.
[0148] In this way, it is possible to self-recover only the hardware layer virtual machine A that has crashed, that is, to restart only the business that the hardware layer virtual machine A is responsible for without affecting other normal businesses.
[0149] 105. After the hardware layer virtual machine where the crash occurred is successfully restarted, the main virtual machine releases the IO resources and interrupt resources of the touch driver and the display driver.
[0150] That is, the control of the IO resources and interrupt resources of the touch driver and display driver is transferred back to the hardware layer virtual machine A.
[0151] 106. After the restarted hardware layer virtual machine obtains the IO resources and interrupt resources of the touch driver and the display driver, it re-establishes the service link and resumes the service.
[0152] In this way, virtual machine A at the hardware layer can resume the business it is responsible for.
[0153] Therefore, based on the self-recovery logic, when a crash occurs in a hardware-layer virtual machine, the primary virtual machine that has not crashed and carries the operating system temporarily takes over the resources occupied by the hardware-layer virtual machine, and restarts the hardware-layer virtual machine separately. After the hardware-layer virtual machine restarts, the resources of the hardware-layer virtual machine taken over by the primary virtual machine are returned to the hardware-layer virtual machine. In this way, only the hardware-layer virtual machine that crashes is self-recovered, and the electronic device is not triggered to restart the entire machine, thereby reducing user perception and improving user experience.
[0154] In addition, when a hardware-layer virtual machine crashes, data in the memory of the electronic device is not collected for dumping during the self-recovery process, thus avoiding the leakage of user data and protecting user privacy.
[0155] In addition, the method for handling virtual machine crash events provided in the embodiments of the present application can be used as a platform solution, that is, the virtual machines in the SOC provided by the chip platform with such problems can be gradually iterated and inherited, and backward compatible with the new platform.
[0156] See also Figure 7 , exemplarily illustrates a method for processing a virtual machine crash event applied to an internal test machine, specifically comprising:
[0157] 201. The virtual machine management service collects dump information and stores it in a specified partition.
[0158] It should be noted that in the computer field, dump is generally translated as dump. Because when the program is running in an electronic device, the data on the memory, CPU, IO, etc. are dynamic (or volatile), that is to say, the data will be lost after it is used up or an exception occurs. Therefore, in the research and development stage, if you want to obtain the data at the moment when the crash occurs, that is, the dump information mentioned in the embodiment of the present application, you need to dump these data into a static form (such as a file, which will be referred to as a dump file later) and store it in a non-volatile storage medium, such as an internal memory, an external memory / card, that is, the disk is dropped to a specified partition as mentioned in the embodiment of the present application. In this way, the R&D personnel can locate and solve the problem according to the content recorded in the dump file.
[0159] 202 , the virtual machine management service notifies the hardware layer virtual machine where the crash event occurs to process current resources.
[0160] Still taking the example that when the hardware layer virtual machine in the crash event is responsible for processing the business, the drivers to be called include touch drivers and display drivers. On this basis, step 202 in the embodiment of the present application is the same as Figure 6 Step 101 in the illustrated embodiment is substantially the same. That is, the hardware virtual machine in which the crash event occurs turns off the IO and interrupt of the touch driver (executes step 2021), cleans up the events and messages in the queue corresponding to the IO and interrupt of the touch driver (executes step 2022), turns off the IO and interrupt of the display driver (executes step 2021'), cleans up the events and messages in the queue corresponding to the IO and interrupt of the display driver (executes step 2022').
[0161] For details on the implementation of the above steps, see Figure 6 The description of step 101 in the illustrated embodiment, and steps 1011, 1011', 1012, and 1012' included in step 101 are not repeated here.
[0162] Continue to see Figure 7For example, after the hardware-layer virtual machine that crashes completes the release of occupied resources, the virtual machine management service will reclaim the memory resources occupied by the hardware-layer virtual machine and hand over the memory resources to the primary virtual machine.
[0163] 203 , the virtual machine management service notifies the primary virtual machine to take over the resources of the hardware layer virtual machine where the crash event occurs.
[0164] 204 , the primary virtual machine takes over the IO resources, interrupt resources, and reclaimed memory resources of the hardware layer virtual machine where the crash event occurs.
[0165] 205, after the master virtual machine takes over the resources of the hardware virtual machine where the crash event occurs, the virtual machine management service coordinates the memory resources of the hardware layer virtual machine, re-establishes the mapping, and calls the kernel driver to reload the image of the hardware layer virtual machine where the crash event occurs, and restarts the hardware layer virtual machine where the crash event occurs.
[0166] 206. After the hardware layer virtual machine where the crash occurred is successfully restarted, the main virtual machine releases the IO resources and interrupt resources of the touch driver and the display driver.
[0167] 207, after the restarted hardware layer virtual machine obtains the IO resources and interrupt resources of the touch driver and the display driver, it re-establishes the service link and resumes the service.
[0168] It should be noted that steps 203 to 207 in the embodiment of the present application are similar to Figure 6 Steps 102 to 106 in the embodiment shown are substantially the same. For specific implementation details, see Figure 6 The description of steps 102 to 106 in the illustrated embodiment will not be repeated here.
[0169] Therefore, based on the self-recovery logic, when a crash occurs in a hardware-layer virtual machine, the primary virtual machine that has not crashed and carries the operating system temporarily takes over the resources occupied by the hardware-layer virtual machine, and restarts the hardware-layer virtual machine separately. After the hardware-layer virtual machine restarts, the resources of the hardware-layer virtual machine taken over by the primary virtual machine are returned to the hardware-layer virtual machine. In this way, only the hardware-layer virtual machine that crashes is self-recovered, and the electronic device is not triggered to restart the entire machine, thereby reducing user perception and improving user experience.
[0170] In addition, when a crash occurs in the hardware layer virtual machine, triggering the hardware layer virtual machine to release resources, the main virtual machine takes over the resources released by the hardware layer virtual machine and the occupied memory resources, and also dumps the current data in the memory of the electronic device, so that during the research and development stage, R&D personnel can quickly locate and solve the problem based on the dump file.
[0171] In addition, it is understandable that, in order to implement the above functions, the electronic device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of the present application.
[0172] In addition, it should be noted that the method for handling virtual machine crash events provided by the above embodiments implemented by electronic devices in actual application scenarios can also be performed by a chip system included in the electronic device, wherein the chip system may include a processor. The chip system can be coupled to a memory so that the chip system calls a computer program stored in the memory when it is running to implement the steps performed by the above electronic device. The processor in the chip system can be an application processor or a processor other than an application processor.
[0173] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the method for handling the virtual machine crash event in the above-mentioned embodiment.
[0174] In addition, an embodiment of the present application further provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the above-mentioned related steps to implement the method for processing a virtual machine crash event in the above-mentioned embodiment.
[0175] In addition, an embodiment of the present application also provides a chip (which may also be a component or module), which may include one or more processing circuits and one or more transceiver pins; wherein the transceiver pins and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the above-mentioned related method steps to implement the method for handling the virtual machine crash event in the above-mentioned embodiment, so as to control the receiving pin to receive the signal, so as to control the sending pin to send the signal.
[0176] In addition, it can be seen from the above description that the electronic device, computer-readable storage medium, computer program product or chip provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0177] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for processing a virtual machine crash event, characterized in that: Applied to an electronic device, a first virtual machine and a second virtual machine are running on the electronic device, the first virtual machine is a virtual machine carrying an operating system of the electronic device, and the second virtual machine is a virtual machine facing a hardware layer of the electronic device; The method comprises: In response to a crash event of the second virtual machine, releasing the first resource occupied by the second virtual machine, and having the first virtual machine take over the first resource; reallocating memory resources for the second virtual machine, and restarting the second virtual machine; After the second virtual machine is successfully restarted, releasing the first resource taken over by the first virtual machine; After the second virtual machine obtains the first resource, a service link between the second virtual machine and the service that the second virtual machine is responsible for is reestablished to restore the processing capability of the service.
2. The method according to claim 1, characterized in that: The releasing the first resource occupied by the second virtual machine includes: Turning off the input / output IO resources and interrupt resources of the first driver, where the first driver is the driver to be called by the second virtual machine when processing the service; Clear the contents in the queues corresponding to the IO resources and the interrupt resources.
3. The method according to claim 2, characterized in that The clearing of the contents in the queues corresponding to the IO resources and the interrupt resources includes: clearing the events in the first-in-first-out queues corresponding to the IO resources and the interrupt resources, as well as, Messages in the message queues corresponding to the IO resources and the interrupt resources are cleared.
4. The method according to claim 2, characterized in that: After clearing the contents in the queues corresponding to the IO resources and the interrupt resources, the method further includes: The memory resources occupied by the second virtual machine are reclaimed.
5. The method according to claim 4, characterized in that The method further comprises: The first virtual machine takes over the reclaimed memory resources.
6. The method according to claim 1, characterized in that The reallocating memory resources for the second virtual machine and restarting the second virtual machine includes: reallocating memory resources for the second virtual machine, and re-establishing a mapping between the second virtual machine and the memory resources; Calling a second driver to load the image of the second virtual machine; Restart the operating system running on the second virtual machine.
7. The method according to any one of claims 1 to 6, characterized in that: In response to a crash event occurring in the second virtual machine, the method further includes: Acquire data in the memory resources corresponding to the second virtual machine, and dump the data to the storage medium of the electronic device.
8. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: In the process of creating the first virtual machine and the second virtual machine, configuring monitoring for the first virtual machine and the second virtual machine; Collecting system event information of the first virtual machine and the second virtual machine by using the monitoring, wherein the system event information includes an event type; When the event type indicates that the system event currently occurring in the second virtual machine is a crash event, the step of responding to the crash event occurring in the second virtual machine is performed.
9. The method according to claim 8, characterized in that The method further comprises: When the event type indicates that the system event currently occurring in the first virtual machine is a crash event, in response to the crash event occurring in the first virtual machine, the electronic device is completely restarted.
10. The method according to claim 9, characterized in that When the event type indicates that the system event currently occurring in the first virtual machine is a crash event, in response to the crash event occurring in the first virtual machine, restarting the entire electronic device includes: When the crash event is a panic exception, in response to the crash event of the first virtual machine, the electronic device is completely restarted.
11. An electronic device, characterized in that: The electronic device includes: a memory and a processor, the memory and the processor are coupled; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method for processing a virtual machine crash event as described in any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that: The method comprises a computer program, which, when executed on an electronic device, enables the electronic device to execute the method for processing a virtual machine crash event as claimed in any one of claims 1 to 10.
Citation Information
Cited By
Handling method for virtual-machine crash event, and device and storage medium
EP4760503A1
Handling method for virtual-machine crash event, and device and storage medium
WO2025097882A1