Hypervisor system operation structure domain failure status monitoring device
The failure status monitoring device for hypervisor systems addresses cascaded failures by sharing failure status information, ensuring operational domains meet their specifications and maintain functional safety.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PERSEUS CO LTD
- Filing Date
- 2022-12-28
- Publication Date
- 2026-05-01
AI Technical Summary
Hypervisors fail to prevent cascaded failures in operational domains due to physical or logical failures, leading to non-compliance with system requirements, especially affecting ASIL-rated domains.
A failure status monitoring device for hypervisor systems that monitors and shares failure status information among operational domains, using a hypervisor function to manage dependent failures through a failure status monitoring manager.
Prevents dependent failures by ensuring operational domains operate according to their specifications, particularly maintaining functional safety for ASIL-rated systems.
Smart Images

Figure 0007854524000001 
Figure 0007854524000002 
Figure 0007854524000003
Abstract
Description
Technical Field
[0004]
[0001] The present invention relates to a failure state monitor device for an operation system domain of a hypervisor system. More specifically, in the present invention, a physical failure or a logical failure situation occurs in some operation system domains operating on a hypervisor, and a cascaded failure phenomenon occurs in the remaining operation system domains that were operating in cooperation with the corresponding operation system domain on the hypervisor, resulting in the inability to operate while satisfying the system requirements. The present invention relates to a failure state monitor device for an operation system domain of a hypervisor system that can prevent such problems.
Background Art
[0002] Generally, hypervisor software is software that makes one computer hardware into multiple virtual computer hardwares. To design / develop hypervisor software, a high technical level such as that for creating general-purpose operation system software like Windows (registered trademark) and Linux (registered trademark) is required.
[0003] Conventionally, hypervisors have been used to operate computing so that hardware efficiency in cloud data centers and banking services are not interrupted during operation system upgrades.
[0004] <00000Such hypervisors efficiently utilize computer hardware to operate multiple operational systems simultaneously, allocating hardware resources according to the given specifications for each operational system.
[0005] On the other hand, among hardware resources, input / output devices such as networks, touchscreens, and mice, which are commonly used by numerous operating systems on the hypervisor, can malfunction by consuming more resources than the hardware resources allocated according to the specifications. For example, in situations like this where hardware resources are used excessively, the hardware resources allocated on the hypervisor according to the original specifications for each operational structure may become insufficient, which can cause malfunctions in the operational structure.
[0006] Since devices such as automobiles, drones, and robots that will be applied in the future will use a real-time operation system, if a shortage of hardware resources occurs as described above, the real-time operation system will be unable to provide normal service, and the systems of automobiles, drones, and robots will malfunction.
[0007] These problems can be explained more specifically and exemplified with reference to Figures 1 and 2 as follows.
[0008] Figure 1 is a diagram illustrating the structure in which control / application software (e.g., AutoSAR) operates on a general-purpose operating system (e.g., Linux®) using conventional technology, while Figure 2 is an illustrative diagram showing a structure in which the control / application software (e.g., AutoSAR) from Figure 1 is separated and configured on a hypervisor using the ISO 26262 functional safety class ASIL (Automotive Safety Integrity Level) decomposition methodology. Referring to Figure 1, conventionally, due to the high complexity of general-purpose operating systems (e.g., Linux), it was not possible to define the system in Figure 1 under higher ASIL grades (e.g., ASIL-B, ASIL-D, etc.). Referring to Figure 2, the ASIL decomposition methodology is applied to the control / application software (e.g., AutoSAR) in Figure 1, and it is structurally distributed and implemented into AutoSAR (QM: Quality Management) which does not require an ASIL grade and AutoSAR (ASIL) which does require an ASIL grade. In this case, AutoSAR (ASIL) is run on an operating system to which a higher ASIL grade is applied (e.g., Safety OS), and AutoSAR (QM) is run on a general-purpose operating system.
[0009] In the structure shown in Figure 2, conventionally, the hypervisor lacked a device to prevent dependent failures between AutoSAR(QM) and AutoSAR(ASIL), which presented problems in applying the ASIL decomposition methodology. [Prior art documents] [Patent Documents]
[0010] [Patent Document 1] Korean Published Patent Gazette No. 10-2021-0127427 (Publication Date: October 22, 2021, Title: CPU Virtualization Method and Apparatus in a Multicore Embedded System) [Patent Document 2] Korean Patent Publication No. 10-2021-0154769 (Publication Date: December 21, 2021, Title: Microkernel-based extensible hypervisor) [Patent Document 3] Korean Published Patent Gazette No. 10-2019-0029977 (Publication Date: March 21, 2019, Title: Control System for Equipment and Method for Driving the Same) [Patent Document 4] Korean Published Patent Publication No. 10-2015-0090439 (Publication Date: August 6, 2015, Title: Method for scheduling on a hypervisor of a many-core system) [Overview of the project] [Problems that the invention aims to solve]
[0011] The technical problem of the present invention is to prevent a problem that occurs when implementing ASIL decomposition running on a hypervisor, where a physical or logical failure occurs in some operational domains, causing a cascaded failure phenomenon in the remaining operational domains that were operating in conjunction with the affected operational domain on the hypervisor, making it impossible to operate while satisfying the system requirements.
[0012] Furthermore, the technical problem of the present invention is to solve the issue in which a situation occurs in which some operational domains operating on the hypervisor experience failures, resulting in dependent failures where the operational domains cannot operate according to their original specifications on the hypervisor, and the remaining operational domains, particularly the operational domains that have received a higher ASIL rating and the control / application software that has received a higher ASIL rating, fail to satisfy the functional safety requirements and become unable to operate.
[0013] Various communication channels (e.g., back-end device drivers, front-end device drivers, logical bus channels) exist between operational domains on the hypervisor, enabling mutually dependent operation between operational domains. A more specific technical objective of the present invention is to fundamentally prevent the problem that occurs when a dependent failure occurs between such operational domains, causing the entire related operational domain to be affected by the failure. [Means for solving the problem]
[0014] To solve these technical problems, the failure status monitoring device for the operational system domain of a hypervisor system according to an embodiment of the present invention includes a hardware physical layer, an operational system layer consisting of multiple operational systems of different types and control / application software, a hypervisor that allocates basic resources for using the system resources of the hardware physical layer so that the multiple operational systems operate in a virtual machine environment, has a function to notify the status of dependent failures for each of the multiple operational system domains corresponding to the multiple operational systems constituting the operational system layer, has a function to monitor the failure status of each of the multiple operational system domains, and includes a failure status monitoring manager that has a failure status monitoring function, the hypervisor collects the failure status for each of the multiple operational system domains via an interface and shares it with the multiple operational system domains.
[0015] The failure status monitoring device for the operational structure domain of a hypervisor system according to the present invention is characterized in that the system resources include one or more of the following: CPU (central processing unit) resources, MCU (microcontroller unit) resources, and memory resources.
[0016] The failure status monitoring device for an operational domain of a hypervisor system according to the present invention is characterized in that the failure status monitoring manager constituting the hypervisor continuously tracks the failure status of an operational domain through an interface from which the hypervisor collects status information of a specific operational domain, shares the operational / malfunction status of the linked operational domains through the hypervisor interface so that it can be understood, makes an event signal public to the entire system when a malfunction occurs in the operational domain of the hypervisor system, and transmits this publicly known event signal to the user.
[0017] The failure status monitoring device for the operational structure domain of a hypervisor system according to the present invention is characterized in that the operational structure hierarchy includes one or more operational structures having an ASIL rating, one or more control / application software having an ASIL rating, and one or more general-purpose operational structures. In the failure status monitoring device for the operational system domain of a hypervisor system according to the present invention, the failure status monitoring manager is characterized in that it registers operational system domains that operate in cooperation with each other via an identification ID assigned to the operational system domain on the hypervisor, and manages the coordinated operation. [Effects of the Invention]
[0018] According to the present invention, a situation occurs in which some operational domains operating on the hypervisor experience failures, resulting in dependent failures where each operational domain on the hypervisor cannot operate according to its original specifications. This solves the problem where the remaining operational domains, especially those with a higher ASIL rating, are unable to operate while satisfying functional safety requirements. Furthermore, the failure status information shared by the hypervisor with the operational domain allows for the identification of the failure status of individual operational systems within the ASIL decomposition software structure, which fundamentally prevents the problem of dependent failures in operational domains operating on the hypervisor. Furthermore, this solution can resolve a problem where a situation occurs where some operational domains running on the hypervisor fail, resulting in dependent failures where the operational domains cannot operate according to their original specifications on the hypervisor. This can cause the remaining operational domains, particularly the operational domains and control / application software that have received a higher ASIL rating, to fail to meet functional safety requirements and become unable to operate. [Brief explanation of the drawing]
[0019] [Figure 1]A structural diagram in which control / application software (e.g., AutoSAR) operates in a general-purpose operation system (e.g., Linux) according to the prior art. [Figure 2] A diagram exemplarily showing a structure in which the control / application software (e.g., AutoSAR) in FIG. 1 is separated and configured on a hypervisor according to the ISO26262 functional safety level ASIL decomposition methodology. [Figure 3] A diagram showing a failure state monitoring device for an operation system domain of a hypervisor system according to an embodiment of the present invention. [Figure 4] A diagram exemplarily showing the configuration of a hypervisor in an embodiment of the present invention. [Figure 5] A diagram showing an exemplification of a system architecture in which a device driver domain is isolated from an operation system domain in an embodiment of the present invention. [Figure 6] A diagram showing an exemplification of a system architecture in which a device driver domain is merged into an operation system domain in an embodiment of the present invention.
Modes for Carrying Out the Invention
[0020] The specific structural or functional description of the embodiments according to the concept of the present invention disclosed in the present invention is merely exemplified for the purpose of explaining the embodiments according to the concept of the present invention. The embodiments according to the concept of the present invention are implemented in various forms and are not limited to the embodiments described in this specification. Embodiments according to the concept of the present invention can be subject to various changes and can have various forms. Therefore, the embodiments are exemplified in the drawings and described in detail herein. However, this is not intended to limit the embodiments according to the concept of the present invention to a specific disclosed form, and includes all changes, equivalents, or alternatives included in the spirit and technical scope of the present invention. Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as those generally understood by a person of ordinary skill in the art to which this invention pertains. Terms that are the same as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not as ideal or overly formal unless expressly defined herein.
[0021] The following describes the technical principles of the present invention, followed by a detailed explanation of embodiments of the present invention.
[0022] Hypervisor 30 efficiently utilizes computer hardware to operate multiple operational systems simultaneously, and it operates according to the specifications given to each operational system. If one of several operational domains malfunctions, other operational systems that operate in conjunction with that specific domain will be unable to function as specified due to dependent failures according to their specifications.
[0023] On the other hand, since devices such as automobiles, drones, and robots that will be applied in the future will use a real-time operation system, if a dependent failure occurs, the real-time operation system will be unable to provide normal service, and the systems of automobiles, drones, and robots will malfunction.
[0024] To resolve such dependent failures, the following describes how to apply ASIL decomposition to the hypervisor.
[0025] Conventional technologies lacked the functionality and interfaces for sharing status information between multiple operational systems running on the hypervisor 30, and therefore had no way to prevent dependent failures. In other words, with conventional technology, if a malfunction occurred in one of the multiple operating systems on the hypervisor, the other operating systems would remain unaware of it. This problem escalated significantly when it was linked to software requiring ASIL certification.
[0026] This invention solves these problems in the following manner. In other words, in order to cope with the proliferation of autonomous operating devices such as future automobiles, drones, and robots, future hypervisors 30 will need to support ASIL-rated operating systems, and therefore will need to operate according to predetermined specifications for each of the multiple operating systems on hypervisor 30. To this end, the present invention provides a hypervisor function and interface that allows the hypervisor 30 to receive or understand the status of multiple operating systems and to share status information such as normal operation or malfunction with the multiple operating systems, thereby providing a mechanism for managing dependent failures by sharing a failure status monitor.
[0027] Preferred embodiments of the present invention will be described in detail below with reference to the attached drawings.
[0028] Figure 3 shows a fault status monitoring device for the operational domain of a hypervisor system according to one embodiment of the present invention, and Figure 4 is a diagram illustrating the configuration of the hypervisor in one embodiment of the present invention.
[0029] Referring to Figures 3 and 4, a failure status monitoring device for the operational structure domain of a hypervisor system according to one embodiment of the present invention comprises a hardware physical layer 10, an operational structure layer 20, and a hypervisor 30.
[0030] The hardware physical layer 10 is an element that constitutes a physical device in which a failure status monitoring device for the operational domain of a hypervisor system according to an embodiment of the present invention is realized. For example, the physical device constituting the hardware physical layer 10 includes a CPU or MCU, memory including DRAM (dynamic random access memory), and input / output devices. The input / output devices include, but are not limited to, storage, network devices, information output devices including touch screens, information input devices including keyboards and mice, and serial input / output devices.
[0031] The operational structure hierarchy 20 consists of multiple operational structures of different types and control / application software. For example, the multiple operational structures constituting the operational structure hierarchy 20 may be configured to include one or more ASIL-grade operational structures and one or more general-purpose operational structures including Unix, Windows, etc., and the control / application software may be configured to include one or more ASIL-grade software.
[0032] The hypervisor 30 is a component that allocates basic resources for using the system resources of the hardware physical layer 10 for each of the multiple operational domains corresponding to the multiple operational systems that make up the operational system layer 20, and enables the multiple operational systems to operate in a virtual machine environment.
[0033] Furthermore, the hypervisor 30 has a function to notify the status of dependent failures for each of the multiple operational system domains corresponding to the multiple operational systems that constitute the operational system hierarchy 20, and has a failure status monitoring function to enable multiple operational systems to operate in a virtual machine environment.
[0034] Furthermore, the hypervisor 30 collects failure status from multiple operational domains via interfaces and shares it with the multiple operational domains. For example, the hypervisor 30 can collect failure status by monitoring via a watchdog and communicating with the operational hierarchy 20 via interfaces. Examples of interfaces include shared memory, device driver software, and hypercall.
[0035] For example, the system resources that hypervisor 30 allocates to each of the multiple operational domains may include one or more of the following: CPU resources, MCU resources, and memory resources.
[0036] Referring to the example in Figure 4, the hypervisor 30 may be configured to include a Resource Allocator 32, a Domain Manager 34, an Access Controller 36, and a Failure Status Monitoring Manager 38.
[0037] The resource allocater 32 allocates the amount of system hardware resources such as CPU and memory for each domain.
[0038] The domain manager 34 schedules the domain in a time-sharing manner based on the amount of resources allocated by the resource allocater 32, and manages context switching during scheduling.
[0039] The access controller 36 controls the access between objects such as domains, hardware system resources, and data.
[0040] The Fault Status Monitor Manager 38 monitors the status information of each domain, device driver, and control / application software, and shares this information with the entire operational structure.
[0041] For example, the failure status monitor manager 38 may be configured to register operational domains that operate in conjunction with each other using an identification ID assigned to the operational domain on the hypervisor 30, and to manage coordinated operations.
[0042] For example, if the hardware system resource is a CPU resource, one embodiment of the present invention uses the failure status monitor manager 38 constituting the hypervisor 30 to accurately grasp and manage the status of each operational system domain (e.g., whether or not there is a malfunction), and by sharing the status with the operational system, the operational system domains operating in cooperation can decide whether to continue or interrupt their coordinated operation. In the end, since the operational system domain is guaranteed to operate according to the specifications when it operates on the hypervisor 30, the service of an ASIL-grade operational system can be stably guaranteed.
[0043] Figure 5 shows an example of a system architecture in which the device driver domain is isolated from the operational structure domain, and Figure 6 shows an example of a system architecture in which the device driver domain is merged into the operational structure domain. One embodiment of the present invention can be applied in common to the system architectures illustrated in Figures 5 and 6.
[0044] As explained in detail above, the present invention has the effect of solving the problem in which a situation occurs in which some operational system domains operating on the hypervisor experience failures, resulting in dependent failures where the operational system cannot operate according to its original specifications on the hypervisor, and the remaining operational system domains, especially those with a higher ASIL rating, are unable to operate while satisfying the functional safety requirements.
[0045] Furthermore, the failure status information shared by the hypervisor with the operational domain allows for the identification of the failure status of individual operational systems within the ASIL decomposition software structure, which fundamentally prevents the problem of dependent failures in operational domains operating on the hypervisor.
[0046] Furthermore, this solution can resolve a problem where a situation occurs where some operational domains running on the hypervisor fail, resulting in dependent failures where the operational domains cannot operate according to their original specifications on the hypervisor. This can cause the remaining operational domains, particularly the operational domains and control / application software that have received a higher ASIL rating, to fail to meet functional safety requirements and become unable to operate. [Explanation of Symbols]
[0047] 10: Hardware physical tier 20: Management Structure Hierarchy 30: Hypervisor 32:Resource Allocator 34: Domain Manager 36: Access Controller 38: Failure Status Monitoring Manager
Claims
1. A failure status monitoring device for the operational structure domain of a hypervisor system, Hardware physical layer, An operational structure hierarchy consisting of multiple operational structures of different types and control / application software, The hypervisor includes a failure status monitor manager that has a failure status monitoring function, which allocates basic resources for using the system resources of the hardware physical layer so that the multiple operational systems can operate in a virtual machine environment, has a function to notify the status of dependent failures for each of the multiple operational system domains corresponding to the multiple operational systems that constitute the operational system layer, and enables the multiple operational systems to operate in a virtual machine environment, The hypervisor collects failure statuses for each of the multiple operational domains via an interface and shares them with the multiple operational domains. The failure status monitor manager, which constitutes the hypervisor, is configured to register operational domains that operate in conjunction with the operational domains on the hypervisor using an identification ID, and to manage coordinated operations. The aforementioned failure status monitor manager is a failure status monitor device for a hypervisor system's operational domain, which continuously tracks the failure status of an operational domain through an interface from which the hypervisor collects status information of a specific operational domain, shares the operational / malfunction status of linked operational domains via the hypervisor interface so that it can be understood, and makes an event signal public to the entire system when a malfunction occurs in the operational domain of the hypervisor system, and ensures that this publicly known event signal is transmitted to users.
2. The failure status monitoring device for the operational system domain of a hypervisor system according to claim 1, characterized in that the system resources include one or more of the following: CPU (central processing unit) resources, MCU (microcontroller unit) resources, and memory resources.
3. The failure status monitoring device for the operational system domain of a hypervisor system according to claim 1, characterized in that the operational system hierarchy includes one or more operational systems having ASIL grades, one or more control / application software having ASIL grades, and one or more general-purpose operational systems.
Citation Information
Patent Citations
Virtual machine migration method, information processing device and program
JP2014142720A
System for vehicle
JP2020187631A
Information processing device, control method, control program, and vehicle
JP2022099008A
Method for scheduling a task in hypervisor for many-core systems
KR1020150090439A
A control system for device and process for operationg the control system
KR1020190029977A