Coordinated multi-mode allocation and runtime switching for systems with dynamic fault tolerance requirements
The method optimizes controller resource utilization in fault-tolerant systems by reallocating functions based on dynamic redundancy needs, addressing inefficiencies in existing systems and reducing costs through dynamic load balancing.
Patent Information
- Application Number
- DE102017119447
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-25
- Filing Date
- 2017-08-24
- Publication Date
- 2025-08-14
- Estimated Expiration
- 2037-08-24
AI Technical Summary
Existing fault-tolerant control systems face inefficiencies in resource utilization due to fixed redundancy modes, leading to unnecessary use of backup controllers and increased costs.
A method for reallocating control functions based on dynamic redundancy requirements, using a look-up table to determine the optimal distribution of functions among controllers, minimizing load and optimizing resource usage.
Achieves cost-effective resource utilization by dynamically reallocating functions, reducing controller load and minimizing the need for backup controllers, thereby enhancing efficiency and reducing system costs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to fault-tolerant control systems and, more particularly, to a method for reallocating control functions based on utilization minimization. BACKGROUND OF THE INVENTION
[0002] For general background information, reference is made at this point to the documents US 2011 / 0 099 553 A1, US 2014 / 0 282 593 A1, US 2007 / 0 143 323 A1 and US 2014 / 0 160 927 A1.
[0003] Further prior art can be found in the documents US 2005 / 0 172 097 A1, EP 3 012 704 A2 and US 6,087,963 A as well as in the articles "Feedback utilization control in distributed real-time systems with end-to-end tasks" by Lu, Chenyang, Xiaorui Wang and Xenofon Koutsoukos (IEEE Transactions on Parallel and Distributed Systems 16, No. 6 (2005): pages 550-561) and "Robust fault tolerant controller in parabolic distributed parameter systems with actuator faults" by Michael A. Demetriou (42nd IEEE International Conference on Decision and Control, pages 324-329).
[0004] Systems that provide safety functions typically use redundant controllers to ensure safety by terminating functions that have experienced a fault or error. Such systems are known as fail-safe systems. If a fault is detected, the controls for the function are disabled, and the function is no longer operational within the system.
[0005] Some systems attempt to implement control systems that employ a fail-safe system, where additional controllers are used to ensure that safe operation can continue for a period of time, such as dual-duplex controllers. If a first controller fails and becomes silent, a second controller is activated, and all actuators switch to rely on requests from the second controller. Because the controllers must perform different functions and redundancies depending on the critical functions running, efficient use of a controller is desirable, as both the functions and the number of backup controllers needed to perform those functions depend on the redundancy mode for the particular function within each controller. SUMMARY OF THE INVENTION
[0006] One advantage of the invention is the reconfiguration of functions based on the executed mode requirements within controllers. For failover applications with mixed and dynamic redundancy requirements, a system architecture pattern and a switching protocol are applied. Cost efficiency is achieved by changing resource utilization at runtime depending on the redundancy requirements of the operating mode. Controller consolidation is achieved on a single subsystem, thus enabling cost-effective architectures.
[0007] According to the invention, a method for reassigning control functions based on minimizing utilization is presented, which is characterized by the features of claim 1.
[0008] A lookup table is created based on functions and operating modes. Each entry in the lookup table contains the number of executions required for a particular function in a particular operating mode. Functions are assigned to controllers based on the number of executions required for a function in a particular operating mode. Each controller is designated as having a primary state, backup state, or non-executing state for each function. A utilization rate is determined for each controller in each operating mode. A minimum utilization rate for each controller is determined across each operating mode. The utilization rates of the different operating modes are compared for each of the controllers. Matching utilization rates between controllers of different operating modes are identified.Multi-mode remapping of function execution is coordinated in the controller by switching a set of pre-mapped functions between different controllers within a respective operating mode to reduce the utilization rate of at least one controller. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is an architectural block diagram of an exemplary integrated control system. Fig. Figure 2 illustrates control configurations based on the operating mode. Fig. Figure 3 illustrates a lookup table that identifies the status mode of a controller based on the operating mode for each function. Fig. Figure 4 illustrates a flowchart for transitioning to a new status mode based on the operating mode. Fig. Figure 5 illustrates control configurations based on operating mode and load. Fig. Figure 6 illustrates controller remapping based on utilization rate coordination. Fig. Figure 7 illustrates a flowchart for a method for coordinating multi-mode allocation for runtime switching. DETAILED DESCRIPTION
[0009] The system and methodology described herein can be used to maintain safety control functions in controllers that cost-effectively execute software functions in control systems. While the approach and methodology are described below with respect to vehicle communications, those skilled in the art will understand that an automotive application is merely exemplary and that the concepts disclosed herein may also be applied to any other suitable communication system, such as general industrial automation applications, manufacturing and assembly applications, avionics, aerospace, and gaming.
[0010] The term “vehicle” as described herein may be broadly interpreted to include not only a passenger car, but all other vehicles, including rail transit systems, aircraft, off-road sports vehicles, robotic vehicles, motorcycles, trucks, sport utility vehicles (SUVs), recreational vehicles (RVs), watercraft, aircraft, agricultural vehicles, self-driving vehicles, shared vehicles, and construction vehicles.
[0011] In Fig. Figure 1 shows an architectural block diagram of an exemplary integrated control system. Such control systems often use two controllers so that if a hardware failure occurs in a primary controller, a backup controller can be easily activated to control a function of the control system or to provide a controller for limited functionality of the feature in the event of a failure.
[0012] In Fig. Figure 1 shows an architectural block diagram of an integrated failover control system. Control systems include vehicles, aircraft, and ships that utilize safety-critical or autonomous systems that require fault-tolerant countermeasures in the event of a fault within the control system. Such control systems often employ two or more controllers so that if a fault (resulting from a defect) occurs in a primary controller, a backup controller can be easily activated to control a function of the control system or to provide a controller for limited functionality of the feature in the event of a fault.
[0013] In Fig. 1 illustrates a respective system including a first controller 12 (e.g., an electronic control unit), a second controller 14, and a third controller 16. The exemplary system described herein is vehicle-based, but the architecture may be applicable to non-vehicle systems, as previously described. Each of the controllers includes at least one central processing unit (CPU) for executing software. Depending on the operating mode of the system, each controller may be enabled to perform functions through its CPU.
[0014] The first controller 12, the second controller 14, and the third controller 16 communicate via a communications network 18. It is understood that the communications network may include a communications area network (CAN), CAN with flexible data rate (CAN-FD), FlexRay, switched networking with Ethernet, wireless communication, or multiple networks with gates. The requirement is that each of the controllers and sensors / actuators can communicate with each other. The first controller 12, the second controller 14, and the third controller 16 use the communications network 18 to receive and transmit data between the sensors 20 and the actuators 22.
[0015] Fig. Figure 2 illustrates the controller status and function execution by controllers in different operating modes. For illustration purposes herein, the first controller 12, the second controller 14, and the third controller 16 are identical, with the same hardware and software. A redundancy table 24 is shown, illustrating the required redundancies for each of the functions given the respective operating mode and the respective functions to be performed. For example, the redundancy table 24 indicates the number of required copies of each function. That is, the redundancy table 24 only indicates the number of controllers required (i.e., whether primary, backup, or no controllers are required) for each function in each respective mode.
[0016] Functions may include lane detection, pedestrian detection, vehicle detection, and path planning. The redundancy table 24 identifies whether a primary controller is required for that function and the number of backup controllers required. In the exemplary redundancy table 24, each function F1, F2, F3, F4 is listed in the rows and the respective operating modes in the column. The operating modes may include, for example, urban daytime driving, urban nighttime driving, highway daytime driving, highway nighttime driving. The redundancy table 24 identifies for each function and each operating mode whether the primary controller is used and the number of backup controllers required. For example, three independent embodiments are required for operating mode (M1) and function (F1).This indicates that one primary controller and two backup controllers are required when executing this corresponding function under this operating mode. Another example in the table is that when the operating mode is (M2) and the function is (F3), independent execution is required. This indicates that one primary controller and no backup controller is required. In yet another example, when the operating mode is (M1) and the function is (F4), no independent executions are required. This indicates that neither a primary controller nor any backup controllers are needed.
[0017] In order to enable each controller for an operating mode in its respective state mode, each controller has a lookup table 25 stored in its memory as shown in Fig. 3. Using lookup table 25, the status of the respective controllers is shown as they are configured when functions are executing for a respective operating mode. Each of the controllers can be designated as either a primary status mode (P), a hot standby status mode (HS), a cold standby status mode (CS), or a non-execution status mode (NE).A primary status mode (P) indicates that a corresponding controller is designated as the primary controller for performing this function; a hot standby status mode (HS) indicates that a corresponding controller is designated as the first backup controller for this function; a cold standby status mode (CS) indicates that a corresponding controller is not active as a backup controller but is ready to enter hot status mode (or directly to the primary if desired by the system designer); and a non-executing status mode (NE) indicates that the controller is not being used in any way in this mode for this function.Switching and reconfiguration between primary status mode, hot standby status mode, and cold standby status mode for controller failures is described in full in the co-pending US application US 2017 / 0 277 607 A1, titled "Fault Tolerance Pattern and Switching Protocol for Multiple Hot and Cold Standby Redundancies." Switching between primary status (P), hot status (HS), cold status (CS), and non-executing (NE) occurs as a result of the mode change. As a result, each respective controller does not know the redundancies required for a respective function for a respective mode; rather, each respective controller only selects its status mode for each respective function under each respective mode.However, it is understood that the controllers can communicate with each other to identify faults and notify other controllers of their status and whether a status mode change is required. The lookup table for each controller is created at design time. That is, when determining the allocation of each function and its replica to the controllers, and for each function and its replica, the execution mode (primary, hot, cold, NE) is determined. Each controller then has a table similar to lookup table 25 stored in its memory to look up the execution mode of each function / replica based on the operating mode (e.g., M1 or M2).
[0018] Fig. Figure 4 is a flowchart for transitioning to a new status mode based on the operating mode. In step 40, a new operating mode is identified. In step 41, the current operating mode is identified.
[0019] In step 42 the current mode is set.
[0020] In step 43, the routine iterates over the set of all functions on a controller. That is, each function is indexed to determine whether the status mode needs to be changed based on the new operating mode. When all functions have been tested between the different modes, the routine ends. If additional functions need to be tested, the routine continues to step 44.
[0021] In step 44, a next function for testing is identified.
[0022] In step 45, a test is performed to determine if the ECU's status mode for executing the identified function from the current operating mode for the new mode is equal to the ECU's status mode for executing the identified functions for the current operating mode for the current mode. If the status modes are equal, a return is made to step 43 to iterate to the next function. If the status modes are not equal, the routine continues to step 46. The lookup table is used to identify if the status modes are equal, comparing the same function between the two different operating modes in the lookup table.
[0023] In step 46, the function's status mode is set to the status mode of the new function, as specified in the lookup table. The routine returns to step 43 to iterate to the next function.
[0024] As in Fig. 2, the function (F1) when operated in an operating mode (M1) requires three controllers (i.e. 1 primary and 2 backups). As shown in Fig. As illustrated in Figure 2, controller 12 is designated as a primary controller (P) for performing function (F1), controller 14 is designated as a backup controller operating in (HS) mode, and controller 16 is designated as a backup controller operating in (CS) mode. This configuration satisfies the redundancy table requirements for (F1) while operating in (M1) mode.
[0025] For function (F4) when operating in system mode M2, where only two controllers are required, controller 12 is designated as the primary controller (P) for executing function (F4), and controller 16 is designated as the backup controller operating in (HS) mode. Controller 14 is not designated as the backup controller (NE) for (F4), which operates in (M2) mode.
[0026] For the function (F3) operating in operating mode (M1), where only two controllers are required, controller 16 is referred to as the primary controller (P), while controller 14 is referred to as the backup controller operating in (HS) mode. Controller 12 is not required and is not referred to as backup. As also described in Fig. As shown in Figure 2, each of the controllers is configured to perform or not perform functions while operating in (M2) and their associated designation identifies whether each is a primary backup or not as used in the manner shown.
[0027] For a distributed approach, a lookup table is customized for each controller, and each customized lookup table is stored in a memory location for each controller. For a centralized approach, the lookup tables for all controllers are stored in a single controller called the coordination controller. In the centralized approach, the coordination controller implements state mode changes for all controllers. The coordinator controller communicates controller messages to each controller via the communication network to switch functions to a different state mode. If the coordinator controller fails, a backup coordinator controller is released to act as the coordinator controller. The backup coordinator controller can be selected through an assigned agreement protocol or through a statically defined sequence.The synchronization of the current coordinator controller and the backup coordinator controller communicate via the communication network described in the co-pending US application US 2017 / 0 277 607 A1 entitled “Tolerance pattern and switching protocol for multiple hot and cold standby redundancies”.
[0028] Fig. Figure 5 illustrates an initial configuration for each of the controllers along with the lookup table that illustrates whether function execution from a respective controller is required for each operating mode. Also shown is a percentage utilization of each controller based on the assignment of each function. The usage of each operating mode is illustrated for each controller. For operating mode (M1), controller 12 is used 40%, controller 14 is used 60%, and controller 16 is used 60%. For operating mode (M2), controller 12 is used 60%, controller 14 is used 40%, and controller 16 is used 80%. The maximum utilization for each controller based on operating modes (M1 and M2) is 60% for controller 12, 60% for controller 14, and 80% for controller 16.The problem presented here is that the respective controllers do not operate at the same utilization rate between the two modes of operation. For example, although controller 12 is only used 40% of the time when performing functions for (M1), controller 12 is still operating 60% of the time when performing functions for (M2). Therefore, the maximum total utilization time for controller 12 is 60%. Similarly, controller 14 operates 60% of the time when performing functions for (M1), but 40% of the time when performing functions for (M2). Therefore, the maximum total utilization time for controller 14 is 60%. The maximum total utilization determines the sizing of the hardware resources for each controller. In . Fig. 5 also shows the redundancy table 24, which includes the usage for each software function. The utilization of each controller can be determined by adding the utilization of each function. The usage of each controller for a function is only added when the function is in a primary state or a hot state. For example, in mode M1 for controller 12, functions F1 and F2 are in primary state mode. Functions F3 and F4 are in NE mode and are not being used. Therefore, F1 and F2 combine using redundancy table 24, for 40% utilization. In another example, in mode M1 for controller 16, functions F2 and F3 are in hot state and primary state mode, respectively. F1 is in cold state mode and F4 is in NE state mode, both unused. As a result, the overall utilization rate for controller 16 in M1 is 60%.
[0029] Fig. Figure 6 illustrates a remapping of the execution of a set of functions between two or more controllers using a coordinated heuristic switching technique. By remapping the execution of the set of functions between two controllers, efficiency can be achieved by minimizing the load within at least one controller. With reference to Fig. 5, the functions executed in the first controller for (M1) involve a 40% utilization, and the set of functions executed in the second controller involve a 60% utilization. In addition, the functions executed in the first controller for (M2) involve a 60% utilization, and the set of functions executed in the second controller involve a 40% utilization. By remapping the set of functions between controller 12 and controller 14 for (M2), controller 12 operates at 40% utilization in mode (M1) and 40% utilization in mode (M2), as shown in Fig. 6. As a result, a maximum utilization rate of 40% is obtained for controller 12 in both modes, which represents a 20% utilization reduction for the overall system. It should be noted that although functions have been reassigned between controllers, the execution requirement of functions in lookup table 25 remains unchanged, and the reconfiguration still satisfies the execution requirement.
[0030] Fig. Figure 7 illustrates a flowchart for a method for coordinating multi-mode allocation for runtime switching.
[0031] In step 50, we choose the permutation (G1,1; G1,2; ...; G1,n) as the mapping for all controllers for the first mode (mode 1). The notation Gi,j indicates the mapping of functions in mode i to controller j. Note that each Gi,j is given as input to the algorithm and can be determined based on any state-of-the-art software mapping algorithm. Thus, i is the index for modes (there are m modes) and j is the index for controllers (there are j controllers). In an improved variation of the algorithm, all permutations (G1,1; G1,2; ...; G1,n) for each controller in mode 1 are identified, and the entire flowchart is executed for each of these permutations, yielding the initial mapping for mode 1.
[0032] In step 51, a load for each of the controllers is determined based on the use of each controller performing the functions in the first mode.
[0033] In step 52, a determination is made as to whether all modes have been evaluated. If all modes have been evaluated, the program continues to step 57, where the routine ends and the lookup tables are generated; otherwise, the routine continues to step 53.
[0034] In step 53, the mode is indexed to the next mode (mode i). The routine starts with mode 1 and then indexes the next mode, as the routine loops.
[0035] In step 54, a permutation of (Gi,1; ...; Gi,n) is selected that satisfies all design constraints and results in the lowest overall utilization, taking into account the current utilization of the controllers based on the already assigned modes. This means that a coordinated determination is performed to determine the minimum utilization for each controller for the mode to be determined and coordinate it with the assignments already determined for the previous operating modes. In this step, all permutations of (Gi,1, Gi,2, ..., Gi,n) are examined.
[0036] In step 55, a coordinated allocation is performed in which sets of allocated functions intended for the controllers are exchanged based on the permutation selected in the previous step (this selected permutation was chosen to obtain the most efficient overall utilization for each controller across each of the indexed modes). As in Fig. As shown in Figure 6, the total utilization for each controller across each mode is determined by swapping the function assignments between the controllers and determining a lowest total utilization for one or more controllers based on the configured assignments between each of the modes. For example, if the permutation (Gi,3, Gi,1, Gi,2) is selected, a swap is performed to obtain the following assignments:
[0037] The assignment of controller 1 is assigned to Gi,3 for mode i (i.e., the original assignment of controller 3 for mode i);
[0038] The assignment of controller 2 is assigned to Gi,1 for mode i (i.e., the original assignment of controller 1 for mode i); and
[0039] The assignment of controller 3 is assigned to Gi,2 for mode i (i.e., the original assignment of controller 2 for mode i).
[0040] A return is made to step 52 to determine if more modes require analysis.
[0041] After the algorithm is complete, which involves assigning functions and determining their respective states on each controller for each operating mode, the lookup table is generated for each controller, and each lookup table is stored in the memory of the respective controller it designates. Alternatively, all lookup tables can be stored in a coordinating controller, with a centralized approach for executing the functions in the various modes.
Claims
[1] A method for reallocating control functions based on minimizing utilization, the method comprising the following steps: generating a lookup table (25) based on functions and operating mode, each entry in the lookup table (25) containing a number of executions required for a respective function in a respective operating mode; allocating functions to execute to the controllers (12, 14, 16) based on the number of executions required for a function in a respective mode of operation, wherein each controller (12, 14, 16) is designated as one of a primary state, a backup state, or a non-execution state for each function; determining a utilization rate for each controller (12, 14, 16) in each operating mode; determining a minimum utilization of each controller (12, 14, 16) over each operating mode; comparing the utilization rates of the different operating modes for each of the controllers (12, 14, 16); identifying the appropriate utilization rates between controllers (12, 14, 16) of different operating modes; and coordinating multi-mode remapping of function execution is coordinated in the controller (12, 14, 16) by switching a set of pre-assigned functions between different controllers (12, 14, 16) within a respective operating mode to reduce the utilization rate of at least one controller (12, 14, 16). [2] Method according to claim 1, wherein the lookup tables (25) are predetermined. [3] Method according to claim 2, wherein the functions within the individual controllers (12, 14, 16) are to be assigned in advance to each controller (12, 14, 16). [4] The method of claim 1, wherein comparing the utilization rates for each individual controller (12, 14, 16) is determined based on the usage of the controllers (12, 14, 16) between the different modes of operation. [5] The method of claim 1, wherein identifying matching utilization rates includes identifying an exact match of utilization rates between controllers (12, 14, 16) of different operating modes. [6] The method of claim 1, wherein identifying appropriate utilization rates includes identifying the lowest utilization rates among the controllers (12, 14, 16) in the different modes of operation. [7] The method of claim 1, wherein each controller (12, 14, 16) stores the lookup tables (25) in a memory. [8] The method of claim 1, wherein lookup tables (25) are stored in a memory of a single controller (12, 14, 16). [9] The method of claim 8, wherein the controller (12, 14, 16) storing the lookup table (25) is referred to as a primary coordinator controller (12, 14, 16), the primary coordinator controller (12, 14, 16) implementing status mode changes for all controllers. [10] The method of claim 1, wherein a uniform lookup table (25) is created for each controller (12, 14, 16) that identifies the status modes of the controller (12, 14, 16) for each function and within each mode of operation, each lookup table (25) within each controller (12, 14, 16) being used by each controller (12, 14, 16) to coordinate a respective status mode change of each controller (12, 14, 16).
Citation Information
Patent Citations
Correlating cross process and cross thread execution flows in an application manager
US20070143323A1
Systems and methods for affinity driven distributed scheduling of parallel computations
US20110099553A1
Methods, systems, and computer readable media for generating simulated network traffic using different traffic flows and maintaining a configured distribution of traffic between the different traffic flows and a device under test
US20140160927A1
Scheduling in a multicore architecture
US20140282593A1