Methods for coordinating access to resources of a distributed computer system, computer system and computer program

A hierarchical resource management system with supervisor managers in each region addresses the challenge of maintaining continuous and safe communication in distributed systems by planning for changing links, ensuring efficient and reliable data transmission.

DE102021111950B4Active Publication Date: 2026-02-19TECHNISCHE UNIVERSITÄT BRAUNSCHWEIG KÖRPERSCHAFT DES ÖFFENTLICHEN RECHTS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102021111950
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-07
Publication Date
2026-02-19
Estimated Expiration
2041-05-07

AI Technical Summary

Technical Problem

Existing methods for managing resources in distributed computer systems, particularly in complex environments like connected vehicles, struggle to ensure continuous and safe transmission of large data sets while adapting to changing communication links, often leading to disruptions and inefficiencies.

Method used

A hierarchical resource management system where each region has a supervisor resource manager to oversee and coordinate other managers, planning and ensuring safe transitions of communication links by anticipating worst-case scenarios and maintaining critical connections.

Benefits of technology

Ensures continuous and safe communication by proactively managing reconfigurations, minimizing disruptions, and maintaining quality of service requirements even in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for coordinating access to resources of a distributed computer system comprising several distributed participant stations (22, 23), wherein the computer system is hierarchically structured into regions (7, 8, 9, 10, 11, 12) and each region (7, 8, 9, 10, 11, 12) is assigned one or more participant stations (22, 23), wherein the participant stations (22, 23) each comprise at least one computer, at least one resource of the computer system, at least one resource manager (20) configured to manage the resources of the computer system assigned to it, and at least one internal communication medium via which the at least one computer, the at least one resource, and the at least one resource manager (20) are coupled internally within the participant station to carry out data communication, wherein the participant stations (22,23) are coupled to each other via at least one external communication medium for the purpose of carrying out data communication, wherein resources and / or resource managers (20) of the computer system may be coupled to an external communication medium, characterized in that in each region (7, 8, 9, 10, 11, 12) a resource manager (RM1, RM2, RM3) is assigned the function of a supervisor (13, 14, 15), who supervises the other resource managers (RM1, RM2, RM3) of his region (7, 8, 9, 10, 11, 12) and the supervisors (13, 14, 15) subordinate to him in the hierarchy, makes decisions on their requests and resolves conflicts between requests, wherein a desired end-to-end communication connection between a requester who wants to access a desired resource and the desired resource is established by the supervisor (13, 14, 15) of the region (7, 8, 9, 10, 11, 12) of the requester, if necessary, cooperatively with other participating supervisors (13, 14,15) and / or other resource managers (RM1, RM2, RM3) is planned and then manufactured.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for coordinating access to resources of a distributed computer system comprising several distributed participant stations, wherein the computer system is hierarchically structured into regions and each region is assigned one or more participant stations, wherein the participant stations each comprise at least one computer, at least one resource of the computer system, at least one resource manager configured to manage the resources of the computer system assigned to it, and at least one internal communication medium via which the at least one computer, the at least one resource, and the at least one resource manager are interconnected within the participant station for the purpose of carrying out data communication, wherein the participant stations are interconnected with each other via at least one external communication medium for the purpose of carrying out data communication.wherein resources and / or resource managers of the computer system may be connected to an external communication medium. The invention further relates to a computer system for executing such a method and a computer program for it.

[0002] Examples of such distributed computer systems include networked computer systems, such as multicore ECUs, within a vehicle or building, or vehicles networked via data connections whose respective local computer systems communicate with each other, or other computer networks, such as the internet. In the first example of a vehicle and its internal, networked systems, this could involve, for example, an ECU (Electronic Control Unit) with a multicore computer system, an internal bus or network, and a resource manager. Several of these ECUs are then connected to each other, for example, via a network (Ethernet, possibly with its own resource manager). The term "participating station" is to be understood in the broadest sense. A participating station can be, for example, a local computer or a computer system, including a multiprocessor computer system, as described in the prior art mentioned above.Participating stations can be, in particular, bus-based computer systems and router-based computer systems, or a combination thereof.

[0003] From WO 2018 / 202446 A1, a method, computer system, and computer program are already known in which a system of distributed resource managers is used. From DE 11 2020 000 054 T5, a modular I / O configuration for edge computing with disaggregated chiplets is known. From US 2021 / 0011765 A1, a time-limited edge resource management system is known.

[0004] The invention is based on the objective of providing a further improved method, computer system and computer program.

[0005] In a procedure of the type mentioned above, this task is solved by assigning a supervisor role to a resource manager in each region. This supervisor oversees the other resource managers in their region and the supervisors subordinate to them in the hierarchy, makes decisions on their requests, and resolves conflicts between requests. A desired end-to-end communication link between a requester who wants to access a desired resource and the desired resource is planned and then established by the supervisor of the requester's region, if necessary in cooperation with other participating supervisors and / or other resource managers.

[0006] This provides an even more universally applicable and broader approach to resource management in complex computer systems, particularly for distributed real-time systems. The invention is suitable, for example, for resource management in connected car applications or other applications where communication takes place via communication links that change during the operation of the computer system, e.g., via wireless communication links such as WLAN and cellular networks.

[0007] The invention takes into account, in particular, the specific characteristics of computer-controlled applications in vehicle systems, which require a complete data set for the continuation of the respective application and not just individual data packets. Such a data set can be, for example, a frame from a vehicle sensor, such as a video sensor, a LiDAR sensor, or a radar sensor, or an object of the operating system, such as ROS, AUTOSAR, or other communication data sets, such as DDS or SOME / IP. Such data sets are often several thousand kilobytes in size, or even in the megabyte range for high-resolution sensors used in automated driving. With conventional communication connections, such data volumes are usually transmitted in packets, and a large number of packets can occur. To ensure the continuous execution of processes in the computer system, i.e.,To ensure the smooth operation of the individual system components, the invention enables the assurance of Quality of Service requirements for the transmission of the complete data set within defined time periods.

[0008] In the inventive method, this is ensured by the hierarchical structure of the resource managers, in which one resource manager is assigned the function of a supervisor who oversees the other resource managers. Furthermore, the inventive method proposes a multi-step procedure for establishing communication connections or for other processes within the computer system. This procedure includes a planning phase in which the processes are planned, followed by an implementation phase for carrying out the processes. During the planning phase, worst-case scenarios regarding communication connections that change during the operation of the computer system can be taken into account.

[0009] According to an advantageous embodiment of the invention, it is provided that a) at least one resource manager checks, if necessary cooperatively with other participating resource managers, whether processes in the computer system require a reconfiguration of at least one resource required for the processes, b) If a reconfiguration is required, the resource managers involved shall prepare a plan in advance for a transition from one configuration to another of the at least one resource required for the processes, such that at least the safety-critical and / or continuous communication links of the desired end-to-end communication link are maintained.

[0010] In this way, automatic reconfiguration management can be implemented using resource managers. Particularly in vehicle applications, it can be assumed in practical operation that external events, such as a change of mobile network cell, will frequently necessitate configuration changes. The associated reconfigurations can be carried out securely by the resource managers, again following the previously mentioned planning phase. During the planning phase, the process plan is created, which is then executed to carry out the reconfiguration. A reason for a reconfiguration could be, for example, the degradation of one or more computer system resources, such as a reduction in the transmission bandwidth of a communication medium.

[0011] Advantageously, each participating station has at least one resource manager. This resource manager can have various functions; for example, it can arbitrate access within the participating station, but this is not a mandatory requirement. It is also possible for a resource manager to use an external communication medium to synchronize access, while communication on the internal communication medium remains unsynchronized.

[0012] By dividing the computer system into regions and establishing a hierarchy among the resource managers, the planning of the computer system's processes, particularly end-to-end communication links, can be collaboratively disentangled into individual control tasks, which are then handled by the respective resource manager of each region. The term "region" is not necessarily to be understood in a geographical sense. For the purposes of this invention, a region can also be defined by function or technical interrelationships. In the context of connected car applications, a region can be formed, for example, by an MSP server, a WLAN network, a cellular network, and a vehicle system. Within the vehicle system itself, a further subdivision into multiple regions is also possible.

[0013] One of the processes in a computer system might be, for example, establishing a desired end-to-end communication connection, or forcibly terminating or modifying it, for example, to ensure the continuity of safety-critical and / or continuous processes within the computer system. The following is an example of establishing an end-to-end communication connection.

[0014] The procedure includes the following steps: a) A requester wishing to access a desired resource, which may be a resource beyond the requester's subscriber station, sends a parameterized connection request to its assigned resource manager or another resource manager of the computer system to establish an end-to-end communication link between the requester and the desired resource. b) The resource manager receiving the connection request checks, if necessary cooperatively with other participating resource managers, whether the desired end-to-end communication connection requires a reconfiguration of at least one resource required for the desired end-to-end communication connection, c) If reconfiguration is required, the resource managers involved shall in advance create a flowchart for a transition from one configuration to another of the at least one resource required for the desired end-to-end communication link, such that at least the safety-critical and / or continuous communication links of the connection request are maintained. d) If the desired end-to-end communication connection to the desired resource can be established while maintaining at least the safety-critical and / or continuous communication links of the connection request, the desired resource will be reserved by the resource manager managing the desired resource according to the parameters of the connection request, and the requester will be signaled that the desired end-to-end communication connection can be established. e) Upon request of the requester, the participating resource managers will then cooperatively establish the desired end-to-end communication connection in accordance with the reservation or at least a part of the reservation, f) The requester accesses the desired resource via the established end-to-end communication connection.

[0015] A requester can be, for example, a computer within the computer system, an application, a monitor, a hypervisor, and / or a client. A requester can therefore consist of hardware components, software components, or a combination thereof. In this context, a monitor is understood as an element of the computer system configured to monitor hardware and / or software components. The monitor can be used, for example, to detect malfunctions in the computer system. A resource manager can then use monitoring information, i.e., data from the monitor, for optimization, switching, and / or fault tolerance.

[0016] The procedure may also provide that the resource manager receiving the connection request sends requests to other resource managers of the computer system to establish the desired end-to-end communication connection, at least when the connection request concerns a third-party resource that is not managed by that resource manager.

[0017] According to an advantageous embodiment of the invention, a safe transition sequence is defined for the transition from one configuration to another of the at least one resource required for the processes. Such a safe transition sequence ensures the maintenance of continuous and / or safety-critical processes in the computer system. The safe transition sequence can, for example, include reliable timing of the reconfiguration operations, taking into account the latency of message transmission and the latency of the reconfiguration process. The transition sequence can be designed to be particularly safe in that it ensures the maintenance of transmission continuity and the quality of service requirements.

[0018] According to an advantageous embodiment of the invention, it is provided that, in cooperation between the participating resource managers, end devices, network components, and applications, a complete or partial reduction of less critical process requirements is negotiated in order to ensure the maintenance of at least the safety-critical and / or continuous communication links. For example, a complete or partial reduction of less critical connection requirements can be negotiated.

[0019] According to an advantageous embodiment of the invention, the process plan is determined before the establishment of a desired end-to-end communication link through cooperation between the supervisors of all regions involved in the desired end-to-end communication link. This ensures the continuous maintenance of at least the safety-critical processes of the computer system during the subsequent execution of the process plan.

[0020] According to an advantageous embodiment of the invention, the process plan is determined prior to the establishment of a desired end-to-end communication connection through cooperation between participating resource managers and one or more terminal devices, network elements and / or applications on these terminal devices.

[0021] According to an advantageous embodiment of the invention, the process flow includes the use of permanent and / or temporary redundancy, at least during the transition from one configuration to another in the computer system, in order to ensure the maintenance of at least the safety-critical and / or continuous processes in the computer system, in particular the maintenance of the continuous operation of safety-critical applications running in the computer system and / or the safety-critical communication links. In this way, the use of redundancy can be minimized, thereby reducing associated costs.

[0022] According to an advantageous embodiment of the invention, the sequence plan is created taking into account permanent and / or transient errors that may occur during reconfiguration. This also ensures safety against such potential issues during reconfiguration and the subsequent end-to-end communication connection.

[0023] According to an advantageous embodiment of the invention, the sequence of events is designed to ensure that, in the event that a communication link is interrupted or a component involved in the communication link fails during the execution of the configuration steps of the reconfiguration, a safe configuration is adopted in which at least the safety-critical and / or continuous communication links of the desired end-to-end communication link are maintained. In this way, the pre-planning, i.e., the sequence of events, also takes such unexpected events into account, which is of great importance, for example, for the continuous operation of autonomous vehicles.

[0024] According to an advantageous embodiment of the invention, the reconfiguration is initiated by at least one resource manager and / or at least one requester. It is particularly advantageous that a reconfiguration can be initiated by a resource manager itself, for example, if the resource manager detects during data transmission monitoring that a transmission bottleneck is imminent.

[0025] According to an advantageous embodiment of the invention, at least one resource manager terminates or adjusts an existing end-to-end communication connection to ensure the continuity of at least the safety-critical and / or continuous processes within the computer system. In this way, for example, newly occurring data transmission requests can be quickly provided with sufficient bandwidth if they are particularly safety-critical or have high priority for other reasons.

[0026] According to an advantageous embodiment of the invention, it is provided that the system components of the distributed computer system can operate autonomously even if there is no connection to a resource manager or if the resource manager fails permanently.

[0027] According to an advantageous embodiment of the invention, the parameters of the connection request for a desired end-to-end communication connection comprise one, several, or all of the following parameters: - a state and / or pattern of previous resource accesses by one or more of the components of the computer system, - an electrical voltage and / or a clock frequency at which one or more of the components of the computer system are operated, - a temperature of one or more of the components of the computer system, - a number of data packets transmitted over an internal and / or external communication medium within a defined period of time, - Type, scope and / or timing of the desired resource.

[0028] The aforementioned task is also solved by a computer system comprising several distributed participant stations, wherein the computer system is hierarchically structured into regions and each region is assigned one or more participant stations, wherein each participant station comprises at least one computer, at least one resource of the computer system, at least one resource manager configured to manage the resources of the computer system assigned to it, and at least one internal communication medium via which the at least one computer, the at least one resource, and the at least one resource manager are interconnected within the participant station for the purpose of carrying out data communication, and wherein the participant stations are interconnected with each other via at least one external communication medium for the purpose of carrying out data communication.wherein resources and / or resource managers of the computer system can be connected to an external communication medium, wherein in each region a resource manager is assigned the function of a supervisor who oversees the resource managers subordinate to him in the hierarchy, makes decisions on their requests, and resolves conflicts between requests, wherein the computer system is configured to execute a method according to one of the preceding claims. The advantages described above can also be realized in this way.

[0029] Where a computer is mentioned, it may be configured to run a computer program, e.g., in the sense of software. The computer may be a standard commercial computer, e.g., a PC, laptop, notebook, tablet, or smartphone, or a microprocessor, microcontroller, or FPGA, or a combination of such elements.

[0030] The computer system can have resources such as memory, communication media, input and output channels, sensors and / or actuators.

[0031] The computer system can use a communication medium such as a data bus, a network on a chip (NoC), or a wireless communication connection, e.g., WLAN, Bluetooth, or a cellular network.

[0032] The aforementioned task can also be solved by a computer program using program code tools, designed to carry out a procedure of the type described above, when the computer program is executed on one or more computers of the computer system. The advantages described above can also be realized in this way.

[0033] The invention is explained in more detail below with reference to exemplary embodiments and drawings.

[0034] They show Fig. 1. Cooperative planning of accesses in a connected car application, Fig. 2 a first variant of a process plan, Fig. 3 a second variant of a process plan, Fig. 4. A schedule that takes latencies into account, Fig. 5 a protocol flow of a reconfiguration, Fig. 6. A change of communication links in a moving vehicle.

[0035] The Fig. Figure 1 shows a typical structure of communication connections in a connected car scenario. At the top level, there is a cloud 1, which is communicatively coupled to an edge server 3 via a company network 2. The edge server 3 is communicatively coupled to a connected vehicle 5 via an access network 4. The vehicle 5 has a multitude of control units, which form the respective participant stations 6 of the computer system. According to the invention, the computer system is structured into various regions 7, 8, 9, 10, 11, 12 to ensure safety-critical processes, e.g., a cloud region 7, an edge region 8, an edge region 9, a vehicle region 10, and a vehicle region 11, which can have further subregions 12, e.g., on-chip (ECU) regions. The entire computer system is thus divided into the various regions. Each region has a resource manager designated as the supervisor.Each region can contain additional resource managers and / or subregions with their own resource managers. The resource manager hierarchy starts at the lowest level, e.g., the on-chip resource manager, and continues through the vehicle's internal network (V5), the infrastructure of the edge server (3), up to the top level of the cloud (1), or the permanently installed telecommunications network.

[0036] Individual regions and resource managers can each define their own rules for connections. These rules take into account various communication characteristics, such as the object size of the data records, security relevance, and other conditions. This necessitates resolving a conflict of objectives in each case, determining whether an application or communication should be permitted and what properties it should possess.

[0037] For example, the rules can ensure that only connections and connection changes that do not violate the locally established rules are permitted, i.e., the rules in the respective region. Furthermore, the hierarchy among resource managers ensures that only permissible decisions are made regarding each end-to-end communication connection.

[0038] Resource managers in regions lower in the hierarchy request an end-to-end communication link for the applications they monitor. For example, applications can send communication link requests to the appropriate resource manager. The resource manager then synchronizes the communication with their supervisor at a higher level, a process that is repeated iteratively.

[0039] In each region, there is only one supervisor who makes decisions regarding conflict resolution between applications belonging to that region, as well as for requests from other resource managers in regions lower in the hierarchy.

[0040] Each region is set up to operate autonomously, e.g. to continue processes if the communication link to the superior region or to regions at lower hierarchical levels is interrupted.

[0041] In this way, the requirements of functional safety, e.g., according to ISO 26262, can be reliably met. If a communication link is interrupted, e.g., in wireless communication, the computer system, and in particular the subsystem in the vehicle 5, is configured to switch to a fail-safe operating mode. The same applies in the event of an interrupted connection between components within the vehicle, e.g., an interruption of communication links between the control units 6. This protects both the driver and the surroundings from hazards. In the case of autonomous vehicles, safe operation can also be ensured in the event of failures ("fail-operational"). The same safe operating modes can be adopted by the other components of the computer system, e.g., in the area of ​​edge infrastructure.

[0042] Edge infrastructure, in this context, refers specifically to the part of the infrastructure located near the vehicle's current position and configured for local communication with the vehicle. For example, if vehicle 5 is at a road intersection, the infrastructure components present there and communicating with the vehicle are considered edge infrastructure.

[0043] To establish an end-to-end communication link, the previously explained rules regarding the sequence of communications can be applied. For example, the negotiation process for providing an end-to-end communication link can begin in the sender's region and then proceed downstream to the receiver's region. Of course, other strategies are also possible.

[0044] End-to-end communication links can also be blocked or modified by resource managers, not only based on application requirements but also on system monitoring requirements. It is advantageous to perform such subsequent blocking or modification of communication links in a way that maintains the established hierarchy and avoids disrupting other critical processes. This is illustrated by the following example.

[0045] If the requirements of an application regarding different communication properties, e.g., object size (according to the data sets explained at the beginning), security relevance, or other requirements can no longer be met, e.g., due to a termination of the connection, a deterioration of the connection quality, e.g., in wireless communication, then the responsible resource managers initiate changes to the communication link that primarily protect the critical connections in the region.

[0046] Provided the operation of the safety-critical connections in the region is ensured, the resource managers begin cooperating with other resource managers in the hierarchy to restore the end-to-end communication link to a state that meets requirements. This process can be performed iteratively in both directions through the hierarchy, e.g., simultaneously with the lowest and highest hierarchy levels. Alternatively, the process can be performed in only one direction of the hierarchy if this direction is predetermined, e.g., in the case of problems with a wireless connection.

[0047] To ensure rapid responses, local behavior, e.g., in each region, can be planned in advance. This includes determining the reliability of connections. Furthermore, a search can be conducted for a plannable alternative in case of unreliable connections. The next higher hierarchical level is informed of the planning and can use it to determine the reliability of an end-to-end communication link and derive guarantees from it. The planning of the end-to-end communication link takes these potential guarantees into account. The network access parameters must be adjusted accordingly.

[0048] Based on the Fig. Section 2 now describes an example that is analogous to the structure of the communication links between Cloud 1 and the vehicle 5. Fig. The hierarchy starts with level 1. At each hierarchy level, a resource manager (RM1, RM2, RM3) is defined as a supervisor for a given region. At the lowest hierarchy level, resource manager RM1 is defined as supervisor 13; at a middle hierarchy level, resource manager RM2 is defined as supervisor 14; and at the highest hierarchy level, resource manager RM3 is defined as supervisor 15.

[0049] Supervisor 13 in vehicle 5 synchronizes all electronic control units (ECUs) at this lowest hierarchy level. The respective ECUs can also have their own resource managers. These then form regions that are even lower in the hierarchy. Supervisor 13 in vehicle 5 receives requests from requesters or other resource managers from lower hierarchy levels. It is also possible that several resource managers exist at the lower hierarchy levels, for example, if the vehicle's internal network is divided into a multitude of subregions 12, as shown by the Fig. 1 described. In this case, each region or subregion has a resource manager who acts as a supervisor for that region and resolves conflicts.

[0050] Resource managers or supervisors 13, 14, 15 can then schedule and execute a communication connection 17. Furthermore, in the event of specific occurrences, such as a deterioration in communication quality, one or more resource managers can adjust or interrupt the communication connection. This allows the vehicle 5 to operate autonomously even if the connection to the edge server 3 is interrupted. The same applies if one or more vehicle components fail, such as a sensor.

[0051] If communication link 17 can be established, vehicle 5 connects to the edge infrastructure, e.g., a 5G edge server, which is configured to conduct wireless communication at a road intersection with all vehicles present there, e.g., those waiting for a signal from a traffic light. In this case, the intersection is considered region 9. The supervisor 14 of this region 9 can, for example, define specific communication parameters, such as assigning defined frequencies to specific vehicles 5, e.g., based on requests from vehicle 5. The supervisor 14 can establish the connections between vehicles 5 and cloud 1, as described in more detail below. The supervisor 14 can operate autonomously, e.g., in the event of a connection failure to cloud 1 or a connection failure to a specific vehicle 5.

[0052] To establish the communication link 17, for example an application running in vehicle 5 can execute the following: First, a request is sent to the resource manager of your own region, e.g., to Supervisor 13.

[0053] Supervisor 13 in vehicle 5 evaluates the request. If the request is executable, the resource manager reconfigures the relevant components in its region, i.e., the vehicle's internal network, to meet the new communication requirements of communication link 17. Supervisor 13 can send requests within the application to the resource manager nearest in the hierarchy, i.e., in this case, Supervisor 14.

[0054] In the next step, Supervisor 14 can allocate and adjust the necessary resources for communication link 17 in its region 9. For example, Supervisor 14 can allocate wireless frequencies as communication resources. Supervisor 14 can resolve conflicts in the event of contradictory requests from different vehicles. To do this, Supervisor 14 can make decisions based on an assessment of how critical the respective communication and how reliable the respective wireless data connection is for each individual vehicle. Supervisor 14 can then decide to postpone or block certain less critical connections. Furthermore, Supervisor 14 can decide to allocate redundant resources to critical connections, such as two frequencies or multiple parallel connections of a WLAN or cellular network.

[0055] Following this, or in parallel, Supervisor 14 can establish a connection with the hierarchically superior Supervisor 15 in the Cloud 1 region to finalize the planning for communication link 17. Supervisor 14 can temporarily store data already received from Vehicle 5 and upload it to Cloud 1, even if Vehicle 5 has already left the intersection at that time.

[0056] Communication link 17 is established when the superior supervisor in the hierarchy has sent a confirmation to the subordinate supervisor. In the example according to Fig. 2 This means that Supervisor 14 must give confirmation to Supervisor 13, and Supervisor 15 must give confirmation to Supervisor 14.

[0057] To terminate a communication connection, the requester, e.g., an application, can send a request to its responsible resource manager, i.e., the resource manager of its region, e.g., its assigned supervisor. This resource manager evaluates the request. If the request is valid, the resource manager reconfigures the necessary resources in its region to fulfill the request. Additionally, the resource manager sends a request on behalf of the requester to the next resource manager in the hierarchy. For example, in the case of Fig. 2. An application in vehicle 5 requests the termination of communication link 17. This application sends a corresponding request to Supervisor 13. Supervisor 13 performs the reconfigurations in its region and sends a corresponding request to the superior Supervisor 14. These steps are also performed in the region of Supervisor 14. Supervisor 14 reconfigures the necessary resources in its region and sends a request to the superior Supervisor 15. Supervisor 15 performs the necessary reconfigurations in its region and sends an acknowledgment to Supervisor 14, the next lower in the hierarchy. Supervisor 14 then sends an acknowledgment to Supervisor 13. Now, communication link 17 can be terminated.

[0058] Another possible scenario in the case of the Fig. 2. Assume that a wireless communication link (radio link) between vehicle 5 and edge infrastructure 3 is interrupted. This interruption is detected by a resource manager in vehicle 5, e.g., Supervisor 13. A corresponding resource manager in edge infrastructure 3, e.g., Supervisor 14, also detects the interruption. The resource manager in vehicle 5 ensures that vehicle 5 remains in a safe state. The resource manager prepares the necessary resources for reconfiguration. In doing so, the resource manager releases as many resources, such as links and / or bandwidth allocations, as are required for critical communications. The resource manager informs the applications running in vehicle 5 so that they can adjust their data transmission settings accordingly.

[0059] Supervisor 14 in the edge infrastructure defines the resources previously tied up by the communication link as free and prepares the necessary resources for reconfiguration. For example, Supervisor 14 can define the frequencies reserved for the previously conducted communication as free. Supervisor 14 also informs its superior supervisor 15 in the hierarchy about the interrupted connection, enabling the latter to perform the necessary reconfigurations in its region. This information about the interrupted connection is propagated through the hierarchy of supervisors until the highest-ranking supervisor is reached.

[0060] Based on the Fig. In section 3, the example will be explained in which an application 16 in vehicle 5 directly requests an end-to-end communication connection 17 to cloud 1. This communication connection request is initially processed by supervisor 14. As mentioned, supervisor 14 is part of the edge infrastructure 3. Supervisor 14 then conducts the further negotiations for establishing the communication connection 17 not with application 16, but with the higher-level supervisor 15. When the communication connection 17 is available, supervisor 14 informs application 16 so that it can use the communication connection 17. When the communication connection can be terminated, application 16 informs supervisor 14. Supervisor 14 then informs supervisor 15. The communication connection 17 is then terminated.

[0061] Using the example of Fig. Section 4 will explain the case where reconfiguring the network used for communication typically requires transitioning the network through various intermediate states. During these intermediate states, the resource managers involved should ensure uninterrupted communication and provide a protocol that does not disrupt the performance and security of other ongoing communications.

[0062] In packet-oriented data transmission, the reconfiguration of network nodes cannot be performed immediately because packets, once introduced into the network, remain in the network buffers.

[0063] The Fig. Figure 4 shows a typical packet-oriented data transmission between subscriber stations 22 and 23, which can be located at different sites. The network contains several nodes 21, where packets are received, buffered, and then forwarded. The responsible resource manager 20 now plans and executes a network reconfiguration. For this purpose, the resource manager 20 sends reconfiguration requests across the network to subscriber stations 22 and 23. Since the path from the resource manager 20 to subscriber station 23 leads through a larger number of nodes 21 than the path to subscriber station 22, the data packets of the reconfiguration requests sent by the resource manager 20 reach subscriber stations 22 and 23 at different times. Subscriber station 23 receives the data packets with a delay T1, and subscriber station 22 with a delay T2, where T1 is greater than T2.

[0064] If the resource manager sends 20 reconfiguration messages to the various participant stations 22 and 23 simultaneously, these stations receive the messages at different times. Therefore, without specific countermeasures, participant station 22, with its shorter packet transit time T2, could adopt a changed setting earlier due to the reconfiguration message, while participant station 23 continues to operate with the old configuration for some time. This can lead to data loss and significant latency. The inventive method solves this problem through appropriate scheduling of the reconfiguration. This scheduling takes the difference in transmission times T1 and T2 into account and compensates for it during execution. In particular, it ensures that both participant stations 22 and 23 switch to a new configuration at the same time.

[0065] The packets fed in at a specific data rate, e.g., in mode 1, remain in the buffers of node 21 for a certain time, e.g., defined by the transmission latency and the interference latency due to interactions with other data streams. This can also happen if the sender has already been reconfigured and begins sending the data packets in the different mode 2.

[0066] When node 21 receives a reconfiguration message, it does not know what type of packets are in its data buffers, i.e., whether they are Mode 1 or Mode 2 data packets. As a result, internal reconfiguration at node 21 can lead to packet loss, changes in priorities, or other unsafe behavior, such as incorrect packet forwarding.

[0067] A secure transition from one mode to another can be ensured by controlling data traffic at the end nodes, for example, by reducing, increasing, or blocking it for the different transmission modes. For instance, certain connections can be blocked to clear queues in the switches in order to reconfigure these switches, such as their port arbiters or queue disciplines. In this way, secure transitions between modes can be performed even in systems with commonly available switches.

[0068] Performing these actions in the correct order can prevent the following effects: - Unreliable behavior of transmissions and data streams, e.g. during reconfiguration and with regard to all other ongoing communications. - Deadlocks or lifelocks during the transition phase, e.g., incorrect switch reconfiguration processes that can block access to certain paths. - In addition, interruptions or other disruptions in communication with the resource manager can be avoided.

[0069] The reconfiguration steps must be performed at a specific time, taking into account the latencies that affect the reconfiguration process. For example, latency is required to ensure that traffic injected into an operating mode has left the network. Latency is also required to reconfigure certain network components, such as switch reconfiguration latency. Additionally, latencies resulting from the forwarding of reconfiguration messages must be considered.

[0070] Based on the Fig. Section 5 describes an example where an error occurs during a protocol-based reconfiguration, such as the failure of a component that is to be reconfigured, messages that do not arrive, or the failure of a resource manager.

[0071] It is advantageous if the system is configured to isolate the fault and maintain or revert to a safe state and / or a fail-operational state. Fig. Figure 5 shows a schematic representation of the timeline of messages exchanged between different components of the computer system. The individual components Cloud, Switch 1, NMU, SG1, SG2, SG3, SG4, and GW are shown in the upper section of the diagram. Fig. 5. The NMU acts as the region's supervisor. The timescale runs from the upper dashed line downwards. Blocks marked with reference 30 indicate message processing, blocks marked with reference 31 indicate reconfiguration, and symbols marked with reference 33 indicate a timeout. The symbol marked with reference 32 indicates an error.

[0072] At time point 34, a request message is initially sent from the cloud to the NMU. This is processed in the NMU. The NMU then sends reconfiguration messages to Switch 1, GW, and SG1. The requested reconfigurations are carried out in blocks 31. After a certain waiting period, i.e., timeout 33, Switch 1 and SG1 send respective acknowledgments to the NMU, which are processed there. The NMU then sends reconfiguration messages to SG3 and SG4. In blocks 31, SG3 and SG4 perform the requested reconfiguration. After a waiting period (timeout 33), SG3 and SG4 acknowledge the completed reconfiguration to the NMU (time point 35). Now the NMU sends reconfiguration messages to SG1, SG2, and GW. After a waiting period (timeout 33), SG1, SG2, and GW send an acknowledgment of the completed reconfiguration to the NMU.The NMU processes these confirmations and in turn sends a confirmation to the cloud (time 36).

[0073] During the phase between times 34 and 35, the system is in a safe mode transition. From time 35 onwards, the new configuration applies to all network components.

[0074] Let us now assume that the first reconfiguration message from the NMU to GW does not arrive there or is corrupted due to a disturbance 32. To compensate for such effects caused by a single disturbance 32, the system can also provide for the multiple transmission of such messages, either sequentially via the same path and / or sequentially or simultaneously via different paths. In this way, temporal and spatial redundancy can be achieved.

[0075] As another example, let's assume that the gateway component fails during reconfiguration. Since the individual components are required to send confirmation of the completed reconfiguration to the NMU, the NMU can detect such a failure by the absence of the confirmation message.

[0076] If the gateway component fails, the network management unit (NMU) can restore the network to its previous state, for example, by sending reconfiguration messages to revert to the previous configuration. The NMU can also execute a backup configuration plan to place the network in a fail-safe or fail-operational mode. Finally, it is possible to isolate the gateway component from the network, for example, by blocking connections to the gateway component in adjacent switches or routers.

[0077] Another scenario considers the failure of the NMU resource manager. Other network components can detect this failure, for example, because the protocol stipulates that the resource manager must send defined heartbeat messages. If these heartbeat messages are absent, this can be recognized as a defect in the resource manager. The individual network components can then switch to a predefined fail-safe or fail-operational operating state.

[0078] Accordingly, the system can also be operated in a safe state even if no resource manager is present or is not in operation.

[0079] In the example according to Fig. 6 considers the case where a vehicle 5 changes its position, which in the Fig. Figure 6 illustrates this based on the vehicle positions 5a, 5b, and 5c. The respective network connection of the vehicle 5 also changes; for example, at positions 5a and 5b, there may be a connection to a cellular network 4a, and at position 5c, to a WLAN 4b. Different connections to various edge servers 3a and 3b can also be established, each connected to one of the communication media 4a and 4b.

[0080] This example is intended to illustrate the difference between reconfiguring the vehicle's internal network and the network outside the vehicle 5.

[0081] Many parameters of the vehicle's internal network are known, such as the number of nodes, their arrangement and running applications, the available system modes, possible variations in the number of nodes, and the number and details of the work performed by the running resource managers. Therefore, it is possible to predetermine various precise scenarios / strategies for cooperation between the resource managers within the vehicle.

[0082] Regarding the vehicle-external network, as in Fig.As can be seen by way of example in Figure 6, the possible changes cannot be determined in advance, as they depend on the vehicle's route. For instance, it is not possible to determine in advance how many vehicles will be at a particular intersection at a given time and which applications they will be using. Therefore, the entire infrastructure must be scalable. The invention achieves such scalability by employing a large number of resource managers that are hierarchically structured, with one resource manager designated as supervisor in each region. This ensures defined resource management even with changing external network configurations.

[0083] As mentioned, resource reconfigurations can also be initiated by one or more resource managers themselves. Therefore, reconfigurations do not need to be initiated by applications. This has advantages in many cases.

[0084] The resource manager can collect data from various sensors on the network, such as monitors that record power and / or temperature, monitors in network interfaces and ports to record data transmission volume, and detectors for security threats. Once the relevant data is collected, the resource manager analyzes it, determines the network status, and adjusts the configuration accordingly.

[0085] The resource manager can collect data from sensors reporting transient or permanent faults in network components, such as connections, cables, and ports, or security threats, such as nodes transmitting excessive amounts of data. Once sufficient data has been collected, the resource manager analyzes it and performs a reconfiguration procedure if necessary. For example, in a moving vehicle, the distance to the nearest station in a cellular network can change, potentially increasing the number of transient faults. The resource manager monitors the available bandwidth of the cellular network and informs and adjusts the transmitters accordingly.

[0086] It is also possible for the resource manager to enforce specific timing or order of data transmissions. For example, a resource manager can send messages to the end nodes in a round-robin manner. When such a message is received at a node, that node is permitted to use the network for a defined period and / or with defined Quality of Service parameters, such as transmission rate, packet rate, data volume, and so on. In such cases, applications do not need to send requests or queries to the resource manager; they simply receive acknowledgments and permissions from the resource manager.

[0087] Other possible criteria that a resource manager might use to initiate a resource reconfiguration include, for example: - According to rules and / or time criteria predetermined by the design of the computer system. - Reconfiguration can be triggered by data from end nodes and sensors. - The reconfiguration can be selected from a set of predefined configurations determined in advance during the design of the computer system.

[0088] When a resource manager initiates a reconfiguration, they can, for example, reconfigure all necessary resources in their own region. Additionally, the resource manager can inform other resource managers higher up in the hierarchy.

[0089] The resource manager, and in particular the supervisor, plays a crucial role in terminating or modifying communication connections. For example, the resource manager may frequently modify or block active communication connections to establish new configuration connections with a higher security level or priority. This can be done permanently or temporarily, such as only for the duration of a reconfiguration process. The resource manager must ensure that modifying or blocking communication connections does not lead to security-critical situations.

Claims

[1] Method for coordinating access to resources of a distributed computer system comprising several distributed participant stations (22, 23), wherein the computer system is hierarchically structured into regions (7, 8, 9, 10, 11, 12) and each region (7, 8, 9, 10, 11, 12) is assigned one or more participant stations (22, 23), wherein the participant stations (22, 23) each comprise at least one computer, at least one resource of the computer system, at least one resource manager (20) configured to manage the resources of the computer system assigned to it, and at least one internal communication medium through which the at least one computer, the at least one resource, and the at least one resource manager (20) are coupled internally within the participant station for the purpose of carrying out data communication, wherein the participant stations (22,23) are coupled to each other via at least one external communication medium for the purpose of carrying out data communication, wherein resources and / or resource managers (20) of the computer system may be coupled to an external communication medium, , characterized by, that in each region (7, 8, 9, 10, 11, 12) a resource manager (RM1, RM2, RM3) is assigned the function of a supervisor (13, 14, 15), who oversees the other resource managers (RM1, RM2, RM3) of his region (7, 8, 9, 10, 11, 12) and the supervisors (13, 14, 15) subordinate to him in the hierarchy, makes decisions on their requests and resolves conflicts between requests, whereby a desired end-to-end communication connection between a requester who wants to access a desired resource and the desired resource is established by the supervisor (13, 14, 15) of the requester's region (7, 8, 9, 10, 11, 12), if necessary cooperatively with other involved supervisors (13, 14, 15) and / or other resource managers. (RM1, RM2, RM3) is planned and then manufactured. [2] Method according to claim 1, characterized by the following procedural steps: a) at least one resource manager (20) checks, if necessary cooperatively with other participating resource managers (20), whether processes in the computer system require a reconfiguration of at least one resource required for the processes, b) If a reconfiguration is required, the resource managers involved (20) shall in advance prepare a plan of action for a transition from one configuration to another of the at least one resource required for the operations in such a way that at least the safety-critical and / or continuous communication links of the desired end-to-end communication link are maintained. [3] Method according to claim 2, characterized by , that a safe transition sequence is determined for the transition from one configuration to the other configuration of at least one resource required for the processes. [4] Method according to any one of claims 2 to 3, characterized by, that in cooperation between participating resource managers (20), end devices, network components and applications a complete or partial reversal of less important process requirements is negotiated in order to ensure the maintenance of at least the safety-critical and / or continuous communication links. [5] Method according to any one of claims 2 to 4, characterized by , that the schedule is determined prior to the establishment of a desired end-to-end communication link through cooperation of the supervisors (13, 14, 15) of all regions (7, 8, 9, 10, 11, 12) involved in the desired end-to-end communication link. [6] Method according to any one of claims 2 to 5, characterized by, that the schedule is determined prior to the establishment of a desired end-to-end communication link through cooperation of participating resource managers (20) with one or more terminal devices, network elements and / or applications on those terminal devices. [7] Method according to any one of claims 2 to 6, characterized by , that the schedule includes the use of permanent and / or temporary redundancy, at least during the transition from one configuration to another in the computer system, to ensure the maintenance of at least the safety-critical and / or continuous operations in the computer system, in particular the maintenance of the continuous operation of safety-critical applications running in the computer system and / or the safety-critical communication links. [8] Method according to any one of claims 2 to 7, characterized by, that the schedule is created taking into account permanent and / or transient errors that may occur during reconfiguration. [9] Method according to any one of claims 2 to 8, characterized by , that the procedure in case a communication link is interrupted or a component involved in the communication link fails during the execution of the configuration steps of the reconfiguration includes that a safe configuration is adopted in which at least the safety-critical and / or continuous communication links of the desired end-to-end communication link are maintained. [10] Method according to any one of the preceding claims, characterized by , that the reconfiguration of at least one resource manager (20) and / or at least one requester is initiated. [11] Method according to any one of the preceding claims, characterized by, that at least one resource manager (20) terminates or modifies an existing end-to-end communication link to ensure the maintenance of at least the safety-critical and / or continuous processes in the computer system. [12] Method according to any one of the preceding claims, characterized by that a requester is a computer of the computer system, an application, a monitor, a hypervisor and / or a client. [13] Method according to any one of the preceding claims, characterized by , that the parameters of the connection request for a desired end-to-end communication connection include one, several, or all of the following parameters: - a state and / or pattern of previous resource accesses by one or more of the components of the computer system, - an electrical voltage and / or a clock frequency at which one or more of the components of the computer system are operated, - a temperature of one or more of the components of the computer system, - a number of data packets transmitted over an internal and / or external communication medium within a defined period of time, - Type, scope and / or timing of the desired resource. [14] Computer system comprising several distributed subscriber stations, wherein the computer system is hierarchically organized into regions (7, 8, 9, 10, 11, 12) and each region (7, 8, 9, 10, 11, 12) is assigned one or more subscriber stations (22, 23), wherein the subscriber stations (22, 23) each comprise at least one computer, at least one resource of the computer system, at least one resource manager (20) configured to manage the resources of the computer system assigned to it, and at least one internal communication medium via which the at least one computer, the at least one resource, and the at least one resource manager (20) are interconnected within the subscriber station for the purpose of carrying out data communication, wherein the subscriber stations (22, 23) are interconnected with each other via at least one external communication medium for the purpose of carrying out data communication.wherein resources and / or resource managers (20) of the computer system can be connected to an external communication medium, wherein in each region (7, 8, 9, 10, 11, 12) a resource manager (RM1, RM2, RM3) is assigned the function of a supervisor (13, 14, 15) who supervises the resource managers (RM1, RM2, RM3) subordinate to him in the hierarchy, makes decisions on their requests and resolves conflicts between requests, characterized by that the computer system is configured to execute a method according to one of the preceding claims. [15] Computer program with program code means, configured to carry out the method according to any one of claims 1 to 13, when the computer program is executed on one or more computers of the computer system.

Citation Information

Patent Citations

  • Resource, security and service management for multiple entities in edge computing applications

    DE112020000054T5

  • Adaptive limited-duration edge resource management

    US20210011765A1