Self-configuring distributed system and method for operating a self-configuring distributed system
The self-configuring distributed system addresses inefficiencies in resource allocation by using an operator assistance system to process diverse sensor data, optimizing resource allocation and reducing waste through automated control actions.
Patent Information
- Application Number
- PCT/EP2025/051860
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-31
AI Technical Summary
Existing distributed systems face inefficiencies in resource allocation due to reactive or manually based scaling methods, leading to wasted resources and potential application crashes, with existing frameworks relying on incomplete forecasts and human intervention.
A self-configuring distributed system and method that utilizes internal and external sensor data through an operator assistance system to automatically generate control actions for optimizing resource allocation, reducing data volume through preprocessing and prioritizing actions based on machine learning and predefined rules.
This approach enables precise resource allocation without initial deficiencies, reducing waste and improving sustainability by leveraging a wide range of sensor data types, ensuring efficient and proactive management of distributed systems.
Smart Images

Figure EP2025051860_31072025_PF_FP_ABST
Abstract
Description
[0001] Self-configuring distributed system and method for operating a self-configuring distributed system
[0002] The present application deals with an operator assistance system for a distributed system, a self-configuring distributed system and a computer-implemented method for operating the same.
[0003] A distributed system is a combination of several independent computers of the same machine type, which present themselves to the user as a single system. By using such a large number of computers, the capacities of a distributed system can in principle be scaled almost arbitrarily. However, using many computers in a distributed system also increases the costs and it is therefore desirable to optimize the number of reserved resources in order to achieve a good balance between utilization and reserves. Resources can be both physical resources and virtual resources. When it comes to costs, both direct and indirect costs must be taken into account. Direct costs can be, for example, power consumption and other costs directly caused by the additional computers. Indirect costs of using many computers include, for example:high CO2 emissions, higher runtime complexity and / or maintenance costs.
[0004] In large, distributed systems, configuration and maintenance are a multidisciplinary and time-consuming problem, which in many cases can no longer be overseen by humans. To test the respective applications running in the clusters, established frameworks that meet specified minimum standards, such as the ISO 25010 standard for assessing software quality, are used. The results of these frameworks are then used to derive the state of the cluster and identify potential for improvement. Based on the assessment, maintenance is also performed and configurations are adjusted. Well-architected frameworks such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), or similar can also be used.
[0005] Distributed systems, such as cloud systems, generally offer the possibility of software-controlled capacity management, such as vertical scaling and / or horizontal scaling. Vertical scaling involves adjusting the capacities of individual computers in the distributed system, for example, by increasing or decreasing CPU power, available storage, memory, etc. Horizontal scaling involves changing the number of virtual machines in the distributed system. For example, scaling can be achieved by changing the number of replicas or threads. Both vertical and horizontal scaling can be performed on an application-by-application basis in existing systems. The scaling triggers are typically defined by humans and are often based on an educated guess as to what performance will be required in the future.For example, a user defines the upper and lower limits of scaling and the possible resources that can be used. The selected limits are usually based on an estimate using historical data, which is incomplete or erroneous. As a result, the selected limits are sometimes incorrect, resulting in wasted resources at best and application crashes or security vulnerabilities at worst. Scaling decisions are often difficult to understand, and manual scaling leads to frequent restarts and, in many cases, excessive and therefore costly resource reservation.
[0006] In other words, scaling in existing systems is either reactive, meaning resources are only increased in response to triggers that have already occurred, such as deficiencies, or resources are manually allocated based on long-term forecasts (cf., for example, traditional IT capacity planning), which typically leads to the reservation of partially unused resources for long periods of time. One of the disadvantages of this is that these resources are then unused or unavailable for other applications.
[0007] It is therefore the object of the present application to provide an operating method for a self-configuring distributed system and a self-configuring distributed system.
[0008] This object is achieved by the method and system according to the independent claims. Further advantageous embodiments are described in the dependent claims and in the following description.
[0009] The application is directed to a computer-implemented method for the automatic configuration of a distributed system. The distributed system comprises a plurality of servers, which can also be referred to as nodes. Furthermore, the terms "distributed system" and "cluster" are used synonymously below. Any plurality of virtual machines can be considered a distributed system. Thus, data centers, for example, can also be considered distributed systems. The method according to the invention initially comprises receiving sensor data from the distributed system, distinguishing between internal sensor data of the distributed system and external sensor data of the distributed system.
[0010] Internal sensor data of the distributed system can be understood to mean, in particular, sensor data that can be measured within the servers of the distributed system, such as utilization or utilization curves of individual servers or CPUs, utilization or utilization curves of storage and memory, amount of data inputs or outputs (I / O), etc.
[0011] Further examples of internal sensor data of a distributed system can be found in the Kubernetes metrics directory, see for example https: / / kubernetes.io / docs / reference / instrumentation / .
[0012] Furthermore, external data of the distributed system can be understood to include, in particular, sensor or measurement data that is measured outside the distributed system, such as utilization data of other systems, environmental data such as weather, power consumption, security reports, etc.
[0013] Further examples of external sensor data from a distributed system can be found at -metrics-to-measure-im-
[0014] -of-cloud-finops?hl=en and at https: / / sre. oo le / sre-book / service-le vel-objectives / , see also ISBN 9181491929124.
[0015] Which sensor data is collected can preferably be defined in advance by a user. Alternatively, the sensor data to be collected and received can also be determined automatically.
[0016] The user can select the range of functions to be used from a curated list, which defines which sensor data or metrics must be collected, but not necessarily by whom (e.g., CPU usage can be collected from different systems, through different implementations, different software solutions, or other providers; the function can include implementations for different providers; or the user selects a function; the system recognizes AWS as the provider and can thus use all AWS metrics; if it recognizes Azure, it can use all Azure metrics; if no providers are present, it can only use standard metrics, etc.). The user can also add functions and must then manually define the metrics and provider implementations within the function.
[0017] The sensor data can be in numerical form, as scale values, counters, time series or histograms, for example.
[0018] In particular, the received and / or recorded sensor data can also be used to record time series of data, which also makes it possible to correlate different and independently measured sensor data with each other over the measurement times.
[0019] The acquired sensor data can first be prepared before processing, meaning the data is cleaned and freed from noise and measurement artifacts. Furthermore, the data can be enriched, for example, by deriving or calculating metrics, histograms, or other statistical properties of the sensor data.
[0020] In a further step, relevant properties of the distributed system are extracted from the received sensor data based on extraction rules.
[0021] Automated computer-implemented processes are used for this purpose. These processes can be implemented rule-based or based on statistical methods. For example, a pre-trained classifier can be used for this purpose. Furthermore, conditions for extracting the relevant information can be defined automatically or by a user, for example, in the form of weights or priorities. For example, thresholds can be used in the classification; for example, a usage load below 30% can be classified as low usage, and above 80% as high usage. The classification can be part of the function or a dependency of the function.
[0022] Functions specify which actions / categories they fulfill. Each action has a base value that is adjusted to the current context through weighting and prioritization. For example, actions that remove defective states can be executed first, followed by cost optimization, and only then sustainability. In each case, actions that fulfill the targeted user-defined goal should be given a higher weighting.
[0023] The extraction step is also referred to as "preprocessing".
[0024] The goal of the extraction step is to pre-process and isolate the relevant data while significantly reducing the amount of data to be processed.
[0025] The extraction rules, which can also be referred to as preprocessing rules, determine which data is considered during extraction or preprocessing and can be defined flexibly. The extraction rules can be defined manually or they can be determined using a learning algorithm and / or training data. In the context of this patent application, however, only the existence of at least one function that produces expected values is important. SWAP-C is particularly relevant in this context; see https: / / www.baesystems.com / en-us / defini- which functions can be used.
[0026] In a further step, control actions are generated based on the relevant properties that can change the configurations and resource capacities of the distributed system.
[0027] The control actions can, for example, reduce or increase allocated resources. For example, the number of replicas, servers, or virtual machines can be changed, services or scalers can be added or removed, tasks can be scheduled, etc. All functions offered by the cluster can be controlled via the included control actions, as well as external functionalities (not accessible by default in the cluster) that are then executed in external services.
[0028] One advantage of the described method is that it can utilize various types of internal and external sensor data. This eliminates the limitations of existing systems that only consider selected data (such as CPU utilization).
[0029] By selecting modules, high-level abstract goals can be linked to technical implementations. For example, a high-level goal could be understood as "My systems must not waste more than 10% of resources." The corresponding module then contains functions and actions that monitor and optimize resource waste. It is therefore possible that not all functions are available; for example, some measurement data or functions related to external training information or CO2 measurements may not be available. In this case, only a subset of the module functions and control actions are used to achieve the goals. The functions are thus technical implementations of the high-level goal and represent a link. This also allows the metrics to be cleaned up again if the high-level goal is no longer met.
[0030] Thus, in the manner described, a method, and analogously a system, can be created that works better than the reactive methods described above in the sense that resources are allocated without deficiencies having to initially arise. At the same time, the method or system also works better than methods based on long-term forecasts because the physical and virtual resources are allocated more precisely and thus fewer resources are wasted. For example, the internal sensor data of the distributed system can include data on the utilization of CPUs, storage, memory and / or input. Furthermore, the external sensor data of the distributed system can include, for example, ambient data, environmental data and / or historical data. Data on the power consumption of the data center can also be recorded as sensor data.
[0031] This allows the measured and used sensor data to be flexibly selected and defined, ensuring that as much sensor data as possible can be considered when generating control actions. "A lot of sensor data" here doesn't necessarily mean mass or a large number, but rather refers to "observability," i.e., the observability of the context in which the sensor data is generated. The larger the visible context, the greater the chances that the system will converge on an optimal value.
[0032] Furthermore, sensor data of different types can preferably be taken into account, so that a plurality, preferably depending on the available resources more than 10, more than 100 or even several thousand or million different types of sensor data are recorded and taken into account when generating the control actions (see Twin Tower Architecture, https: / / cloud.google.com / blog / products / ai-machine-learn- ing / scaling-deep-retrieval-tensorflow-two-towers-architecture?hl=en).
[0033] The extraction step can serve to preprocess the external and internal sensor data, preferably by classifying the sensor data, fusing the sensor data and / or processing the sensor data with machine learning methods, whereby the relevant properties of the distributed system are extracted from the sensor data.
[0034] Preferably, the relevant properties have a significantly reduced data volume compared to the received sensor data, for example, by at least a factor of 10 or preferably by at least a factor of 100 or 1000. This allows subsequent processing to be carried out efficiently based on the relevant properties of the distributed system.
[0035] Furthermore, the generation of control actions may include prioritization and / or weighting of possible actions.
[0036] The present application further comprises an operator assistance system for the automatic configuration of a distributed system.
[0037] The operator assistance system includes the following modules:
[0038] A sensor module configured to receive internal sensor data and external sensor data of the distributed system;
[0039] An extraction module configured to extract relevant properties of the distributed system from the external and internal sensor data based on extraction rules; and
[0040] An action module configured to generate control actions for the distributed system based on the relevant properties of the distributed system.
[0041] Optionally, the operator assistance system may further comprise the following modules: a weighting module configured to weight the control actions generated by the action module; a selection module configured to select the weighted control actions; a presentation module configured to present the actions generated by the action module or to present the actions selected by the selection module.
[0042] In the implementation of the operator assistance system, the functions used know the actions they can generate. An action is assigned to the IS025010 / Wellarchitektur framework and has defined positive and negative effects.
[0043] For example, actions can be weighted using the following guidelines.
[0044] 1: Distance between the current state and the desired state of the system. Actions that improve goals already achieved are less important. Goals that improve goals not achieved are more important.
[0045] 2: Actions should be preferred that have a large (potentially cumulative) overall impact through few changes.
[0046] 3: Actions are weighted differently depending on the state and need. Need means, for example, that the system to be configured is first brought into a state that is as error-free as possible. Then, security actions are selected, followed by cost actions, and finally, sustainability actions.
[0047] After weighting the actions, the selection process takes place, which limits the changes to be made to highly weighted actions. For example, the top 10, the top 15, or a random selection of 10 actions from the top 15 can be performed.
[0048] Finally, user approval of the changes may be required.
[0049] Rejected recommendations can be considered by the system for future weighting and selection.
[0050] Features described with regard to the method can also be applied analogously in connection with the operator assistance system and vice versa.
[0051] Furthermore, the present application also encompasses a self-configuring distributed system comprising at least a first distributed system comprising a plurality of servers and an operator assistance system as described above. The servers of the first distributed system are configured to execute the actions generated by the action module of the operator assistance system.
[0052] Optionally, the self-configuring distributed system can also include a second distributed system. In this case, the operator assistance system is configured to jointly process the sensor data from the first distributed system and the sensor data from the second distributed system to generate joint control actions for the first and second distributed systems.
[0053] Further detailed embodiments are described below with reference to the figures.
[0054] Figure 1 is a schematic representation of an operator assistance system and a distributed system;
[0055] Figure 2 is a flow chart of the described method;
[0056] Figure 3 shows a detailed representation of an exemplary design of the operator assistance system.
[0057] In a distributed system there are various components, which are first defined below.
[0058] A distributed system, also called a cluster, comprises multiple nodes that can correspond to physical or virtual machines.
[0059] Furthermore, a cluster consists of a multitude of namespaces, each of which installs deployments or replica sets. A deployment consists of a replica set, where each replica set contains a multitude of identical pods. Each pod, in turn, consists of a multitude of non-identical containers. The smallest unit that can be controlled or orchestrated by the cluster is the pod. Although containers are generally known to the cluster, they cannot be directly orchestrated.
[0060] So-called autoscalers are known from the literature, which implement vertical or horizontal scaling of the cluster based on individual container metrics (e.g., CPU or memory). Details can be found, for example, at https: / / www.kubecost.com / kubernetes-autoscaling / kubernetes-vpa, where the "Vertical Pod Autoscaler" (VPA) and "Horizontal Pod Autoscaler" (HPA) of the Kubecost system are described.
[0061] However, the system described there still needs to be improved, as vertical and horizontal scaling can currently only be combined with great risks or users often plan too large amounts of resources, which leads to waste of resources.
[0062] Figure 1 shows a distributed system 200 comprising a plurality of servers 210. The servers are networked and can cooperate in processing orders.
[0063] Furthermore, Figure 1 shows an operator assistance system 100, which includes a sensor module 110, an extraction module 120, and an action module 130. The actions generated by the action module can be used to change or influence a configuration of the distributed system 200.
[0064] In particular, sensor data can be received by the sensor module 110, wherein a distinction is made between sensor data measured in the distributed system 200 (so-called internal sensor data) and sensor data measured outside the distributed system (so-called external sensor data).
[0065] It should be noted that modules represent logical separations and do not necessarily have to be implemented separately. The system that processes the sensors can, for example, also implement extraction. For example, the sensor module 110 can receive or collect data on the utilization of the distributed system, in particular on the utilization of the individual servers 210 in the distributed system 200. Utilization data can include the utilization of CPUs, storage, RAM, etc. In addition, the sensor module can measure or receive other data, such as environmental data (temperature, humidity, etc.), training information, safety reports, etc.
[0066] Relevant properties of the distributed system 200 are then extracted from the received sensor data by the extraction module 120. Extraction rules are used for the extraction, based on which the relevant properties are extracted. Extracting the relevant properties corresponds to preprocessing of the sensor data and serves in particular to reduce the data in order to enable more efficient subsequent processing. The goal of extracting the relevant properties is to extract the sensor data that is particularly relevant for the further configuration of the distributed system. The extraction rules can also be used to define or capture relationships between the sensor data, with particular relevance being found in the interaction of the respective sensor data.
[0067] The relevant properties are extracted in such a way that the data volume is significantly reduced compared to the received sensor data, for example, by at least a factor of 10, preferably by at least a factor of 100 or 1000. This means that processing of the relevant properties and the generation of control actions based on the relevant properties can be carried out much faster than if all sensor data were taken into account.
[0068] Based on the relevant properties of the distributed system, the action module then generates control actions for the distributed system. For example, the control actions can include increasing or decreasing resources.
[0069] All or at least some of the generated control actions can then be made available to a user for selection. The control actions can then be weighted in a weighting step to enable prioritization of the control actions. Furthermore, control actions can also be selected automatically, for example, through a selection step that automatically executes the highest-weighted control actions or that executes all control actions above a specified weighting.
[0070] Figure 2 shows a flowchart of the steps performed by the described method. The method is carried out by a computer or computer system that performs the tasks of the operator assistance system in Figure 1.
[0071] In step S301, the internal sensor data of the distributed system is received. In a further step S302, the external sensor data of the distributed system is received. Steps S301 and S302 can be executed in any order or simultaneously.
[0072] After receiving the sensor data, the relevant properties of the distributed system are extracted from the received internal and external sensor data in step S303. The data volume is reduced using extraction rules or an extraction method, so that the extracted relevant properties have a significantly smaller data volume and can thus be processed and analyzed much faster than the total amount of sensor data.
[0073] In a further step S304, control actions for the distributed system are then generated based on the extracted relevant properties of the distributed system.
[0074] When extracting the relevant properties of the distributed system from the sensor data, a pre-trained classifier or another system based on machine learning principles is preferably used. When generating actions by the action module 130, the generated action can be based on the initial state or the current state of the distributed system and can further utilize at least part of the classified information.
[0075] Actions include completed changes in an implementation, context, and explanation of why they are appropriate, e.g., projections and impacts, and / or a classification in IS025010 (what is being improved here, e.g., costs, complexity, etc.). Both positive and negative impacts are possible. For example, reducing the number of replicas has the positive effect of reducing costs and the negative effect of reducing availability.
[0076] In addition to implementing actions, the system can also issue recommendations. Recommendations are potential actions that must be approved by an authorized user.
[0077] The information from the classification and the selective display of the actions to the user can create a basis for a particularly simple insight into properties of the distributed system and its environment, which can then be used to justify the proposed actions.
[0078] In an exemplary alternative embodiment of an operator assistance system 400 shown in Figure 3, the modules of the operator assistance system 100 from Figure 1 are grouped in an alternative manner. The information sources 410 correspond to the sensor data acquired by the sensor module 110. Furthermore, the embodiment of Figure 3 comprises a processing module 422, an enrichment module 424, and a classification module 426. These modules preprocess the values obtained from the sensor data so that the relevant properties of the distributed system necessary for determining the actions to be performed can be extracted. In particular, the processing module 422 is configured to clean the values, for example, to remove noise from the data, and to store the data in a manner that allows for further processing.The enrichment module 424 is further configured to derive secondary metrics from the data and thus enrich the received and cleaned values with metrics, histograms and / or connections between the values.
[0079] The action module 430 then generates the possible control actions for the distributed system 200, as in the action module 130 of Figure 1.
[0080] Furthermore, in addition to the action module 430, the operator assistance system 400 further comprises a weighting module 440 and a selection module 450. The weighting module 440 is configured to weight and thereby sort the actions, and the selection module 450 is configured to select the actions that are presented to the user or the operator. The selection can be made, for example, based on a threshold value, so that only actions that exceed the threshold value after the weighting are presented and suggested.
[0081] Alternatively or additionally, it is also possible that the number of actions is already limited during the generation of the actions, so that the weighting by the weighting module 440 and selection by the selection module is then only carried out based on the limited number of actions generated.
[0082] Alternatively, the highest-weighted actions can also be implemented autonomously and without prior selection by a user by the operator assistance system 400. However, in this case, too, the autonomously performed operations are traceable, so that a user can understand the actions even afterward.
[0083] Generated actions are documented, and it is recorded whether they were applied or not. Based on this, a feedback loop can be built that can be used for future action generation or for audits.
[0084] It is also possible for advantageous, i.e., highly weighted and / or selected, control actions of the operator assistance system to be displayed to a user via a human-machine interface. A presentation module can be used for this purpose.
[0085] The presented control actions can be further enriched with supplementary information, such as the result of a performed classification, to enable users to gain a deeper understanding of the proposed actions. In particular, the supplementary information can also include positive or negative effects as well as impact projections of the control action, thereby improving the user's decision-making ability.
[0086] Furthermore, and particularly preferred, is the implementation in which the operator assistance system autonomously and automatically selects and executes actions based on the properties of the distributed system in order to transfer the distributed system to a favorable state. For example, unused resources can be automatically removed. Due to the large amount of sensor data used and the properties and information generated from it, actions are selected with a high probability that contribute to improving the operational sustainability and reliability of the distributed system. External sensor data, in particular, and the combination with internal sensor data, allows previously opaque resource and scaling decisions to be specified declaratively.
[0087] For the system described, a variety of different properties are conceivable, which can be collected or extracted using internal or external sensor data.
[0088] For example, carbon emissions could be extracted as a property of the applications running in the distributed system by combining usage data from CPU, memory, persistent storage, I / O operations and / or the power consumption of the hardware components.
[0089] Alternatively or additionally, dependencies between applications or the states of different applications can be taken into account during classification. The properties listed in this description are not exhaustive in the sense that additional properties are also conceivable in addition to those described.
Claims
Patent claims 1. A computer-implemented method for automatically configuring a distributed system (200), wherein the distributed system (200) comprises a plurality of servers (210), the method comprising the following steps: Receiving (S301) internal sensor data of the distributed system (200); Receiving (S302) external sensor data of the distributed system (200); Extracting (S303) relevant properties of the distributed system (200) from the external and internal sensor data based on extraction rules; Generating (S304) control actions for the distributed system (200) based on the relevant properties of the distributed system (200).
2. The computer-implemented method of claim 1, wherein the internal sensor data of the distributed system (200) comprises data on CPU utilization, memory, storage, and / or input; and / or wherein the external sensor data of the distributed system (200) comprises environmental data, environmental data, and / or historical data.
3. A computer-implemented method according to any one of claims 1 or 2, wherein the internal and / or external sensor data comprises a plurality of different sensor data types.
4. Computer-implemented method according to one of claims 1 - 3, wherein extracting relevant properties of the distributed system (200) from the external and internal sensor data comprises preprocessing the sensor data, preferably classifying the sensor data, fusing the sensor data and / or processing the sensor data with machine learning methods.
5. Computer-implemented method according to one of claims 1 - 4, where the relevant properties comprise a significantly reduced data volume compared to the received sensor data.
6. Computer-implemented method according to one of the preceding claims, wherein generating the control actions further comprises prioritizing and / or weighting possible actions.
7. Operator assistance system (100, 400) for the automatic configuration of a distributed system, comprising A sensor module (110) configured to receive internal sensor data and external sensor data of the distributed system; An extraction module (120) configured to extract relevant properties of the distributed system (200) from the external and internal sensor data based on extraction rules; An action module (130) configured to generate control actions for the distributed system (200) based on the relevant properties of the distributed system (200).
8. The operator assistance system (100, 400) according to claim 7, further comprising at least one of the following modules: a weighting module (440) configured to weight the control actions generated by the action module (130, 430); a selection module (450) configured to select the weighted control actions; a presentation module (460) configured to present the actions generated by the action module (130, 430) or to present the actions selected by the selection module (450).
9. Self-configuring distributed system (200) comprising an operation assistance system (100, 400) according to claim 7 or 8 and further comprising: A first distributed system (200), the first distributed system (200) comprising a plurality of servers (210); Wherein the servers (210) of the first distributed system (200) are configured to execute actions generated by the action module (130, 430) of the operator assistance system (100, 400).
10. Self-configuring distributed system (200) according to claim 9, further comprising a second distributed system (200); - wherein the operator assistance system (100, 400) is configured to To process sensor data of the first distributed system (200) and the sensor data of the second distributed system (200) together in order to generate common control actions for the first and second distributed system (200).
Citation Information
Patent Citations
Data Center Collective Environment Monitoring and Response
US20210311509A1