Resource management method and electronic equipment

By building an event-driven collaborative control bus and a lightweight time series prediction model between OpenBMC and the container orchestration platform, the problem of lagging fault response in the collaborative management of OpenBMC and containerized applications is solved, enabling real-time state synchronization and predictive container migration, and improving the system's energy efficiency, performance and reliability.

CN121255367APending Publication Date: 2026-01-02INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511815154.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In the field of collaborative management of the basic management controller OpenBMC and containerized applications, the fault response mode in existing technologies is passive and lagging, resulting in a high risk of business interruption and a lack of ability to predictively avoid potential hardware failures.

Method used

By constructing an event-driven collaborative control bus, real-time, bidirectional state synchronization and event linkage between the controller and containerized applications are achieved. Combined with intent-driven unified resource orchestration and hardware health prediction, a lightweight time series prediction model is used to identify potential hardware failures in advance, and predictive container migration is performed through a rule engine-driven cross-domain automated response mechanism.

Benefits of technology

It significantly improves the system's ability to coordinate and optimize energy efficiency, performance and reliability, effectively avoids service interruptions caused by hardware degradation, enhances system availability, and realizes closed-loop collaborative management of controllers and containerized applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121255367A_ABST
    Figure CN121255367A_ABST
Patent Text Reader

Abstract

The invention discloses a resource management method and electronic equipment, and relates to the technical field of cloud platforms, and the method comprises the steps: receiving a business intention instruction of a user; converting the business intention of the user into a corresponding management strategy type through an intention analysis model; obtaining a hardware strategy parameter according to the management strategy type, obtaining a container scheduling strategy parameter according to the management strategy type, and constructing a unified objective function according to the hardware strategy parameter and the container scheduling strategy parameter; generating a controller configuration command and a container scheduling strategy request through an optimization solution result of the unified objective function; and performing resource configuration between the controller and the container arrangement platform according to the controller configuration command and the container scheduling strategy request. According to the method, bidirectional state synchronization and event linkage between the controller and the containerized platform are realized, and an active migration mechanism of intention-driven uniform resource arrangement and hardware health prediction is fused, so that the collaborative optimization capability of the system among energy efficiency, performance and reliability is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud platform technology, and in particular to a resource management method and electronic device. Background Technology

[0002] In the field of collaborative management between the Open Baseboard Management Controller (OpenBMC) and containerized applications, the related fault response mode is passive and lagging, with a high risk of business interruption. Operation and maintenance activities usually begin after the hardware has already experienced an unrecoverable error and then the controller issues an alarm. If the container platform is forced to migrate containers at this time, it will lead to business interruption. At this time, the fault-alarm-manual intervention-recovery process cannot meet the high availability requirements of modern applications and lacks the ability to predictively avoid potential hardware failures. Summary of the Invention

[0003] This application provides a resource management method and electronic device. The method includes: a controller receiving a user's business intent instruction; converting the user's business intent into a corresponding management policy type through an intent parsing model; obtaining hardware policy parameters based on the management policy type, wherein the hardware policy parameters include power operating mode, fan speed control parameters, and power limit parameters; obtaining container scheduling policy parameters based on the management policy type, wherein the container scheduling policy parameters include the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters; constructing a unified optimization objective function based on the hardware policy parameters and the container scheduling policy parameters; generating a controller configuration command and a container scheduling policy request based on the optimization solution of the unified optimization objective function; and performing resource configuration optimization between the controller and the container orchestration platform based on the controller configuration command and the container scheduling policy request. This application achieves bidirectional state synchronization and event linkage between the controller and containerized applications, and integrates intent-driven unified resource orchestration and a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability.

[0004] This application provides a resource management method applied to a resource management system, the system including a server controller and a container orchestration platform, the controller and the container orchestration platform being interconnected, the method including: The controller receives the user's business intent instructions; The intent parsing model is used to translate the user's business intent into the corresponding management strategy type; The hardware policy parameters are obtained based on the management policy type. These hardware policy parameters include one or more of the following: power supply operating mode parameters, fan speed control parameters, and power limit parameters. The container scheduling policy parameters are obtained based on the management policy type. These parameters include one or more of the following: the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. Construct a unified objective function based on hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated by optimizing the solution of the unified objective function. Resource configuration is performed between the controller and the container orchestration platform based on controller configuration commands and container scheduling policy requests.

[0005] This application also provides an electronic device, which includes a server controller and a container orchestration platform, wherein the controller is interconnected with the container orchestration platform, and the controller is used for: The controller receives the user's business intent instructions; The intent parsing model is used to translate the user's business intent into the corresponding management strategy type; The hardware policy parameters are obtained based on the management policy type. These hardware policy parameters include one or more of the following: power supply operating mode parameters, fan speed control parameters, and power limit parameters. The container scheduling policy parameters are obtained based on the management policy type. These parameters include one or more of the following: the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. Construct a unified objective function based on hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated by optimizing the solution of the unified objective function. Resource configuration is performed between the controller and the container orchestration platform based on controller configuration commands and container scheduling policy requests.

[0006] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of a resource management method, the method comprising: The controller receives the user's business intent instructions; The intent parsing model is used to translate the user's business intent into the corresponding management strategy type; The hardware policy parameters are obtained based on the management policy type. These hardware policy parameters include one or more of the following: power supply operating mode parameters, fan speed control parameters, and power limit parameters. The container scheduling policy parameters are obtained based on the management policy type. These parameters include one or more of the following: the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. Construct a unified objective function based on hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated by optimizing the solution of the unified objective function. Resource configuration is performed between the controller and the container orchestration platform based on controller configuration commands and container scheduling policy requests.

[0007] This application describes a method that includes: a controller receiving a user's business intent instruction; converting the user's business intent into a corresponding management policy type using an intent parsing model; obtaining hardware policy parameters based on the management policy type, including power operating mode, fan speed control parameters, and power limit parameters; obtaining container scheduling policy parameters based on the management policy type, including the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters; constructing a unified optimization objective function based on the hardware policy parameters and the container scheduling policy parameters; generating controller configuration commands and container scheduling policy requests based on the optimization results of the unified optimization objective function; and performing resource configuration optimization between the controller and the container orchestration platform based on the controller configuration commands and container scheduling policy requests. This application achieves bidirectional state synchronization and event linkage between the controller and containerized applications, and integrates intent-driven unified resource orchestration with a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability.

[0008] This application's technical solution achieves real-time, bidirectional state synchronization and event linkage between the controller and containerized applications by constructing an event-driven collaborative control bus. It also integrates intent-driven unified resource orchestration and a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability. By using a hardware health scoring model to identify potential hardware failures in advance and trigger predictive container migration, it effectively avoids service interruptions caused by hardware degradation and greatly enhances system availability. At the same time, the rule engine-driven cross-domain automated response mechanism realizes closed-loop management of hardware events triggering container scheduling actions and container anomalies triggering hardware data collection. This breaks down the siloed management between traditional controllers and cloud-native platforms, providing solid support for intelligent operation and maintenance and root cause analysis. In short, the overall technical solution has a high degree of integration, foresight, and engineering feasibility. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1A structural diagram of the resource management system provided in the embodiments of this application; Figure 2 A first flowchart of a resource management method provided in an embodiment of this application; Figure 3 A second flowchart of a resource management method provided in an embodiment of this application; Figure 4 A detailed flowchart of the resource management method provided in the embodiments of this application; Figure 5 A schematic diagram of the modules of the resource management system provided in the embodiments of this application; Figure 6 Exemplary systems provided for embodiments of this application that can be used to implement the various embodiments of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0012] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] Controller and containerized application collaborative management refers to the deep integration of the server's basic management controller, OpenBMC, with the upper-layer containerized application platform (such as Kubernetes and other container orchestration systems) to build a unified management system across hardware and software boundaries. Therefore, how to use advanced technologies to improve the intelligence and security of file encryption has become one of the urgent problems to be solved.

[0015] Embodiments of this application provide a resource management method, such as... Figure 1 , Figure 2As shown, the method is applied to a resource management system, which includes a server controller and a container orchestration platform. The controller and the container orchestration platform are interconnected. The method includes: The controller receives the user's business intent instructions; The intent parsing model is used to translate the user's business intent into the corresponding management strategy type; The hardware policy parameters are obtained based on the management policy type. These hardware policy parameters include one or more of the following: power supply operating mode parameters, fan speed control parameters, and power limit parameters. The container scheduling policy parameters are obtained based on the management policy type. These parameters include one or more of the following: the number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. Construct a unified objective function based on hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated by optimizing the solution of the unified objective function. Resource configuration is performed between the controller and the container orchestration platform based on controller configuration commands and container scheduling policy requests.

[0016] It is understood that this application provides a method for co-managing controllers and containerized applications, which includes: building a containerized runtime environment on a server where the basic management controller OpenBMC is deployed, and establishing a bidirectional event-driven co-control bus between the controller OpenBMC and the container orchestration system (container orchestration platform) to achieve real-time communication between hardware status events and container events; The system receives user input business intents, transforms them into a unified optimization objective function that includes hardware policies and container scheduling policies through an intent parsing model, and dynamically adjusts the controller BMC configuration and container deployment based on this function. A lightweight time series forecasting model runs on the OpenBMC side, which calculates the hardware health score (HHS) in real time based on the collected hardware sensor data. When the HHS value is lower than the preset health threshold, a risk warning event is generated and the container orchestration system is notified via the cooperative control bus. The container orchestration system responds to hardware risk warnings, performs predictive migration of containers on affected nodes, and dynamically adjusts the container's Quality of Service (QoS) during the migration process. The collaborative control bus, through a preset rule engine, enables bidirectional automated responses where hardware events trigger container actions and container events trigger hardware data acquisition, thus completing closed-loop collaborative management between OpenBMC and containerized applications.

[0017] Embodiments of this application provide a resource management method, such as... Figure 3As shown, the method is applied to a resource management system, which includes a server controller and a container orchestration platform. The controller and the container orchestration platform are interconnected. The method includes: Step S01: Set up a unified event data structure, wherein the unified event data structure includes one or more of the following: event source domain, event type, node identifier, timestamp, and additional data fields. Set up a unified event data structure, including: In response to the controller detecting a change in hardware state, the hardware state event information is converted into a unified event data structure and published to the corresponding hardware event platform; In response to the container orchestration platform detecting container lifecycle changes or resource anomalies, the container lifecycle change or resource anomaly event information is transformed into a unified event data structure and published to the corresponding container application event platform.

[0018] Specifically, a lightweight message middleware that supports the publish / subscribe pattern is deployed in the server host operating system. The message middleware provides low-latency and highly reliable message transmission services. Integrate a message client module into the OpenBMC firmware and configure it to establish a connection with the host-side message middleware via network protocol, forming a communication path from the controller BMC to the host; Deploy a collaborative management agent on the container orchestration system control node. The collaborative management agent registers as a subscriber to the message middleware and listens for hardware event topics published by OpenBMC and application event topics published by the container platform. Define a unified event data structure, adopt a serializable format, and include event source domain, event type, node identifier, timestamp, and additional data fields; When OpenBMC detects a change in hardware status, it encapsulates the event information according to a unified data structure and publishes it to the corresponding hardware event topic. When the container orchestration system detects a container lifecycle change or resource anomaly, it encapsulates the event information according to a unified data structure and publishes it to the corresponding application event topic. After receiving an event, the collaborative management agent module parses the event source and type, determines whether a cross-domain response action needs to be triggered, and realizes bidirectional, asynchronous, and real-time communication between OpenBMC and the container orchestration platform. Understandably, the message middleware runs in privileged container mode in the host operating system, is bound to the host network namespace, and receives encrypted data packets sent by the OpenBMC message client module via UDP or TCP protocol. The message client module is initialized during the OpenBMC startup phase, establishing a persistent connection channel based on the pre-configured host IP address and port number; After the connection is established, the message client module periodically sends a link keep-alive signal. If no response is received three times in a row, it switches to the backup communication path and triggers a local alarm log recording. The unified event data structure adopts a binary compact encoding format, with a fixed field arrangement order to ensure cross-platform parsing consistency; the event publishing topic adopts a hierarchical naming structure. The hardware event topic format is / event / hardware / <node_id> / <event_type> The application event topic format is / event / application / <cluster_id> / <namespace>The subscriber realizes selective receiving according to a wildcard matching mechanism; After receiving the event, the cooperative management agent first verifies the message integrity check code, and then judges the belonging execution domain according to the event source domain. If it is a hardware event, it enters the matching process of the rule response module. If it is an application event, it enters the health prediction correlation analysis process, forming the starting point of cross-domain processing driven by events.

[0019] Here, the communication mechanism of the bidirectional event-driven cooperative control bus is independent of the running state of the host operating system. OpenBMC can still normally publish hardware events when the host is down or the operating system is not started, ensuring that critical alarms are not lost. The message middleware supports message persistence and retry mechanism, which can guarantee that events can be finally reached when the network is temporarily interrupted, improving the robustness of the system. The design of the unified event model shields the differences between the underlying platforms, so that the scheme can adapt to server hardware and various container orchestration systems of different manufacturers, and has good universality and portability.

[0020] It can be understood that the message middleware configures a non-volatile storage area for caching high-priority hardware events. The event priority is determined according to the fault severity classification. Power failure and temperature overrun events are marked as first priority and are forced to be persistent. The message retransmission mechanism is executed by the message client module on the OpenBMC side. The retransmission interval uses an exponential backoff strategy, with an initial interval of 1 second and an upper limit of 32 seconds. Event deduplication is based on a three-tuple hash value of node identifier + event type + timestamp to prevent repeated responses due to network jitter. The communication path is independent of the host PCIe bus at the physical layer and is transmitted through a dedicated management network port or an out-of-band channel to ensure uninterrupted communication when the host system crashes.

[0021] Step S02, the controller receives the user's business intent instruction; The user's business intent is converted into a corresponding management policy type through an intent analysis model.

[0022] Step S021, according to the preset mapping relationship, the user's business intent is converted into a corresponding management policy type; According to the preset mapping relationship, the user's business intent is converted into a corresponding management policy type, including: Extracting intent keywords from the user's business intent through regular expressions or semantic templates; Creating a user's business intent-management policy type mapping table according to the user's business intent keywords; According to the user's business intent-management policy type mapping table, the user's business intent is converted into a corresponding management policy type; According to the classification algorithm, the business intention of the user is converted into a corresponding management strategy type: According to the classification algorithm, the business intention of the user is converted into a corresponding management strategy type, including: The business intention of the user is converted into a field feature vector; Obtain a classification algorithm model, wherein the classification algorithm model includes one or more of a logistic regression model, a random forest model, and a graph neural network model; The field feature vector is converted into a corresponding management strategy type by the classification algorithm model.

[0023] Step S03, obtaining a hardware strategy parameter according to the management strategy type, wherein the hardware strategy parameter includes one or more of a power supply working mode, a fan speed parameter, and a power limit parameter; According to the management strategy type, obtain a container scheduling strategy parameter, wherein the container scheduling strategy parameter includes one or more of a container instance number, a resource allocation limit parameter, and a container node deployment constraint parameter.

[0024] Step S04, constructing a unified optimization objective function according to the hardware strategy parameter and the container scheduling strategy parameter.

[0025] Step S041, obtaining a controller-side energy consumption value e, a container service delay l, and a hardware health risk index r; Obtain normalized weight coefficients a, b, and c; and a decision variable set x; The mathematical expression of the unified optimization objective function is: min x (a×e+b×l+c×r), wherein the unified optimization objective function is a multi-objective weighted optimization model.

[0026] Step S05, generating a controller configuration command and a container scheduling strategy request through the optimization solution of the unified optimization objective function.

[0027] Step S051, obtaining a controller-side energy consumption value e, a container service delay l, and a hardware health risk index r; Obtain normalized weight coefficients a, b, and c; Calculate the comprehensive cost function f by the formula: f=a×e+b×l+c×r, wherein the comprehensive cost function is the objective function to be minimized; Minimize the comprehensive cost function f through the unified optimization objective function to obtain a decision variable set x, wherein the decision variable set x includes a controller configuration command and a container scheduling strategy request.

[0028] Specifically, obtain a controller-side energy consumption value e, a container service delay l, and a hardware health risk index r; a weight coefficient a of an energy consumption value of a controller side, a weight coefficient b of a container service delay, a weight coefficient c of a hardware health risk index; a decision variable set x; wherein a+b+c=1; An integrated cost function f is calculated through a formula: f=a×e+b×l+c×r; A mathematical expression of a unified objective function is min x (f), wherein the unified objective function is a multi-objective weighted optimization model; The decision variable set x is obtained by minimizing the integrated cost function f through the unified objective function, wherein the decision variable set x includes a controller configuration command and a container scheduling strategy request.

[0029] It can be understood that a high-level business intent is received, and the business intent is a semantic description representing a management target; The business intent is input into an intent analysis model, and the model determines a corresponding management strategy category according to a preset mapping relationship or a classification algorithm; Hardware strategy parameters are generated according to the strategy category, and the hardware strategy parameters include a power supply working mode, a fan speed regulation strategy, and a power limitation strategy; Container scheduling strategy parameters are generated according to the strategy category, and the container scheduling strategy parameters include a container instance number, a resource allocation limitation, and a node deployment constraint condition; The hardware strategy parameters and the container scheduling strategy parameters are taken as inputs to construct a unified optimization objective function for jointly optimizing hardware and software resource configurations; Here, the unified optimization objective function is a multi-objective weighted optimization model, and a mathematical expression thereof is: min x (a×e+b×l+c×r); Wherein e represents a BMC side energy consumption index, the unit is watt, l represents a container service delay, the unit is millisecond, r represents a hardware health risk index, and a, b, and c are normalized weight coefficients; An optimization solving result of the unified optimization objective function is used to generate a BMC configuration command and a Kubernetes scheduling strategy update request: A current system power consumption value is obtained from OpenBMC and set as e, representing a hardware side resource consumption level; a response delay value of a key service is obtained from a container monitoring system and set as l, representing an application side service quality level; A hardware health risk index is obtained and set as r, representing a current reliability state of a server hardware, which is output by a prediction model; A weight distribution relationship is determined according to a current business intent, and weight coefficients corresponding to e, l, and r are set as a, b, and c, respectively; A comprehensive cost function is constructed, and the function expression is: f = a x e + b x l + c x r; Wherein, f is the target function to be minimized, reflecting the overall operating cost of the system; The decision variable set x that minimizes f is solved by an optimization algorithm, x includes BMC configuration parameters and container scheduling parameters; The solution is converted into BMC control instructions and container orchestration system configuration update requests to complete the collaborative resource configuration; It can be understood that the intent analysis model is constructed based on a pre-trained text classification model. After inputting the business intent string, the policy category label is output. The label is mapped to the entry in the policy parameter template library; The policy parameter template library is stored in the local configuration file, and each template defines a combination of a set of hardware policy parameters and container scheduling policy parameters; The weight coefficients a, b, and c are obtained from the preset weight table according to the policy category to ensure the semantic consistency of multi-objective optimization; The solution process of the comprehensive cost function is executed inside the collaborative management agent. The decision variable space is limited by physical constraints. The BMC configuration parameters cannot exceed the allowed range of device specifications, and the container scheduling parameters cannot exceed the total amount of available resources in the cluster; The solving algorithm uses a gradient descent method with constraints. In the discrete variable space, a heuristic search strategy is used to output a set of feasible solutions; The BMC control instructions are submitted to the controller OpenBMC through the Redfish RESTful interface, and the container orchestration system configuration update request is implemented by calling the Patch interface of the Kubernetes API Server, ensuring the atomicity and rollback of the configuration change.

[0030] Here, the intent analysis model supports dynamic expansion and hot loading of strategies. Users can add business intent types and their corresponding strategy mapping relationships through configuration files or management interfaces without modifying the core logic. The construction process of the unified optimization objective function is transparent and auditable. All weight coefficients and input parameters are recorded in the system log, supporting post-tracing and strategy tuning to ensure that the resource scheduling decision-making process meets the operation and maintenance specifications and security policy requirements.

[0031] Wherein, the strategy hot loading is realized by monitoring the file system events of the strategy configuration directory. After a new strategy file is written, the collaborative management agent verifies its digital signature and syntax structure, and loads it into the memory strategy table after confirming that it is correct; Each time the target function is constructed and solved, a structured operation log is generated. The log includes input parameters, weight values, solution time, and result status code. The log entries are broadcast to the centralized audit system through the collaborative control bus and form a traceable decision chain.

[0032] Step S052: Obtain the idle power consumption Piidle, the full load power consumption Pipeak, the overall utilization rate ui of node i, and the total number of nodes N; Through the formula: ; Calculate the energy consumption value e on the controller side; Obtain the latency of LBase under ideal, contention-free use, the overall utilization rate ui of node i, and the sensitivity coefficient k; The container service latency l is calculated using the formula: l=lbase× (1+(k-1)ui) / (1-ui); Obtain the overall utilization rate ui of node i and the single node risk index Ri, where Ri = P(failure in next24h | sensor data); Through the formula: r=Σ i ui×R i Calculate the hardware health risk index r.

[0033] Step S06: Optimize resource configuration between the controller and the container orchestration platform based on the controller configuration command and the container scheduling policy request.

[0034] Step S061: The controller sets up a lightweight time series prediction model, obtains the probability value of hardware failure through the lightweight time series prediction model, obtains a hardware health score based on the probability value of hardware failure, compares the hardware health score with a preset health threshold, generates a hardware risk warning event based on the comparison result, and sends the hardware risk warning event to the container orchestration platform. The container orchestration platform receives hardware risk warning events, performs predictive migration of containers on nodes affected by hardware risks based on these events, and dynamically adjusts the container resource configuration of the nodes.

[0035] Step S0611: Set up a lightweight time series prediction model in the controller firmware; The controller acquires hardware sensor data at a first interval threshold (5 minutes). The hardware sensor data includes at least one of temperature, voltage, fan speed, power status, and memory error count. Convert hardware sensor data into time-series input vectors; The probability value p of hardware failure is obtained from the time series input vector using a lightweight time series prediction model. The hardware health score H is calculated using the formula: H = 100 × (1 - p). The hardware health score is compared with a preset health threshold. If the number of times the hardware health score is less than the preset health threshold (60 points) reaches the second threshold (which is set according to the hardware type, monitoring frequency and false alarm tolerance, for example, when the hardware is a disk, the second threshold is 3 to 5 times), then it is determined that there is a hardware deterioration risk and a hardware risk warning event is generated. Transform hardware risk warning events into a unified event data structure; Hardware risk warning events with a unified event data structure are sent to the container orchestration platform.

[0036] Specifically, a lightweight time series forecasting model is deployed in the OpenBMC firmware, which is capable of running in resource-constrained environments; The server hardware sensor data is collected at fixed intervals. The hardware sensor data includes temperature, voltage, fan speed, power status, and memory error count. The collected server hardware sensor data is preprocessed to form a time series input vector, which is then input into a lightweight time series prediction model. The lightweight time series forecasting model outputs the probability of hardware failure within a future period. Based on this probability, a hardware health score is calculated, expressed as: H = 100 × (1 - p); Where p is the probability value of the predicted anomaly, and H is the hardware health score; Compare H with a preset health threshold. If H is consistently lower than the health threshold for a set number of times, it is determined that there is a risk of hardware degradation. Generate hardware risk warning events, encapsulate them into a unified event format, and publish them to the container orchestration system via the collaborative control bus; Here, the lightweight time series prediction model is a compressed deep neural network model with fewer than 500KB of parameters, which supports inference execution on the ARM Cortex-M or A-series processors of the OpenBMC controller. The sensor data acquisition cycle is 10 seconds, and data from 30 time points in the last 5 minutes are continuously collected to form the input sequence. Preprocessing includes linear interpolation of missing values, 3σ filtering of outliers, and feature normalization; The lightweight time series forecasting model outputs the probability value of a hardware failure occurring within the next 5 minutes; A risk warning event is triggered when the Hardware Health Score (HHS) falls below 0.6 for three consecutive times. Risk warning events include fields such as node identifier, fault prediction type, confidence level, and suggested response action, and are published to the / event / hardware / risk topic via the collaborative control bus.

[0037] Understandably, the lightweight time series prediction model uses historical operation and maintenance data for offline training during the training phase, and updates the model parameters regularly through an incremental learning mechanism to adapt to server aging trends and environmental changes. The calculation process of hardware health score considers the correlation of multi-dimensional sensor data to avoid misjudgment caused by fluctuations in a single indicator, which can improve prediction accuracy and system stability. The model inference process is completed locally at the controller BMC without relying on external computing resources, ensuring real-time performance and data security.

[0038] In step S0612, the collaborative management agent module of the container orchestration platform receives a hardware risk warning event with a unified event data structure. Extract the identifier information of nodes affected by hardware risks based on hardware risk warning events; Set the identified nodes to a maintenance ready state to prevent new container instances from being scheduled to that node; Obtain the preset migration priority strategy; The execution order of container migration for identified nodes is determined according to a preset migration priority strategy; Predictive migration is performed on the containers of the identified nodes according to the container migration execution order; Create new container instances on other unmarked nodes and monitor the resource load status of the source node, dynamically adjusting the resource allocation limit of containers on non-critical nodes based on the resource load status of the source node.

[0039] Specifically, the collaborative management agent of the container orchestration system receives hardware risk warning events and extracts the identification information of nodes affected by hardware risks. Set a maintenance standby state for nodes affected by hardware risks to prevent new container instances from being scheduled to those nodes. The execution order of containers to be migrated is determined based on the preset migration priority strategy. Perform graceful eviction for each container to be migrated to ensure a smooth termination of service connections; Recreate the evicted container instance on other healthy nodes to complete the service migration; During the migration process, the resource load status of the source node is monitored, and the resource allocation limit of non-critical containers is dynamically adjusted to achieve adaptive adjustment of service quality. Here, maintaining the ready state is achieved by registering node taints with the container orchestration system. The taint key is node-health / unstable:NoSchedule, which prevents new node Pods from being scheduled. The migration priority strategy is determined based on the priority.class field in the container tag, with higher priority services being migrated first; Graceful eviction is achieved by sending a SIGTERM signal to the Pod and waiting for a preset grace period, during which new connections are denied. When a new container instance is created, it inherits the tags, annotations, resource requests, and limits of the original Pod. The resource allocation limit is adjusted by modifying the CPU quota and memory limit of the container cgroup, and the adjustment range is dynamically calculated based on the current node load.

[0040] Meanwhile, the predictive migration process follows the standard interface specifications of the container orchestration platform, and the graceful eviction operation is compatible with the Pod eviction policies of mainstream systems such as Kubernetes, ensuring a smooth switch of application services. The QoS adjustment policy distinguishes between critical and non-critical loads based on container tags or namespaces to avoid impacting core business. The migration process supports batch execution and rate limiting to prevent large-scale concurrent migration from causing network congestion or target node overload.

[0041] Step S062: Implement bidirectional automated response to hardware events and container events through a preset rule engine, and manage resource configuration between the controller and the container orchestration platform; A pre-defined rules engine enables bidirectional automated response to hardware and container events, and manages resource configuration between the controller and container orchestration platform, including: Configure the rule engine module in the collaborative management agent of the container orchestration platform; When a hardware failure or resource shortage event is received, the corresponding container event is triggered through the rules engine module. Container events include limiting container resources, degrading resource services, or suspending them. When a container resource anomaly event is received, the corresponding hardware event is triggered through the rule engine module. The hardware event includes collecting hardware sensor snapshot data at the current moment through the controller. The collected hardware sensor snapshot data is time-aligned and stored in association with the container event log; Generate audit logs for the triggering and execution process of each hardware event and container event.

[0042] Specifically, a rules engine module is deployed in the collaborative management agent to load a pre-defined set of cross-domain response rules; When an event indicating hardware failure or resource shortage is received, the rules engine triggers the corresponding container action, which may include resource limiting, service degradation, or suspension. When an event indicating an abnormality in container resources is received, the rules engine triggers OpenBMC to collect sensor hardware snapshot data at the current moment; The collected hardware snapshot data is time-aligned and correlated with the container event logs for subsequent root cause analysis. An audit log is generated for the triggering and execution of each rule and published through the collaborative control bus to ensure that the operation is traceable. Here, the rule engine uses the Rete algorithm to achieve efficient pattern matching. The rule set is stored in JSON format, and each rule contains a condition expression and a list of actions. Conditional expressions support multi-event combination logic, and action lists support cross-domain call instructions; Sensor snapshot data acquisition is achieved by sending synchronization requests to OpenBMC, and the acquisition timestamp is aligned with the container event timestamp to the millisecond level; The associated storage is implemented using a time-series database, establishing an index mapping between event IDs and data blocks; Audit logs include rule ID, trigger time, matching conditions, execution action, and execution result, and are published via the / event / audit topic.

[0043] Finally, the rules engine supports user-defined cross-domain response rules. Rule entries can be flexibly added, modified, or disabled through the configuration interface to meet the operational needs of different scenarios. The rule matching and execution process has priority control and conflict detection mechanisms to prevent multiple rules from being triggered simultaneously and causing abnormal system status. The audit log includes rule number, triggering condition, execution action, and timestamp, and supports integration with centralized log systems to provide complete evidence for security auditing and fault analysis.

[0044] Here, as Figure 4 As shown, the technical solution of this application includes: building a containerized runtime environment on a server where OpenBMC is deployed, and establishing a bidirectional event-driven collaborative control bus between OpenBMC and the container orchestration platform to achieve real-time communication between hardware status and container events; The system receives user input business intents, transforms them into a unified optimization objective function that includes hardware policies and container scheduling policies through an intent parsing model, and dynamically adjusts BMC configuration and container deployment based on this unified optimization objective function. A lightweight time series forecasting model runs on the OpenBMC side, which calculates the hardware health score (HHS) in real time based on the collected hardware sensor data. When the hardware health score (HHS) is lower than the health threshold, a risk warning event is generated and the container orchestration system is notified via the cooperative control bus. The container orchestration system responds to risk warning events, performs predictive migration of containers on affected nodes, and dynamically adjusts container QoS during the migration process; The collaborative control bus, through a preset rule engine, enables bidirectional automated responses where hardware events trigger container actions and container events trigger hardware data acquisition, thus completing closed-loop collaborative management between OpenBMC and containerized applications.

[0045] In addition, the resource load status of the source node is monitored, and the resource allocation limit of non-critical node containers is dynamically adjusted based on the resource load status of the source node, including: The collaborative management agent periodically (e.g., every 5 seconds) obtains the resource load status indicators of the source node, which include node CPU utilization, node memory pressure, and container actual CPU usage vslimit. The overall load score is calculated by weighting the resource load status indicators of the source nodes. The resource allocation limit for non-critical containers is dynamically adjusted based on the overall load score. The resource allocation cap for non-critical containers is dynamically adjusted based on the overall load score, including: Retrieve non-critical containers by tag; Modify the controller to which a non-critical container belongs, trigger a rolling update, and dynamically adjust the resource allocation limit of the non-critical container; Alternatively, deploy a lightweight agent on each node; The lightweight agent listens for instructions from the collaborative management agent, calls the cgroup interface of the container on the node, modifies the CPU quota or memory limit of the cgroup interface, and dynamically adjusts the resource allocation limit of non-critical containers. Dynamically adjusting the resource allocation cap for non-critical containers based on the overall load score also includes: Built-in degradation switches in non-critical container services include gRPC rate limiting and log level adjustment; The collaborative management agent sends a "load reduction signal" via the Sidecar or API; By reducing resource consumption (such as reducing the number of threads and cache size), the resource allocation limit of non-critical containers can be dynamically adjusted. Alternatively, enable the "Off" mode for the Vertical Pod Autoscaler (VPA); The recommended value of VPA is read through the collaborative management agent, and an update is proactively triggered based on the current resource load. Extend the VPA controller to include a "migration mode" strategy to dynamically adjust the resource allocation cap for non-critical containers.

[0046] By monitoring the resource load status of the source node and dynamically adjusting the resource allocation limit of non-critical containers accordingly, the stability and success rate of the migration process can be improved; adaptive hierarchical guarantee of Quality of Service (QoS) can be achieved; the overall resource utilization efficiency of the cluster can be optimized, the waste caused by "over-reservation" can be avoided, the impact of migration on the overall resource pool of the cluster can be reduced, and the resource turnover efficiency can be improved.

[0047] The resource management method provided in this application embodiment can be improved and optimized in several ways without departing from the technical solution of this application, and these improvements and optimizations should also be considered within the scope of protection of this application.

[0048] The technical solution of this application can also be applied to similar scenarios in open-source firmware projects (OPENBIOS).

[0049] The beneficial effects of the technical solutions provided in this application are: This application's technical solution achieves real-time, bidirectional state synchronization and event linkage between the controller and containerized applications by constructing an event-driven collaborative control bus. It also integrates intent-driven unified resource orchestration and a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability. By using a hardware health scoring model to identify potential hardware failures in advance and trigger predictive container migration, it effectively avoids service interruptions caused by hardware degradation and greatly enhances system availability. At the same time, the rule engine-driven cross-domain automated response mechanism realizes closed-loop management of hardware events triggering container scheduling actions and container anomalies triggering hardware data collection. This breaks down the siloed management between traditional controllers and cloud-native platforms, providing solid support for intelligent operation and maintenance and root cause analysis. In short, the overall technical solution has a high degree of integration, foresight, and engineering feasibility.

[0050] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0051] Embodiments of this application also provide an electronic device, which includes a server controller and a container orchestration platform, wherein the controller is interconnected with the container orchestration platform, and the controller is used for: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0052] The controller is used to: obtain the controller-side energy consumption value e, container service latency l, and hardware health risk index r; The weighting coefficients are: a for obtaining the controller-side energy consumption value, b for obtaining the container service latency value, and c for obtaining the hardware health risk index value; and a set of decision variables x, where a+b+c=1. The comprehensive cost function f is calculated using the formula: f=a×e+b×l+c×r; The mathematical expression for the unified objective function is: min x (f), where the unified objective function is a multi-objective weighted optimization model; The decision variable set x is obtained by minimizing the comprehensive cost function f through the unified objective function, wherein the decision variable set x includes controller configuration commands and container scheduling policy requests.

[0053] The controller is used to: obtain the node's idle power consumption Piidle, full load power consumption Pipeak, the overall utilization rate ui of node i, and the total number of nodes N; Through the formula: ; Calculate the energy consumption value e on the controller side; Obtain the latency of LBase under contention-free conditions, the overall utilization rate ui of node i, and the sensitivity coefficient k; The container service latency l is calculated using the formula: l=lbase× (1+(k-1)ui) / (1-ui); Obtain the overall utilization rate ui of node i and the single node risk index Ri, where Ri = P(failure in next24h | sensor data); Through the formula: r=Σ i ui×R i Calculate the hardware health risk index r.

[0054] The controller is used to: set up a lightweight time series prediction model, obtain a probability value of hardware failure through the lightweight time series prediction model, obtain a hardware health score based on the probability value of hardware failure, compare the hardware health score with a preset health threshold, generate a hardware risk warning event based on the comparison result, and send the hardware risk warning event to the container orchestration platform. The container orchestration platform is used to: receive the hardware risk warning event, perform predictive migration of containers on nodes affected by hardware risks based on the hardware risk warning event, and dynamically adjust the node container resource configuration.

[0055] The controller is used to: set up a lightweight time series prediction model in the firmware of the controller; The controller acquires hardware sensor data at a first interval threshold time, and the hardware sensor data includes at least one of temperature, voltage, fan speed, power status, and memory error count. The hardware sensor data is converted into a time-series input vector; The lightweight time series prediction model obtains the probability value p of hardware failure based on the time series input vector. The hardware health score H is calculated using the formula: H = 100 × (1 - p). The hardware health score is compared with a preset health threshold; If the number of times the hardware health score is less than the preset health threshold reaches a second threshold, then it is determined that there is a risk of hardware degradation, and a hardware risk warning event is generated. Transform the hardware risk warning events into a unified event data structure; Hardware risk warning events with a unified event data structure are sent to the container orchestration platform.

[0056] The container orchestration platform is used to receive hardware risk warning events based on the unified event data structure. Extract the identifier information of nodes affected by hardware risks based on the hardware risk warning events; Set the identified nodes to a maintenance ready state to prevent new container instances from being scheduled to that node; Obtain the preset migration priority strategy; The execution order of container migration for the identified nodes is determined according to the preset migration priority strategy; Predictive migration is performed on the containers of the identified nodes according to the container migration execution order; Create new container instances on other unmarked nodes and monitor the resource load status of the source node. Dynamically adjust the resource allocation limit of containers on non-critical nodes based on the resource load status of the source node.

[0057] Specifically, such as Figure 5 As shown, an OpenBMC and containerized application collaborative management system is provided, including: Event bus module, intent parsing module, health prediction module, migration execution module, and rule response module; The event bus module is used to establish a bidirectional event-driven communication channel between the server deploying OpenBMC and the container orchestration system, enabling real-time communication between hardware status and container events, and supporting the encapsulation, publication, and subscription of a unified event model. The intent parsing module is used to receive the business intent input by the user, convert it into hardware policy parameters and container scheduling policy parameters through a preset mapping relationship or classification algorithm, and construct a unified objective function for joint optimization. The health prediction module is used to run a lightweight time series prediction model on the OpenBMC side, calculate the hardware health score based on periodically collected hardware sensor data, and generate a risk warning event when the hardware health score is continuously lower than the set health threshold and send it to the container orchestration system. The migration execution module is used to perform predictive migration of containers on affected nodes after receiving a hardware risk warning. This includes setting a maintenance ready state, evicting containers according to priority, rebuilding container instances on healthy nodes, and dynamically adjusting the resource quotas of non-critical containers to adapt to the current load during the migration process. The rule response module is used to achieve cross-domain automated response through a preset rule engine. When a hardware event is detected, it triggers container resource regulation actions. When an abnormal container event is detected, it triggers hardware sensor data acquisition and records the operation process as an audit log to support traceability analysis.

[0058] The beneficial effects of the technical solutions provided in this application are: This application's technical solution achieves real-time, bidirectional state synchronization and event linkage between the controller and containerized applications by constructing an event-driven collaborative control bus. It also integrates intent-driven unified resource orchestration and a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability. By using a hardware health scoring model to identify potential hardware failures in advance and trigger predictive container migration, it effectively avoids service interruptions caused by hardware degradation and greatly enhances system availability. At the same time, the rule engine-driven cross-domain automated response mechanism realizes closed-loop management of hardware events triggering container scheduling actions and container anomalies triggering hardware data collection. This breaks down the siloed management between traditional controllers and cloud-native platforms, providing solid support for intelligent operation and maintenance and root cause analysis. In short, the overall technical solution has a high degree of integration, foresight, and engineering feasibility.

[0059] For a description of the features in the corresponding embodiments of the resource management system, please refer to the relevant descriptions in the corresponding embodiments of the resource management method, which will not be repeated here.

[0060] Embodiments of this application also provide a resource management system, including a server controller and a container orchestration platform, wherein the controller and the container orchestration platform are interconnected. The controller is used to receive user business intent instructions; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0061] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in the resource management method embodiment, the method including: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0062] like Figure 6 As shown, embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in the resource management method embodiments at runtime, the method including: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0063] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0064] Embodiments of this application also provide a computer program product, which includes a computer program. When executed by a processor, the computer program implements the steps in the resource management method embodiments, the method including: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0065] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps in the resource management method embodiments, the method including: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

[0066] This application's technical solution achieves real-time, bidirectional state synchronization and event linkage between the controller and containerized applications by constructing an event-driven collaborative control bus. It also integrates intent-driven unified resource orchestration and a proactive migration mechanism based on hardware health prediction, significantly improving the system's collaborative optimization capabilities in energy efficiency, performance, and reliability. By using a hardware health scoring model to identify potential hardware failures in advance and trigger predictive container migration, it effectively avoids service interruptions caused by hardware degradation and greatly enhances system availability. At the same time, the rule engine-driven cross-domain automated response mechanism realizes closed-loop management of hardware events triggering container scheduling actions and container anomalies triggering hardware data collection. This breaks down the siloed management between traditional controllers and cloud-native platforms, providing solid support for intelligent operation and maintenance and root cause analysis. In short, the overall technical solution has a high degree of integration, foresight, and engineering feasibility.

[0067] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0068] The resource management method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.< / namespace>

Claims

1. A resource management method, characterized in that, The method is applied to a resource management system, the system including a server controller and a container orchestration platform, the controller and the container orchestration platform being interconnected, and the method comprising: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

2. The resource management method according to claim 1, characterized in that, Before receiving the user's business intent instruction, the process includes: A unified event data structure is set up, wherein the unified event data structure includes one or more of the following: event source domain, event type, node identifier, timestamp, and additional data fields; The setting of a unified event data structure includes: In response to the controller detecting a change in hardware status, the hardware status event information is converted into a unified event data structure and published to the corresponding hardware event platform. In response to the container orchestration platform detecting a container lifecycle change or resource anomaly, the container lifecycle change or resource anomaly event information is transformed into a unified event data structure and published to the corresponding container application event platform.

3. The resource management method according to claim 1, characterized in that, The process of converting the user's business intent into a corresponding management strategy type through an intent parsing model includes: The user's business intent is converted into a corresponding management strategy type based on a preset mapping relationship; The step of converting the user's business intent into a corresponding management strategy type according to a preset mapping relationship includes: Extract intent keywords from the user's business intent using regular expressions or semantic templates; Create a user's business intent-management strategy type mapping table based on the user's business intent keywords; Based on the user's business intent-management strategy type mapping table, the user's business intent is converted into the corresponding management strategy type; Based on the classification algorithm, the user's business intent is transformed into a corresponding management strategy type: The step of converting the user's business intent into a corresponding management strategy type based on a classification algorithm includes: Transform the user's business intent into a field feature vector; Obtain a classification algorithm model, wherein the classification algorithm model includes one or more of a logistic regression model, a random forest model, and a graph neural network model; The classification algorithm model transforms the field feature vector into the corresponding management strategy type.

4. The resource management method according to claim 1, characterized in that, The step of constructing a unified objective function based on the hardware policy parameters and container scheduling policy parameters, and generating controller configuration commands and container scheduling policy requests based on the optimization results of the unified objective function, includes: Obtain the controller-side energy consumption value e, container service latency l, and hardware health risk index r; The weighting coefficients are: a for obtaining the controller-side energy consumption value, b for obtaining the container service latency value, and c for obtaining the hardware health risk index value; and a set of decision variables x, where a+b+c=1. The comprehensive cost function f is calculated using the formula: f=a×e+b×l+c×r; The mathematical expression for the unified objective function is: min x (f), where the unified objective function is a multi-objective weighted optimization model; The decision variable set x is obtained by minimizing the comprehensive cost function f through the unified objective function, wherein the decision variable set x includes controller configuration commands and container scheduling policy requests.

5. The resource management method according to claim 4, characterized in that, The acquisition of controller-side energy consumption values, container service latency, and hardware health risk indices includes: Get the idle power consumption Piidle, the full load power consumption Pipeak, the overall utilization rate ui of node i, and the total number of nodes N; Through the formula: ; Calculate the energy consumption value e on the controller side; Obtain the latency of LBase under contention-free conditions, the overall utilization rate ui of node i, and the sensitivity coefficient k; The container service latency l is calculated using the formula: l=lbase×(1+(k-1)ui) / (1-ui); Obtain the overall utilization rate ui of node i and the single node risk index Ri, where Ri = P(failure in next 24h | sensor data); Through the formula: r=Σ i ui×R i Calculate the hardware health risk index r.

6. The resource management method according to claim 2, characterized in that, The step of configuring resources between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request includes: The controller is configured with a lightweight time series prediction model. The probability value of hardware failure is obtained through the lightweight time series prediction model. A hardware health score is obtained based on the probability value of hardware failure. The hardware health score is compared with a preset health threshold. A hardware risk warning event is generated based on the comparison result and the hardware risk warning event is sent to the container orchestration platform. The container orchestration platform receives the hardware risk warning event, performs predictive migration of containers on nodes affected by hardware risks based on the hardware risk warning event, and dynamically adjusts the node container resource configuration.

7. The resource management method according to claim 6, characterized in that, The process includes setting a lightweight time series prediction model, obtaining a probability value for hardware anomalies using the lightweight time series prediction model, obtaining a hardware health score based on the probability value of hardware anomalies, comparing the hardware health score with a preset health threshold, generating a hardware risk warning event based on the comparison result, and sending the hardware risk warning event to the container orchestration platform. A lightweight time series prediction model is set in the firmware of the controller; The controller acquires hardware sensor data at a first interval threshold time, and the hardware sensor data includes at least one of temperature, voltage, fan speed, power status, and memory error count. The hardware sensor data is converted into a time-series input vector; The lightweight time series prediction model obtains the probability value p of hardware failure based on the time series input vector. The hardware health score H is calculated using the formula: H = 100 × (1 - p). The hardware health score is compared with a preset health threshold; If the number of times the hardware health score is less than the preset health threshold reaches a second threshold, then it is determined that there is a risk of hardware degradation, and a hardware risk warning event is generated. Transform the hardware risk warning events into a unified event data structure; Hardware risk warning events with a unified event data structure are sent to the container orchestration platform.

8. The resource management method according to claim 6, characterized in that, The container orchestration platform includes a collaborative management agent module. Upon receiving the hardware risk warning event, the agent performs predictive migration of containers on nodes affected by hardware risks based on the warning event and dynamically adjusts the node container resource configuration, including: The collaborative management agent module of the container orchestration platform receives the hardware risk warning event of the unified event data structure; Extract the identifier information of nodes affected by hardware risks based on the hardware risk warning events; Set the identified nodes to a maintenance ready state to prevent new container instances from being scheduled to that node; Obtain the preset migration priority strategy; The execution order of container migration for the identified nodes is determined according to the preset migration priority strategy; Predictive migration is performed on the containers of the identified nodes according to the container migration execution order; Create new container instances on other unmarked nodes and monitor the resource load status of the source node. Dynamically adjust the resource allocation limit of containers on non-critical nodes based on the resource load status of the source node.

9. The resource management method according to claim 2, characterized in that, The step of configuring resources between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request also includes: A pre-defined rules engine enables bidirectional automated response to hardware and container events, and resource configuration management is performed between the controller and the container orchestration platform. The method of achieving bidirectional automated response to hardware and container events through a preset rule engine, and managing resource configuration between the controller and the container orchestration platform, includes: Configure the rule engine module in the collaborative management agent module of the container orchestration platform; When a hardware failure or resource shortage event is received, the corresponding container event is triggered through the rule engine module. The container event includes limiting container resources, degrading resource services, or suspending them. When a container resource anomaly event is received, the corresponding hardware event is triggered through the rule engine module. The hardware event includes collecting hardware sensor snapshot data at the current moment through the controller. The collected hardware sensor snapshot data is time-aligned and stored in association with the container event log; Generate audit logs for the triggering and execution process of each hardware event and container event.

10. An electronic device, characterized in that, The electronic device includes a server controller and a container orchestration platform, the controller being interconnected with the container orchestration platform, and the controller being used for: The controller receives the user's service intent command; The intent parsing model is used to convert the user's business intent into a corresponding management strategy type; Hardware policy parameters are obtained based on the management policy type, wherein the hardware policy parameters include one or more of power operating mode parameters, fan speed control parameters, and power limiting parameters; The container scheduling policy parameters are obtained according to the management policy type, wherein the container scheduling policy parameters include one or more of the following: number of container instances, resource allocation limit parameters, and container node deployment constraint parameters. A unified objective function is constructed based on the hardware policy parameters and container scheduling policy parameters; The controller configuration command and container scheduling policy request are generated based on the optimized solution of the unified objective function; Resource configuration is performed between the controller and the container orchestration platform according to the controller configuration command and container scheduling policy request.

Citation Information

Cited By

  • Server task management system, method and equipment based on BMC (Baseboard Management Controller) collaboration and medium

    CN121478362A