Bare metal starting optimization scheduling method under distributed architecture based on machine learning
By migrating management functions to network cards under a distributed architecture and optimizing scheduling strategies using the Q-Leaning model, the problem of too long startup time of bare metal servers is solved, and the startup efficiency and computing resource utilization are improved.
Patent Information
- Application Number
- CN202510028453.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
The startup time of bare metal servers is too long, especially in the case of high load and high concurrent task scheduling, resulting in waste of computing resources and system performance bottlenecks.
The bare metal startup optimization scheduling method is adopted under a distributed architecture based on machine learning. By migrating management functions to the network card, using the Q-Leaning model to dynamically adjust the scheduling strategy, reducing the latency of cloud disk acquisition, and optimizing the data loading path.
It significantly reduces the startup delay of bare metal servers, improves the rapid utilization efficiency of computing resources, avoids resource waste and performance bottlenecks, and meets the requirements of low latency and high efficiency of big data and high-performance computing tasks.
Smart Images

Figure CN119938330A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud computing and task scheduling algorithms, and specifically relates to a bare metal startup optimization scheduling method in a distributed architecture based on machine learning. Background Art
[0002] With the rapid development of big data and cloud computing, enterprises and users have a growing demand for efficient data processing and storage. Bare metal servers, with their advantage of direct access to hardware resources and avoiding the overhead of the virtualization layer, have become an ideal choice for high-performance computing, big data analysis and other fields. Bare metal servers can provide higher computing power and lower latency, and are particularly suitable for tasks that require high computing performance and response speed. However, despite the obvious performance advantages of bare metal servers, the long startup time of bare metal servers still significantly affects their overall performance, especially in the case of high load and high concurrent task scheduling. Long startup time will lead to a waste of computing resources and become a bottleneck for system performance.
[0003] The boot process of a bare metal server usually involves loading the operating system and related application data. If you rely on cloud storage to load this data, storage access latency or disk read and write speed limitations may significantly extend the boot time. Especially when cloud storage access is limited or network bandwidth is insufficient, the bare metal server needs to download the complete operating system image from the cloud and establish a connection with the cloud. This process is very time-consuming, resulting in the server being unable to start computing work in the shortest time, thereby wasting efficient computing resources. In addition, long startup delays may cause significant performance losses for real-time data processing and large-scale computing tasks that require fast responses, affecting the overall throughput of the platform and the response speed of tasks.
[0004] Currently, the startup process of bare metal servers mostly relies on the traditional management model, in which all management functions run on the CPU. During the startup process, the bare metal server needs to be configured through a simple operating system and download the complete operating system image through the cloud. This reliance on the configured operating system and the connection method with the cloud is simple but inefficient, especially when the access latency of cloud storage is large, the startup time is often too long. This design not only increases the startup latency, but also further aggravates the performance bottleneck in the case of high concurrency, making it difficult to meet the low latency and high efficiency requirements of big data and high-performance computing tasks. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a bare metal startup optimization scheduling method in a distributed architecture based on machine learning, which reconstructs the process of bare metal, reduces the delay of cloud disk acquisition, and improves the startup efficiency of bare metal.
[0006] In order to achieve the above object, the present invention is achieved through the following technical solutions:
[0007] The present invention is a bare metal startup optimization scheduling method under a distributed architecture based on machine learning. The bare metal startup optimization scheduling method under a distributed architecture is implemented by a service node scheduling model based on machine learning, and includes the following steps:
[0008] Step 1: Establish a distributed bare metal architecture, migrate bare metal management functions to the network card, and collect host deployment information through the bare metal management platform OpenStack;
[0009] Step 2: Initialize the modules and related parameters in the service node scheduling model, and use the gradient descent algorithm to process the deployment information collected in step 1;
[0010] Step 3: The service node scheduling model combines the network communication delay between the host and the service node, and uses the ARP protocol to find the set of available service nodes;
[0011] Step 4: Update the performance prediction value of each host node in real time through the Q-Leaning model, and dynamically adjust the scheduling strategy according to the network status and service node scheduling model. In addition, adopt a continuous learning mechanism to regularly train and update the Q-Leaning model to generate an optimized scheduling plan.
[0012] Step 5: According to the optimized scheduling plan generated by the Q-Leaning model, the host obtains the startup image from the service node with the startup image after model training and completes the startup process. The relevant data of each startup completion is collected and fed back to the bare metal management platform OpenStack.
[0013] A further improvement of the present invention is that in step 1, the established distributed bare metal architecture includes multiple IronConductors, HTTP servers, TFTP servers, hosts, and DHCP servers. Ironic Conductor obtains information about host deployment, the host configures the network and provides a TFTP server address, the host downloads a lightweight service from the TFTP server and starts the proxy service, the host selects the optimal Ironic Conductor to establish a connection based on the network status, the host obtains a complete image from the optimal HTTP server and restarts, and the management function for bare metal is migrated to a network card connected via PCIe. The network card is a remote hard disk, and the deployment information of the host is collected, including the network status between the host and the server and the historical startup data of each host. The bare metal does not need to download the operating system.
[0014] A further improvement of the present invention is that in step 1, the management function for bare metal is migrated to the network card connected via PCIe, the network card is a remote hard disk, and the deployment information of the host is collected, including the network status between the host and the server and the historical startup data of each host. The bare metal no longer needs to download the operating system.
[0015] A further improvement of the present invention is that in step 3, the network communication delay calculation formula between the host and the service node is:
[0016]
[0017] Among them, L b is the fixed delay for communication startup, d i,j is the amount of data transmitted between the host and the service node, B t The bandwidth between the host and the service node is used to perform network perception and discover adjacent service nodes through the ARP protocol.
[0018] A further improvement of the present invention is that in step 4, the Q-Leaning model is regularly trained and updated, and the updating process is expressed by the formula:
[0019] R(S t , α t )=ω1·(-L t )+ω2·B t +ω3·P+ω4·(1-R t )
[0020] ω1+ω2+ω3+ω4=1
[0021] Among them, ω1, ω2, ω3 and ω4 represent the importance weights of delay, bandwidth, historical selection times ratio and node load status. If the startup time is sensitive, the weight of ω1 is increased, emphasizing the importance of delay. If high bandwidth is required to support data transmission, the weight of ω2 is increased. The weights of ω3 and ω4 are adjusted according to the requirements for quality communication and node load. R(S t ,α t ) is used to evaluate the current action α t The larger the reward value, the better the currently selected service node is in terms of delay, bandwidth, packet loss rate and node load status, and it is suitable for priority selection. t Indicates the network delay between the current host and the service node. The lower the delay, the better the performance, so it is represented by a negative value in the reward. t Indicates the bandwidth between the current host and the service node. The higher the bandwidth, the stronger the network throughput and the greater the reward. P indicates the historical selection ratio of the current service node. The more historical selections, the higher the ratio. tIndicates the load status of the service node, indicating the resource usage of the node. The lower the load, the more available resources, so 1-R t Represents reward. The regression coefficient is used here to quantify the impact of different indicators on the overall performance of the system.
[0022] A further improvement of the present invention is that in step 4, the performance prediction is achieved by updating the Q value according to the following formula:
[0023] Q(S,α)=Q(S,α)+α[R(S,α)+γ·maxQ(s′,a′)-Q(S,α)]
[0024] Among them, Q is a two-dimensional array. The rows in the two-dimensional array represent all selectable states S, and the columns in the two-dimensional array represent all possible actions A, that is, candidate service nodes. The ∈-greedy strategy is adopted to select the next action. The action is randomly selected with probability ∈, and new service nodes are tried. The action selection, execution, observation of new states, and update of Q value are repeated until Q converges.
[0025] The further improvement of the present invention is that: a distributed architecture design module: a distributed bare metal architecture is established;
[0026] System initialization module: initializes the parameters in the service node scheduling model and the distributed bare metal architecture;
[0027] Network perception module: performs network perception through ARP protocol and searches for adjacent service nodes;
[0028] Action selection and scheduling decision module: Based on the ∈-greedy strategy of Q-Leaning, a service node with the largest Q value in the Q table is selected for scheduling, and the Q value is updated after execution;
[0029] Reward calculation module: Combined with the formula R(S t , α t ) Calculate the reward after service node selection;
[0030] Q value update module: updates the Q value in the Q table according to the formula Q(S,α);
[0031] Scheduling strategy evaluation module: regularly evaluates the effectiveness of scheduling strategies, checks whether the Q table converges, and adjusts strategies;
[0032] Service node selection and scheduling execution module: selects the best service node to perform bare metal startup, collects startup information, and records the selected node.
[0033] The beneficial effects of the present invention are:
[0034] The present invention migrates the bare metal management function to the hardware layer, especially transfers part of the management tasks to hardware devices such as network cards, significantly reduces the management task overhead that relies on CPU processing in the traditional mode, reduces the delay in the startup process, thereby speeding up the server startup speed and improving the rapid utilization efficiency of computing resources.
[0035] The present invention combines machine learning technology to dynamically optimize the data loading path based on real-time information such as system load, network bandwidth, storage performance, etc., and reduce the delay in the cloud disk access process. This intelligent optimization solution can adjust the storage access strategy in real time according to task requirements and system status, further improving startup efficiency.
[0036] The present invention can maintain the stability of the startup process in high-concurrency and high-load scenarios, avoid resource waste and performance bottlenecks caused by traditional startup methods, ensure that computing resources can be put into use in a timely manner, and improve system response speed and platform throughput. The experimental results of the present invention prove that service nodes with higher service efficiency can be preferentially selected. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart of the bare metal startup optimization scheduling method under the distributed architecture of the present invention.
[0038] Figure 2 It is a structural schematic diagram of a distributed bare metal architecture in the bare metal startup optimization scheduling method under the distributed architecture of the present invention.
[0039] Figure 3 This is a diagram of experimental results of the bare metal startup optimization scheduling method under the distributed architecture of the present invention. DETAILED DESCRIPTION
[0040] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0041] The present invention is a bare metal startup optimization scheduling method in a distributed architecture based on machine learning. The method is implemented by a service node scheduling model based on machine learning, and the model includes:
[0042] Distributed architecture design module: Established distributed bare metal architecture;
[0043] System initialization module: Initializes the parameters in the service node scheduling model and the distributed bare metal architecture. It initializes all necessary parameters and data structures of the system, including Q table initialization, learning parameters such as learning rate, discount factor, and network status between the host and service node such as latency, bandwidth, etc.
[0044] Network perception module: performs network perception through ARP protocol and searches for adjacent service nodes;
[0045] Action selection and scheduling decision module: Based on the ∈-greedy strategy of Q-Leaning, a service node with the largest Q value in the Q table is selected for scheduling, and the Q value is updated after execution;
[0046] Reward calculation module: Combined with the formula R(S t , α t ) Calculate the rewards after the service node is selected, mainly considering factors such as delay and bandwidth. The reward calculation module includes the delay calculation submodule and the bandwidth calculation submodule.
[0047] Q value update module: Update the Q value in the Q table according to Q(S, α); use the reinforcement learning formula to update the current state action;
[0048] Scheduling strategy evaluation module: Regularly evaluate the effectiveness of the scheduling strategy, check whether the Q table converges, and adjust the strategy; here, it decides whether to stop the training process by judging whether the Q table converges.
[0049] Service node selection and scheduling execution module: Executes scheduling decisions based on the final learned strategy, that is, selects the best service node to perform bare metal startup, collects startup information, and records the selected node.
[0050] like Figure 1 As shown, the present invention is a bare metal startup optimization scheduling method in a distributed architecture based on machine learning, and the scheduling method specifically includes the following steps:
[0051] Step 1: Establish a distributed bare metal architecture and migrate the bare metal management functions to the network card connected via PCIe. The network card is a remote hard disk. The deployment information of the host is collected, including the network status between the host and the server and the historical startup data of each host. These data are used as input for subsequent scheduling strategy optimization and machine learning models to reflect the current network status. Bare metal no longer needs to download the operating system.
[0052] like Figure 2As shown, the distributed bare metal architecture of the present invention includes multiple management schedulers Ironic Conductor, HTTP server, TFTP server, host, DHCP server, Ironic Conductor obtains the host deployment information, the host configures the network and provides the TFTP server address, the host downloads the lightweight service from the TFTP server and starts the proxy service, the host selects the best Ironic Conductor to establish a connection according to the network status, the host obtains the complete image from the best HTTP server and restarts, and migrates the management function for bare metal to the network card connected via PCIe. The network card is a remote hard disk, and the deployment information of the host is collected, including the network status between the host and the server and the historical startup data of each host. The bare metal does not need to download the operating system.
[0053] Step 2: Build a service node scheduling model based on the Q-Leaning model, initialize the modules and related parameters in the service node scheduling model, and use the gradient descent algorithm to process the deployment information collected in step 1;
[0054] Step 3: The service node scheduling model combines the network communication delay between the host and the service node, and uses the ARP protocol to find the set of available service nodes.
[0055] In this step, the network communication delay between the host and the service node is calculated as:
[0056]
[0057] Among them, L b is the fixed delay for communication startup, d i,j is the amount of data transmitted between the host and the service node, B t The bandwidth between the host and the service node is used to perform network perception and discover adjacent service nodes through the ARP protocol.
[0058] Step 4: Update the performance prediction value of each host node in real time through the Q-Leaning model, and dynamically adjust the scheduling strategy according to the network status and service node scheduling model. In addition, adopt a continuous learning mechanism to regularly train and update the Q-Leaning model to generate an optimized scheduling plan.
[0059] In this step, the Q-Leaning model is trained and updated regularly, that is, the instant reward calculation formula after taking an action in the current state is:
[0060] R(S t , α t )=ω1·(-L t )+ω2·B t+ω3·P+ω4·(1-R t )
[0061] ω1+ω2+ω3+ω4=1
[0062] In the embodiment of the present invention, ω1: delay weight coefficient 0.4; ω2: bandwidth weight coefficient 0.3; ω3: historical selection times ratio weight coefficient 0.1; ω4: load state weight coefficient 0.2.
[0063] If the startup time is sensitive, the weight of ω1 is increased to emphasize the importance of latency. If high bandwidth is required to support data transmission, the weight of ω2 is increased. The weights of ω3 and ω4 are adjusted according to the requirements for quality communication and node load. R(S t , α t ) is used to evaluate the current action α t The larger the reward value, the better the currently selected service node is in terms of delay, bandwidth, packet loss rate and node load status, and it is suitable for priority selection. t Indicates the network delay between the current host and the service node. The network delay is randomly generated in the range of 10ms to 100ms. The lower the delay, the better the performance, so it is represented by a negative value in the reward. t Indicates the bandwidth between the current host and the service node. The randomly generated range is 100Mbps to 10Gbps. The higher the bandwidth, the stronger the network throughput and the greater the reward. P indicates the historical selection ratio of the current service node. The more times the historical selection is made, the initial value is 0. The higher the ratio, the higher the reward. t Indicates the load status of the service node, indicating the resource usage of the node. The lower the load, the more available resources, so 1-R t Represents reward. The regression coefficient is used here to quantify the impact of different indicators on the overall performance of the system.
[0064] In this step, performance prediction is achieved by updating the Q value according to the following formula. The goal is to select actions in different states to maximize the cumulative reward. Specifically, the matching quality between the host and the service node is evaluated through the reward function, and the optimal scheduling strategy is learned so that the host can always select the optimal service node in a dynamic environment, thereby optimizing the startup time and success rate.
[0065] After each action is performed, observe the system enter a new state S t+1 , record the network performance of the next state. Update the Q value each time according to the formula:
[0066] Q(S,α)=Q(S,α)+α[R(S,α)+γ·maxQ(s′,a′)-Q(S,α)]
[0067] Among them, Q is a two-dimensional array, the rows in the two-dimensional array represent all selectable states S, and the columns in the two-dimensional array represent all possible actions A, that is, candidate service nodes. At initialization, all values in the Q table are 0. The learning rate α determines the speed of the model's learning progress each time the Q value is updated. The value is set to 0.1. The discount factor γ is used to control the impact of future rewards on current decisions. The value is set to 0.9. The ∈-greedy strategy is used to select the next action. The action is randomly selected with probability ∈, and a new service node is tried. The action selection, execution of the action, observation of the new state, and update of the Q value are repeated until Q converges.
[0068] Step 5: According to the optimized scheduling plan generated by the Q-Leaning model, the host obtains the startup image from the service node with the startup image after model training and completes the startup process. The relevant data of each startup completion is collected and fed back to the bare metal management platform OpenStack.
[0069] like Figure 3 As shown in the figure, the data comparison of correctly selecting the optimal service node under the influence of different background traffic is compared. Specifically, it is assumed that each link adopts equal cost multi-path routing (ECMP) and randomly generates background traffic, and the range of background traffic is 2-10Mbps, 2-50Mbps and 2-100Mbps respectively. The present invention has conducted multiple experiments, with a total of 100 samples in each group. The experimental results are shown in the figure. Figure 3 Figure 2 shows incorrect choices under different conditions.
[0070] Through experimental comparative analysis, it can be concluded that the Q-Leaning algorithm using gradient descent optimization shows higher stability and accuracy under background traffic fluctuations, and effectively reduces the wrong selection caused by noise and peak fluctuations.
[0071] The present invention formalizes the problem of bare metal startup scheduling under a distributed architecture, avoids congestion problems in traditional network transmission, and maximizes the bare metal startup rate and stability. First, in the bare metal information collection stage of the distributed architecture, the current state space information is obtained through network testing tools and converted into machine learning input to provide support for subsequent scheduling and optimization. Secondly, in the stage of selecting a service node for the host, machine learning is used to obtain a link that transmits smoothly and quickly in the next period of time; finally. Utilizing the operating system function of the network card and combining it with remote hard disk technology, the bare metal host startup process is simplified and the startup time is shortened.
[0072] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A bare metal startup optimization scheduling method in a distributed architecture based on machine learning, characterized in that: The bare metal startup optimization scheduling method under the distributed architecture is implemented by a service node scheduling model based on machine learning, and specifically includes the following steps: Step 1: Establish a distributed bare metal architecture, migrate bare metal management functions to the host network card, and collect host deployment information through the bare metal management platform OpenStack; Step 2: Initialize the service node scheduling model and use the gradient descent algorithm to process the deployment information collected in step 1; Step 3: The service node scheduling model combines the network communication delay between the host and the service node, and uses the ARP protocol to find the set of available service nodes; Step 4: Update the performance prediction value of each host node in real time through the machine learning model, and dynamically adjust the scheduling strategy according to the network status and service node scheduling model. In addition, adopt a continuous learning mechanism to regularly train and update the machine learning model to generate an optimized scheduling plan. Step 5: According to the optimized scheduling plan generated by the machine learning model, the host obtains the boot image from the service node with the boot image after model training and completes the boot process. The host collects the relevant data of each boot completion and feeds the data back to the bare metal management platform OpenStack.
2. The bare metal startup optimization scheduling method under a distributed architecture based on machine learning according to claim 1 is characterized in that: In step 1, the distributed bare metal architecture established includes multiple management schedulers (IronicConductor), HTTP servers, TFTP servers, hosts, and DHCP servers. Ironic Conductor obtains host deployment information, the host configures the network and provides the TFTP server address, the host downloads the lightweight service from the TFTP server and starts the proxy service, the host selects the best Ironic Conductor to establish a connection based on the network status, the host obtains the complete image from the best HTTP server and restarts, and migrates the bare metal management function to the network card connected via PCIe. The network card is a remote hard disk. The host deployment information is collected, including the network status between the host and the server and the historical startup data of each host. Bare metal no longer needs to download the operating system.
3. The bare metal startup optimization scheduling method under a distributed architecture based on machine learning according to claim 1 is characterized in that: In step 3, the network communication delay between the host and the service node is calculated as: Among them, L b is the fixed delay for communication startup, d i,j is the amount of data transmitted between the host and the service node, B t The bandwidth between the host and the service node is used to perform network perception and discover adjacent service nodes through the ARP protocol.
4. The bare metal startup optimization scheduling method under a distributed architecture based on machine learning according to claim 1 is characterized in that: The machine learning model is a Q-Leaning model.
5. The bare metal startup optimization scheduling method under a distributed architecture based on machine learning according to claim 4 is characterized in that: In step 4, the Q-Leaning model is trained and updated regularly, and the updating process is expressed as follows: R(S t ,a t )=ω1·(-L t )+ω2·B t +ω3·P+ω4·(1-R t ) ω1+ω2+ω3+ω4=1 Among them, ω1, ω2, ω3 and ω4 represent the importance weights of delay, bandwidth, historical selection times ratio and node load status. If the startup time is sensitive, the weight of ω1 is increased to emphasize the importance of delay. If high bandwidth is required to support data transmission, the weight of ω2 is increased. The weights of ω3 and ω4 are adjusted according to the requirements for quality communication and node load. R(S t ,α t ) is used to evaluate the current action α t The larger the reward value, the better the currently selected service node is in terms of delay, bandwidth, packet loss rate and node load status. t Indicates the network delay between the current host and the service node, B t Indicates the bandwidth between the current host and the service node. P indicates the historical selection ratio of the current service node. The more times the historical selection is made, the higher the ratio is. R t Indicates the service node load status.
6. The bare metal startup optimization scheduling method based on machine learning in a distributed architecture according to claim 5 is characterized in that: In step 4, performance prediction is achieved by updating the Q value according to the following formula: Q(S,α)=Q(S,α)+α[R(S,α)+γ·maxQ(s′,a′)-Q(S,α)] The Q table is a two-dimensional array. The rows in the two-dimensional array represent all selectable states S, and the columns in the two-dimensional array represent all possible actions A, that is, candidate service nodes. The ∈-greedy strategy is adopted to select the next action. The action is randomly selected with probability ∈, and new service nodes are tried. The action selection, execution, observation of new states, and updating of Q values are repeated until the Q table converges.
7. The bare metal startup optimization scheduling method based on machine learning in a distributed architecture according to any one of claims 4 to 6, characterized in that: The service node scheduling model based on machine learning includes: The distributed bare metal architecture established by the distributed architecture design module; System initialization module: initializes the parameters in the service node scheduling model and the distributed bare metal architecture; Network perception module: performs network perception through ARP protocol and searches for adjacent service nodes; Action selection and scheduling decision module: Based on the ∈-greedy strategy of Q-Leaning, a service node with the largest Q value in the Q table is selected for scheduling, and the Q value is updated after execution; Reward calculation module: Combined with the formula R(S t ,α t ) Calculate the reward after service node selection; Q value update module: updates the Q value in the Q table according to the formula Q(S, α); Scheduling strategy evaluation module: regularly evaluates the effectiveness of scheduling strategies, checks whether the Q table converges, and adjusts strategies; Service node selection and scheduling execution module: selects the best service node to perform bare metal startup, collects startup information, and records the selected node.