Computing power resource scheduling method of intelligent ship

By employing multi-agent deep reinforcement learning algorithms, the application scenarios of resources have been intelligentized and optimized. This enables the dynamic allocation of resources and their application scenarios, significantly improving resource utilization and reducing operating costs. This provides key technical support for the reliable, efficient, and autonomous operation of intelligent ships.

CN120849096APending Publication Date: 2025-10-28ZHENDUI IND ARTIFICIAL INTELLIGENCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510876709.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

During navigation, intelligent ships are constrained by space, energy consumption, and complex marine environments, making it difficult to flexibly expand hardware resources. Traditional static resource allocation strategies are unable to meet dynamic computing needs, resulting in low resource utilization and affecting navigation safety.

Method used

A multi-agent deep reinforcement learning algorithm is used to collect computing resource data in real time, build a Markov decision process, and construct a scheduling model based on task execution cost and resource utilization. The solution is solved through a multi-agent deep reinforcement learning algorithm to optimize resource allocation.

Benefits of technology

It enables dynamic and precise resource allocation under limited resources, avoids resource waste, improves system reliability and task completion time, significantly improves resource utilization, reduces costs, optimizes system performance and resource utilization, improves system reliability and task completion efficiency, significantly expands the application scenarios of resources, reduces operating costs, and provides key technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849096A_ABST
    Figure CN120849096A_ABST
Patent Text Reader

Abstract

The invention relates to a computing power resource scheduling method of an intelligent ship, and belongs to the technical field of machine learning. The method comprises the following steps: collecting computing power resource data of shipborne computing equipment of an intelligent ship in real time; predicting the computing power resource demand of each task based on the computing power resource data and the computing tasks; constructing a shipborne computing power resource scheduling model by taking the minimum task execution cost as a target and taking the task completion period as a constraint; constructing a Markov decision process, and solving the shipborne computing power resource scheduling model by using a multi-agent deep reinforcement learning algorithm based on the computing power resource demand to obtain an optimal scheduling strategy; and allocating computing power resources based on the optimal scheduling strategy. According to the method, under the condition that the total amount of shipborne computing power resources is limited, more shipborne task instances are served, elastic expansion and contraction of the resources are achieved, and an efficient and self-adaptive computing power resource management scheme is provided for the intelligent ship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to a method, system and apparatus for scheduling computing resources of intelligent ships. Background Technology

[0002] With the rapid development of intelligent ship technology, the demand for computing resources from shipboard intelligent applications is increasing dramatically. However, as edge computing nodes, ships are limited by space, energy consumption, and the complex marine environment, making it difficult to scale hardware resources as flexibly as cloud data centers. Meanwhile, the dynamic changes in load demands during navigation (such as sudden tasks and decisions made in severe weather) make traditional static resource allocation strategies inadequate to meet the needs of efficient and elastic computing, resulting in low resource utilization and even impacting navigation safety. Different tasks have significantly different computing resource requirements, and task loads change dynamically over time. Traditional resource scheduling methods often struggle to respond to these changes in real time and accurately, leading to low resource utilization, wasted idle resources, and performance limitations for some tasks due to insufficient resources. Furthermore, as business grows, the number of instances requiring service continues to increase. A significant challenge currently faces is how to serve more shipboard task instances with limited total resources, and ensure that the cluster can automatically adjust the number of shipboard task instances based on application load to achieve elastic resource scaling. Summary of the Invention

[0003] Based on the above analysis, the present invention aims to disclose a method, system and device for scheduling computing resources of intelligent ships. By utilizing a multi-agent deep reinforcement learning algorithm, under the condition of strictly limited ship hardware resources, it breaks through the limitations of traditional static scheduling, dynamically adjusts the resource allocation strategy, and ensures that critical navigation tasks obtain computing resources first.

[0004] On the one hand, the present invention provides a method for scheduling computing resources of intelligent ships, specifically including the following steps:

[0005] Real-time acquisition of computing power resource data from intelligent shipboard computing equipment;

[0006] Based on the computing resource data and computing tasks, predict the computing resource requirements of each task;

[0007] A shipboard computing resource scheduling model is constructed with the goal of minimizing task execution cost and the task completion deadline as a constraint.

[0008] A Markov decision process is constructed, and the shipborne computing resource scheduling model is solved using a multi-agent deep reinforcement learning algorithm based on the computing resource requirements to obtain the optimal scheduling strategy.

[0009] Allocate computing resources based on the optimal scheduling strategy.

[0010] Furthermore, the task execution cost is calculated based on the task completion time, resource utilization rate, and resource overload penalty.

[0011] Furthermore, the construction of the Markov decision process, based on the computing resource requirements, utilizes a multi-agent deep reinforcement learning algorithm to solve the shipborne computing resource scheduling model, including:

[0012] Identify key elements, including a set of states, a set of actions, a set of policies, a reward function, and a reward.

[0013] Based on the aforementioned key elements, the process of solving the shipborne computing power resource scheduling model is transformed into a Markov decision process; wherein, the computing power resource data and the real-time status of the task queue are used as inputs to the Markov decision process, and the optimal scheduling strategy is used as the output; the optimal scheduling strategy includes the priority of each task and the corresponding resource allocation.

[0014] Construct and train a multi-agent deep reinforcement learning algorithm model;

[0015] The Markov decision process is solved using a trained multi-agent deep reinforcement learning algorithm model to obtain the optimal scheduling strategy.

[0016] Furthermore, the state set is constructed based on state variables; wherein the state variables include the average waiting time of the task, the standard deviation of the average waiting time of the task, the CPU load change rate, the memory fragmentation rate, the disk IOPS load rate, and the bandwidth utilization rate.

[0017] Furthermore, the construction and training of the multi-agent deep reinforcement learning algorithm model includes:

[0018] Construct a multi-agent deep reinforcement learning algorithm model, wherein the multi-agent deep reinforcement learning algorithm model includes agents and environment;

[0019] The intelligent agents include priority intelligent agents and resource allocation intelligent agents;

[0020] The environment includes the task queue and computing resource data;

[0021] The agent interacts with the environment to update its network parameters, thereby training a multi-agent deep reinforcement learning algorithm model.

[0022] Furthermore, both the priority agent and the resource allocation agent include a main network, a target network, and an experience replay pool; the process of updating the network parameters of each agent through interaction with the environment to train a well-trained multi-agent deep reinforcement learning algorithm model includes:

[0023] Initialize each of the experience replay pools, target networks, and main networks of each of the aforementioned intelligent agents;

[0024] Based on the task queue, determine whether the current moment is a decision point;

[0025] If the current moment is determined as the decision point, each agent selects an action from the action set according to a preset reinforcement learning strategy; among them, the priority agent selects the action of determining the task priority, and the resource allocation agent selects the action of allocating computing power to the task based on the computing power resource requirements.

[0026] Each intelligent agent executes the action a. t Calculate the reward based on the reward function and update the state set;

[0027] Each agent stores its experience in its respective experience replay pool; the experience includes the current state set, the next state set, the action performed at the current moment, and the reward for performing the action;

[0028] Check whether the number of samples stored in each experience pool exceeds the minimum batch size. If it does not exceed the minimum batch size, continue collecting experience. If it does exceed the minimum batch size, randomly select a batch of samples from the experience pool.

[0029] Based on the samples in the aforementioned batch, the parameters of each of the corresponding main networks are updated;

[0030] When the number of execution time steps is determined to be greater than the preset number, the parameters of the corresponding target networks are updated according to the network parameters of each main network.

[0031] Once the set number of training iterations is reached, the network parameters are saved to obtain the preset agent deep reinforcement learning algorithm model.

[0032] Furthermore, the step of using a trained multi-agent deep reinforcement learning algorithm model to solve the Markov decision process and obtain the optimal scheduling strategy includes:

[0033] The current state set is obtained based on the computing resource data;

[0034] The optimal scheduling strategy is obtained by using a trained multi-agent deep reinforcement learning algorithm model based on the real-time state of the task queue and the state set at the current moment.

[0035] Furthermore, the reward function is expressed as:

[0036] rt=-α*Task_Time+β*Utilization-γ*Overload_Penalty

[0037] Where Task_Time represents the task completion time, Utilization represents the resource utilization rate, Overload_Penalty represents the resource overload penalty, and α, β, and γ represent weight parameters.

[0038] This invention also provides a computing resource scheduling system for intelligent ships, comprising:

[0039] The data acquisition module is used to collect computing resource data of the intelligent ship's onboard computing equipment in real time;

[0040] The computing power prediction module is used to predict the computing power resource requirements of each task based on the computing power resource data and computing tasks.

[0041] The computing resource scheduling module is used to construct an onboard computing resource scheduling model with the goal of minimizing task execution cost and the task completion deadline as a constraint; construct a Markov decision process, and use a multi-agent deep reinforcement learning algorithm to solve the onboard computing resource scheduling model based on the computing resource requirements to obtain the optimal scheduling strategy; and allocate computing resources based on the optimal scheduling strategy.

[0042] The present invention also provides a computing resource scheduling device for intelligent ships, comprising:

[0043] Memory, used to store computer programs;

[0044] A processor, when executing a computer program, implements the computing resource scheduling method as described in any one of claims 1-8.

[0045] The present invention can achieve at least one of the following beneficial effects:

[0046] By accurately predicting task requirements and employing reinforcement learning-based optimal scheduling strategies, the system can dynamically and precisely allocate limited onboard computing resources. This avoids resource idleness or contention bottlenecks caused by static resource allocation or simple polling, maximizing resource utilization and directly reducing the computational energy and time costs required to complete tasks. By using task completion deadlines as the core constraint, the optimization process prioritizes ensuring that critical tasks or tasks with tight deadlines receive the necessary resources. It can respond to changes in system task status in real time, maximizing the chances that all tasks can be completed before the deadline, thus improving the reliability and determinism of system services.

[0047] Through reinforcement learning models trained online or offline, the system can automatically and in real time generate near-optimal scheduling decisions based on real-time collected data and prediction results. In the highly dynamic and uncertain ship computing environment, it realizes intelligent and optimized resource allocation, significantly improves resource utilization efficiency, task completion rate and overall system performance, while reducing operating costs, and provides key technical support for the reliable, efficient and autonomous operation of intelligent ships.

[0048] By performing task prediction and scheduling locally onboard rather than relying on the cloud, the impact of network latency on decision-making during navigation is effectively avoided for intelligent ships. By employing lightweight neural network models to predict computing resource requirements, the computational resources consumed in prediction are reduced while ensuring sufficiently small prediction errors. This allows intelligent ships to prioritize the allocation of computing resources for urgent tasks.

[0049] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0050] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0051] Figure 1 This is a flowchart of the method of the present invention;

[0052] Figure 2 This is a schematic diagram of the deep reinforcement learning algorithm architecture of the present invention;

[0053] Figure 3 This is a schematic diagram of the process for training the multi-agent deep reinforcement learning algorithm model of the present invention. Detailed Implementation

[0054] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0055] Method Implementation Examples

[0056] An embodiment of the present invention discloses a method for scheduling computing resources of intelligent ships, specifically including steps S1 to S4.

[0057] S1. Real-time acquisition of computing resource data from intelligent shipboard computing equipment.

[0058] Specifically, computing resources include CPU, memory, disk IPIOS, and network bandwidth, with corresponding computing resource data including CPU load, memory usage, disk IPIOS, and network bandwidth consumption. It should be noted that when intelligent ships navigate complex waterways (such as narrow straits or ice zones), the navigation system may increase processor load (real-time path planning), while environmental sensor data surges, driving up memory usage. Unexpected situations can also increase the computational load of ship tasks (collision avoidance, storm avoidance, etc.). Therefore, during ship navigation, computing resource data may fluctuate significantly, and in special circumstances, very frequently. In practical applications, the data acquisition equipment supports high-frequency data acquisition and ensures stable operation even under harsh conditions such as ship vibration and salt spray, providing real-time and reliable computing resource data for system decision-making.

[0059] S2. Based on the computing resource data and computing tasks, predict the computing resource requirements of each task.

[0060] During the navigation of an intelligent ship, exemplary computational tasks include: navigation route planning, heading calculation, turning point calculation, collision avoidance, excess water depth calculation, propulsion performance calculation, fuel management, and load optimization. The urgency of each computational task varies, and the required computing resources also differ.

[0061] Specifically, before practical application, a training set is constructed based on the computing tasks of the ship and the corresponding historical computing resource data at multiple time points; a neural network prediction model is trained based on the training set (for example, a convolutional neural network, graph neural network, etc. can be used) to obtain the trained neural network prediction model.

[0062] During implementation, the computing resource requirements of each task are predicted based on the computing tasks in the task queue of the current observation period and the current computing resource data.

[0063] S3. Construct a shipboard computing resource scheduling model with the goal of minimizing task execution cost and the task completion deadline as a constraint.

[0064] Specifically, the task execution cost is calculated based on task completion time, resource utilization rate, and resource overload penalty.

[0065] Among them, task completion time refers to the time from the start time to the end of all task execution; resource utilization rate is the average utilization rate of each computing resource.

[0066] Specifically, resource overload includes GPU computing power overload, memory overflow, bandwidth congestion, and task latency exceeding thresholds. Furthermore, overload penalties are determined based on the percentage of actual usage exceeding resource limits, the duration of overload, and the criticality of the task.

[0067] Optionally, the overload penalty can be calculated using the following formula:

[0068] Overload_Penalty=(e λP -1)×(1+μT)×C

[0069] Wherein, Overload_Penalty is the overload penalty; λ is the overload sensitivity coefficient, which can be determined based on historical experience or learning iteration, with an initial preferred value of 5; P is the proportion of usage exceeding the resource limit, which is the average of the overload proportions of each resource; T is the duration of overload; μ is the time penalty coefficient, which is determined based on historical experience or learning iteration, with an initial preferred value of 0.01; C is the criticality of the task that caused the overload, and when multiple tasks cause the overload, the highest criticality value is taken.

[0070] Furthermore, with the objective of minimizing task execution cost and the task completion deadline as a constraint, it can be expressed as follows:

[0071] minf = mincost(U);

[0072] stmksp(U)≤Z;

[0073] Where minf represents the objective of minimizing task execution cost, U is the feasible solution for task execution, mincost(U) represents the task execution cost, mksp(U) represents the maximum completion time for the task set to be executed, and Z represents the given completion time constraint value.

[0074] S4. Construct a Markov decision process, and based on the computing resource requirements, use a multi-agent deep reinforcement learning algorithm to solve the shipborne computing resource scheduling model to obtain the optimal scheduling strategy. This specifically includes S41 to S44.

[0075] S41. Determine the key elements, which include a set of states, a set of actions, a set of strategies, a reward function, and a reward.

[0076] The state set is constructed based on state variables, which include the average waiting time of the task, the standard deviation of the average waiting time of the task, the CPU load change rate, the memory fragmentation rate, the disk IOPS load rate, and the bandwidth utilization rate.

[0077] Average wait time for a task (SWT) ave The calculation method for (t) is as follows:

[0078]

[0079] Where SEN represents the task sequence that has been completed at time step t, SSN represents the task sequence that is being executed at time step t, SAN represents the task sequence that has arrived at time step t but has not yet started execution, SN(t) represents the sum of all task sequences at time step t, i.e., the total number of tasks at time step t, x i This represents the actual execution time of the i-th task. This represents the arrival time of the i-th task.

[0080] The standard deviation of the mean waiting time for a task (SWT) std The calculation method for (t) is as follows:

[0081]

[0082] The method for calculating the CPU load change rate ΔCPU(t) is as follows:

[0083]

[0084] Where CPU(t) represents the CPU utilization rate at the current time step t, CPU(t-1) is the CPU utilization rate at the previous time step, and Δt represents the time interval.

[0085] The memory fragmentation rate (Fragmentation(t)) is calculated as follows:

[0086]

[0087] Disk IOPS load rate IOPS Load(t) The calculation method is as follows:

[0088]

[0089] Where IOPS(t) represents the disk IOPS load rate at the current time step t. max This indicates the disk's maximum IOPS capacity.

[0090] The method for calculating bandwidth utilization is as follows:

[0091]

[0092] Where BW(t) represents the network bandwidth usage at the current time step, BW max This represents the maximum bandwidth capacity of the network.

[0093] Furthermore, an action set is constructed based on action variables; wherein the action variables include determining task priority and allocating computing power to tasks.

[0094] Furthermore, the strategy set refers to the set of all possible strategies, where each strategy is the corresponding action to be taken in any state. In other words, the strategy set is the set of all possible actions that can be taken in the Markov decision-making process.

[0095] The reward function is used to quantify the immediate benefit of performing the corresponding action and transitioning to the next state in any given state. The reward function is expressed as:

[0096] r t =-α*Task_Time+β*Utilization-τ*Overload_Penalty

[0097] Here, Task_Time represents the task completion time based on the current time step and future actions, Utilization represents resource utilization, Overload_Penalty represents resource overload penalty, and α, β, and τ represent weight parameters. Furthermore, the values ​​of each weight parameter can be determined according to the target requirements. For example, if the target requirement is to minimize task completion time, then α is set to a larger value; if resource utilization is the primary consideration, then β is set to a larger value; if the target requirement is to avoid resource overload as much as possible, then τ is set to a larger value. The values ​​of each weight parameter can also be determined based on historical data or through iterative adjustments.

[0098] Furthermore, the reward is the cumulative amount of all future rewards starting from the current time step.

[0099] S42. Based on the key elements, the process of solving the shipborne computing power resource scheduling model is converted into a Markov decision process; wherein, the computing power resource data and the real-time status of the task queue are the inputs of the Markov decision process, and the optimal scheduling strategy is the output; the optimal scheduling strategy includes the priority of each task and the corresponding resource allocation.

[0100] The priority of each task refers to the order in which each task is executed in the task queue.

[0101] S43. Construct and train a multi-agent deep reinforcement learning algorithm model, wherein the multi-agent deep reinforcement learning algorithm model includes agents and environment;

[0102] The intelligent agent includes a priority intelligent agent and a resource allocation intelligent agent; both the priority intelligent agent and the resource allocation intelligent agent include a main network, a target network, and an experience replay pool.

[0103] The environment includes the task queue and computing resource data;

[0104] The agent interacts with the environment to update its network parameters, thereby training a multi-agent deep reinforcement learning algorithm model.

[0105] Furthermore, the step of updating the network parameters of each agent by interacting with the environment to train a well-trained multi-agent deep reinforcement learning algorithm model includes:

[0106] Step 1. Initialize the experience replay pools, target networks, and main networks of each agent; specifically, in each agent, create an experience pool D with a capacity of N to store scheduling experience, and randomly initialize the parameters θ and θ' of the Main network and Target network. - Set the training count to episode=0 as the starting point for training;

[0107] Step 2. Set the current time step t to 0, and calculate each state variable based on the current computing power resource data to form the initial state, providing a starting point for the scheduling process;

[0108] Step 3. Determine whether the current time is a decision point based on the task queue; specifically, if there are unassigned tasks in the task queue, the current time is a decision point, proceed to step 4; otherwise, jump directly to step 9.

[0109] Step 4. Each agent selects an action from the action set using a preset reinforcement learning strategy;

[0110] Specifically, the preset reinforcement learning strategy is the ε-greedy strategy, which selects the action with the highest Q value when the probability is 1-ε; otherwise, it randomly selects an action.

[0111]

[0112] Among them, the priority agent is based on the ε-greedy policy, when the probability is 1-ε,

[0113] Specifically, when the probability is 1-ε, the priority agent selects the action of determining the priority of each task that can have the highest Q value (for example, task priority can be represented numerically, such as an integer from 1 to 100, where the larger the value, the higher the priority, i.e., the earlier it is executed), and the resource allocation agent selects the action of allocating computing power (including CPU utilization, memory usage, disk IOPS, and network bandwidth usage) to each task based on the predicted computing power resource requirements; when the probability is other, the priority agent selects the action of randomly determining the task priority for each task, and the resource allocation agent selects the action of allocating computing power to each task based on the predicted computing power resource requirements.

[0114] Step 5. Each agent executes its selected action: the priority agent allocates computing power to each task, and the resource allocation agent calculates rewards based on the reward function. After each agent executes its action, it updates its state set based on the computing power resource data after the action. The state sets before and after the action are represented as s. t and s t+1 ;

[0115] Step 6. Each agent stores its experience into its respective experience replay pool. The experience includes the current state set, the next state set, the action performed at the current moment, and the reward for performing the action. The experience representations for each agent are (s...). t ,a t ,r t ,s t+1 ) and (s t ,a′ t ,r t ,s t+1 ), where a t and a′ t These are the actions performed by the priority agent and the resource allocation agent, respectively.

[0116] Step 7. Check whether the number of samples stored in each experience pool exceeds the minimum batch size. If it does not exceed the minimum batch size, continue collecting experience. If it exceeds the minimum batch size, randomly select a batch of samples from the experience pool. If it does not exceed the minimum batch size, proceed to step 9.

[0117] Step 8. For each agent, update the parameters of the corresponding main network based on the samples of the batch; specifically, update the Main network, setting the state s t Input the Main network, estimate its Q value, and set the state s t+1 Input the target network, calculate the target value y by combining the reward value r(t), update the parameters θ of the main network using the gradient descent method, and introduce the DoubleDQN strategy to reduce the overestimation of Q value and improve the stability and accuracy of training.

[0118] The specific method for calculating the target value of Double DQN is as follows:

[0119]

[0120] Where θ and θ - These are the Main and Target parameters, respectively; γ is the discount factor. Represented as the next state s of the Main network t+1 Choose the action with the highest value, a.t+1 ; Indicates that the Target network is in state s t+1 Take action a t+1 The Q-value of the action is evaluated. That is, the optimal action for the next state is selected in the Main network as an online decision, while the value of the work is evaluated in the Target to prevent the Main network from overestimating the Q-value.

[0121] Furthermore, the Main network is trained by minimizing the loss function L(θ):

[0122]

[0123] Where L(θ) is the loss function, representing the difference between the predicted Q-value and the target value y of the Main network; (s t ,a t ,r t ,s t+1 (done) represents the data extracted from the sample; such as Figure 2 A schematic diagram of a deep reinforcement learning algorithm architecture;

[0124] Step 9. Update the Target network. After each C-step decision, copy the parameters θ from the Main network to the Target network θ. - To maintain the stability of the Target network; C is the specified number of steps;

[0125] Step 10. When each agent determines that the number of time steps to be executed is greater than the preset number, update the parameters of the corresponding target networks according to the network parameters of each main network.

[0126] Step 11. Once the set number of training iterations is reached, save the network parameters for each agent to obtain the preset agent deep reinforcement learning algorithm model. For example... Figure 3 A flowchart illustrating the process of training a multi-agent deep reinforcement learning algorithm model.

[0127] S44. Solve the Markov decision process using the trained multi-agent deep reinforcement learning algorithm model to obtain the optimal scheduling strategy.

[0128] The optimal scheduling strategy includes the priority of each task, i.e., the execution order of each task and the corresponding resource allocation.

[0129] S5. Allocate computing resources based on the optimal scheduling strategy.

[0130] Specifically, after solving the shipborne computing resource scheduling model using a multi-agent deep reinforcement learning algorithm based on the computing resource requirements to obtain the optimal scheduling strategy, computing resources are allocated to each task in the task queue according to the optimal scheduling strategy.

[0131] This embodiment discloses a computing resource scheduling method for intelligent ships. By accurately predicting task requirements and employing an optimal scheduling strategy based on reinforcement learning, it can dynamically and precisely allocate limited onboard computing resources. This avoids resource idleness or contention bottlenecks caused by static resource allocation or simple polling, maximizing resource utilization. It directly reduces the computing energy consumption and time costs required to complete tasks. By using task completion deadlines as the core constraint, the optimization process prioritizes ensuring that critical tasks or tasks with tight deadlines obtain the necessary resources. It can respond to changes in system task status in real time, maximizing the guarantee that all tasks can be completed before the deadline, thus improving the reliability and determinism of system services.

[0132] Through reinforcement learning models trained online or offline, the system can automatically and in real time generate near-optimal scheduling decisions based on real-time collected data and prediction results. In the highly dynamic and uncertain ship computing environment, it realizes intelligent and optimized resource allocation, significantly improves resource utilization efficiency, task completion rate and overall system performance, while reducing operating costs, and provides key technical support for the reliable, efficient and autonomous operation of intelligent ships.

[0133] System Implementation Examples

[0134] Another specific embodiment of the present invention discloses a computing power resource scheduling system for intelligent ships, including a data acquisition module, a computing power prediction module, and a computing power resource scheduling module.

[0135] The data acquisition module is used to collect computing resource data from the onboard computing equipment of intelligent ships in real time.

[0136] The computing power prediction module is used to predict the computing power resource requirements of each task based on the computing power resource data and computing tasks. Optionally, given the limited resources of onboard computing equipment in intelligent ships, in order to achieve a balance between prediction accuracy and computing / storage resource consumption, the computing power prediction module in this embodiment uses a simplified 1D-CNN model to predict the computing power resource requirements of each task based on the computing power resource data and computing tasks.

[0137] The computing resource scheduling module is used to construct a shipborne computing resource scheduling model with the goal of minimizing task execution cost and the task completion deadline as a constraint; construct a Markov decision process, and use a multi-agent deep reinforcement learning algorithm to solve the shipborne computing resource scheduling model based on the computing resource requirements to obtain the optimal scheduling strategy, and allocate computing resources based on the optimal scheduling strategy.

[0138] This embodiment discloses a computing resource scheduling system for intelligent ships. Task prediction and scheduling are performed locally onboard, rather than relying on the cloud, effectively avoiding the impact of network latency on decision-making during navigation. By employing a lightweight neural network model to predict computing resource requirements, the system reduces the computational resources consumed in prediction while ensuring sufficiently small prediction errors. This allows intelligent ships to prioritize the allocation of computing resources for urgent tasks.

[0139] Device Examples

[0140] Another specific embodiment of the present invention discloses a computing resource scheduling device for intelligent ships, the device comprising:

[0141] Memory, used to store computer programs;

[0142] A processor, when executing a computer program, implements the computing resource scheduling method as described in any one of claims 1-8.

[0143] Compared with the prior art, the beneficial effects of the intelligent ship computing resource scheduling device provided in this embodiment are basically the same as those provided in the method embodiment and the system embodiment, and will not be described in detail here.

[0144] It should be noted that the above embodiments are based on the same inventive concept, and any parts not described repeatedly can be referenced from each other.

[0145] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. A method for scheduling computing resources of intelligent ships, characterized in that, Includes the following steps: Real-time acquisition of computing power resource data from intelligent shipboard computing equipment; Based on the computing resource data and computing tasks, predict the computing resource requirements of each task; A shipboard computing resource scheduling model is constructed with the goal of minimizing task execution cost and the task completion deadline as a constraint. A Markov decision process is constructed, and the shipborne computing resource scheduling model is solved using a multi-agent deep reinforcement learning algorithm based on the computing resource requirements to obtain the optimal scheduling strategy. Allocate computing resources based on the optimal scheduling strategy.

2. The computing resource scheduling method according to claim 1, characterized in that, The task execution cost is calculated based on task completion time, resource utilization rate, and resource overload penalty.

3. The computing resource scheduling method according to claim 2, characterized in that, The construction of the Markov decision process, based on the computing resource requirements, utilizes a multi-agent deep reinforcement learning algorithm to solve the shipborne computing resource scheduling model, including: Identify key elements, including a set of states, a set of actions, a set of policies, a reward function, and a reward. Based on the aforementioned key elements, the process of solving the shipborne computing power resource scheduling model is transformed into a Markov decision process; wherein, the computing power resource data and the real-time status of the task queue are used as inputs to the Markov decision process, and the optimal scheduling strategy is used as the output; the optimal scheduling strategy includes the priority of each task and the corresponding resource allocation. Construct and train a multi-agent deep reinforcement learning algorithm model; The Markov decision process is solved using a trained multi-agent deep reinforcement learning algorithm model to obtain the optimal scheduling strategy.

4. The computing resource scheduling method according to claim 3, characterized in that, The state set is constructed based on state variables; wherein the state variables include the average waiting time of the task, the standard deviation of the average waiting time of the task, the CPU load change rate, the memory fragmentation rate, the disk IOPS load rate, and the bandwidth utilization rate.

5. The computing resource scheduling method according to claim 4, characterized in that, The construction and training of the multi-agent deep reinforcement learning algorithm model includes: Construct a multi-agent deep reinforcement learning algorithm model, wherein the multi-agent deep reinforcement learning algorithm model includes agents and environment; The intelligent agents include priority intelligent agents and resource allocation intelligent agents; The environment includes the task queue and computing resource data; The agent interacts with the environment to update its network parameters, thereby training a multi-agent deep reinforcement learning algorithm model.

6. The computing resource scheduling method according to claim 5, characterized in that, Both the priority agent and the resource allocation agent include a main network, a target network, and an experience replay pool; the process of updating the network parameters of each agent through interaction with the environment to train a well-trained multi-agent deep reinforcement learning algorithm model includes: Initialize each of the experience replay pools, target networks, and main networks of each of the aforementioned intelligent agents; Based on the task queue, determine whether the current moment is a decision point; If the current moment is determined as the decision point, each agent selects an action from the action set according to a preset reinforcement learning strategy; among them, the priority agent selects the action of determining the task priority, and the resource allocation agent selects the action of allocating computing power to the task based on the computing power resource requirements. Each agent performs the action, calculates the reward based on the reward function, and updates the state set; Each agent stores its experience in its respective experience replay pool; the experience includes the current state set, the next state set, the action performed at the current moment, and the reward for performing the action; Check whether the number of samples stored in each experience pool exceeds the minimum batch size. If it does not exceed the minimum batch size, continue collecting experience. If it does exceed the minimum batch size, randomly select a batch of samples from the experience pool. Based on the samples in the aforementioned batch, the parameters of each of the corresponding main networks are updated; When the number of execution time steps is determined to be greater than the preset number, the parameters of the corresponding target networks are updated according to the network parameters of each main network. Once the set number of training iterations is reached, the network parameters are saved to obtain the preset agent deep reinforcement learning algorithm model.

7. The computing resource scheduling method according to claim 6, characterized in that, The optimal scheduling strategy is obtained by solving the Markov decision process using a trained multi-agent deep reinforcement learning algorithm model, including: The current state set is obtained based on the computing resource data; The optimal scheduling strategy is obtained by using a trained multi-agent deep reinforcement learning algorithm model based on the real-time state of the task queue and the state set at the current moment.

8. The computing resource scheduling method according to any one of claims 3-7, characterized in that, The reward function is expressed as follows: r t =-α*Task_Time+β*Utilization-γ*Overload_Penalty Where Task_Time represents the task completion time, Utilization represents the resource utilization rate, Overload_Penalty represents the resource overload penalty, and α, β, and γ represent weight parameters.

9. A computing resource scheduling system for intelligent ships, characterized in that, include: The data acquisition module is used to collect computing resource data of the intelligent ship's onboard computing equipment in real time; The computing power prediction module is used to predict the computing power resource requirements of each task based on the computing power resource data and computing tasks. The computing resource scheduling module is used to construct an onboard computing resource scheduling model with the goal of minimizing task execution cost and the task completion deadline as a constraint; construct a Markov decision process, and use a multi-agent deep reinforcement learning algorithm to solve the onboard computing resource scheduling model based on the computing resource requirements to obtain the optimal scheduling strategy; and allocate computing resources based on the optimal scheduling strategy.

10. A computing resource scheduling device for intelligent ships, characterized in that, The device includes: Memory, used to store computer programs; A processor, when executing a computer program, implements the computing resource scheduling method as described in any one of claims 1-8.

Citation Information

Cited By

  • Self-adaptive scheduling method and system for computing network integration of unmanned equipment cluster

    CN121940395A