Multi-agent reinforcement learning method and device for spatial constraint job scheduling, electronic equipment and computer readable medium

By constructing a spatially constrained flexible job shop scheduling environment and using a multi-objective, multi-agent reinforcement learning method, the problems of transportation time and collision risks caused by ignoring spatial constraints in existing technologies are solved, achieving efficient and safe job scheduling.

CN120891801APending Publication Date: 2025-11-04BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510906303.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing flexible workshop scheduling methods typically ignore spatial constraints, leading to increased transportation time and collision risks, which affect scheduling efficiency and safety.

Method used

A flexible job shop scheduling environment with spatial constraints is constructed. A multi-objective, multi-agent reinforcement learning method is adopted. Pareto optimal solutions are trained through conditional agent networks and multi-objective hybrid networks. The reward function considering total completion time, movement cost, and spatial density is used to achieve collaborative scheduling among agents.

Benefits of technology

It improves the efficiency and safety of job scheduling, reduces transportation time and collision risks, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120891801A_ABST
    Figure CN120891801A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent reinforcement learning method and device for spatial constraint job scheduling, electronic equipment and a computer readable medium. The method comprises the following steps: constructing a job scheduling environment comprising an operation sequence of jobs, resource configuration and spatial positions of machines and spatial constraints of a job moving speed; modeling a spatially constrained job scheduling problem into a multi-target partially observable Markov decision process, wherein the multi-target partially observable Markov decision process comprises a state space, a joint action space, an observation space of an intelligent agent, a state transfer function and a reward function; designing a multi-target award function comprising a total completion time award, a movement cost award and a spatial density award; and training a parameter capable of obtaining the Pareto optimal solution of the job scheduling problem facing the spatial constraint through interaction between the conditional agent network and the multi-target hybrid network and the job scheduling environment by utilizing a centralized training decentralized execution framework.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure belongs to the field of job scheduling, and particularly relates to a multi-agent reinforcement learning method and device for spatial constraint job scheduling, an electronic device and a computer readable medium. BACKGROUND

[0002] Job shop scheduling problem (JSP) is an important combinatorial optimization problem in the field of production scheduling that has been studied for a long time. Specifically, this problem describes a processing system in which a set of jobs with a determined operation sequence needs to be processed on a set of machines, each operation of a job has only one machine that can be used, and the processing sequence of jobs on machines needs to be determined to achieve scheduling goals, such as minimizing the total time to complete all jobs.

[0003] Flexible job shop scheduling problem (FJSP) is a common extension of JSP, which assumes that for each operation of a job, there is at least one or more machines that can perform it. Therefore, to solve FJSP, not only the processing sequence of jobs on machines needs to be determined, but also an appropriate machine needs to be assigned to each operation. As an important branch of JSP, FJSP has been widely applied in various industries, including aerospace, manufacturing, and logistics warehouses.

[0004] Current research methods can be mainly divided into three categories: rule-based methods, heuristic methods, and reinforcement learning methods. Rule-based methods are the simplest, but when faced with large-scale and complex scheduling problems, they often cannot obtain optimal solutions and are heavily dependent on human intervention and the accumulation of historical experience. In contrast, heuristic methods have improved in intelligence and can provide better solutions to a certain extent. However, they can usually only produce approximate solutions and are mainly suitable for fixed scheduling environments, lacking broad adaptability and generalization ability. With the development of artificial intelligence technology, scheduling methods based on reinforcement learning have gradually become a research hotspot. This method, with the characteristics of reinforcement learning, can effectively solve the optimal solution of the problem and perform well in large-scale and random environments, showing stronger flexibility and practicality. Most reinforcement learning-based methods generally assume a virtual agent, i.e., a central scheduler, which can use global environmental state information to control all jobs and machines to complete scheduling. However, in a real intelligent scheduling system, robots that control jobs can only determine their own actions through local observations received by their own sensors. Therefore, it is necessary to use multi-agent reinforcement learning with local observations to solve FJSP.

[0005] Efficient scheduling can not only improve the success rate of work, but also optimize the use of resources, whether in factory manufacturing or logistics warehouse scenarios. With the development of industrial technology, scheduling problems are becoming increasingly complex, with more diverse factors and constraints to consider. However, most research on FJSP focuses on optimizing processing time, machine utilization, and other objectives, while spatial constraints are often overlooked. Current research needs to address two key spatial problems. The first problem is the transportation of jobs in the workshop. In actual production, the transportation time between processes often accounts for a large proportion of the total processing time, and the transportation time is different when different machines are selected. Most methods assume that transportation time is included in job processing time or introduce a time lag due to transportation. While this approach can reduce the difficulty of solving the problem, simple assumptions or fixed delays cannot reflect the actual situation, and ignoring this factor may affect the overall scheduling efficiency due to transportation time. In addition, the second problem is that when the density of jobs in the workshop increases, the risk of collision during transportation also increases. Mobile jobs or robots transporting jobs need to slow down or change their path due to the risk of collision, which increases the actual movement time and reduces the efficiency of the schedule and poses a safety hazard. In actual application scenarios, how to ensure efficient and safe transportation is a consideration, which also affects the efficiency of the entire scheduling process.

[0006] In summary, job scheduling faces problems such as low overall scheduling efficiency, and therefore a multi-agent reinforcement learning method for spatial constraint job scheduling is needed to address at least one of the above deficiencies. SUMMARY

[0007] To solve the job scheduling problem with the above challenges, improve the efficiency and / or safety of the scheduling process, and reduce operating costs, the present disclosure constructs a spatially constrained FJSP, taking into full consideration the distribution of entities in space, job transportation between machines, the limited nature of resources, and uncertain factors in the environment. The job shop is treated as a multi-agent system, and the scheduling task is modeled as a multi-objective partially observable Markov decision process. A multi-objective multi-agent reinforcement learning method is proposed to address this complex scheduling problem. Each job agent, based on its own local observations, collaborates with other agents by balancing time, space, and resource constraints, and learns decentralized strategies to maximize global rewards.

[0008] Aspects of the present disclosure provide a multi-agent reinforcement learning method for spatial constraint job scheduling.

[0009] Aspects of the present disclosure also provide a multi-agent reinforcement learning device for spatial constraint job scheduling.

[0010] The first aspect of the present disclosure provides a multi-agent reinforcement learning method for space-constrained job scheduling, the method comprising: constructing a space-constrained job scheduling environment comprising an operation sequence of a job, a resource configuration and a spatial position of a machine, and a job moving speed; modeling the space-constrained job scheduling problem as a multi-objective partially observable Markov decision process, the multi-objective partially observable Markov decision process comprising a state space, a joint action space, an observation space of an agent, a state transition function, and a reward function; designing a multi-objective reward function comprising a total completion time reward, a moving cost reward, and a spatial density reward; and training a parameter capable of achieving a Pareto optimal solution to the space-constrained job scheduling problem by interacting with the job scheduling environment through a conditional agent network and a multi-objective hybrid network using a centralized training and decentralized execution framework.

[0011] In an example embodiment, constructing a space-constrained job scheduling environment comprising an operation sequence of a job, a resource configuration and a spatial position of a machine, and a job moving speed comprises: setting a set of N jobs to be processed, each job having an operation sequence comprising P operations, where N and P are natural numbers greater than 1; setting a set of L machines, setting a resource configuration and a spatial position of each machine, where L is a natural number greater than 1; and setting a job moving speed.

[0012] In an example embodiment, the job scheduling objective is to minimize the maximum completion time T total :

[0013]

[0014] wherein, represents a waiting time of the i-th job before processing the p-th operation, represents a moving time of the i-th job from a position of the machine M i,p-1 processing the (p-1)-th operation to a position of the machine M i,p processing the p-th operation, the moving time being obtained by calculating a quotient of a distance between the two machines and the job moving speed, represents a processing time of the p-th operation, and i is a natural number greater than or equal to 1 and less than or equal to N, and p is a natural number greater than or equal to 1 and less than or equal to P.

[0015] In an example embodiment, the state space comprises an occupancy state of a machine, a remaining working time of a machine, a set of processable operations of a machine, a remaining operation number of a job, an identifier of a next operation of a job, and a processing time of the next operation of the job; the joint action space comprises a selection action of a machine, a waiting action, a busy action, and a completion action; and the observation space of an agent comprises a moving time of an agent from a current position of the agent to a machine and a set of other jobs within an observation range of the agent.

[0016] In an example embodiment, the total completion time reward is inversely proportional to the maximum remaining time of all jobs in the scheduling process, and inversely proportional to the total completion time of all jobs after the end of the scheduling; the movement cost reward is inversely proportional to the movement time of the agent from the current position to the target machine; and the space density reward is inversely proportional to the occupancy rate of the adjacent machines around the machine where the agent is currently located.

[0017] In an example embodiment, the parameters capable of obtaining the Pareto optimal solution of the space constraint oriented job scheduling problem are trained by the conditional agent network and the multi-objective hybrid network interacting with the job scheduling environment through the centralized training and decentralized execution framework, including: the conditional agent network outputs a Q-value vector including the objectives of total completion time, movement cost and space density based on the observation of the agent, historical action and preference vector; and the multi-objective hybrid network processes the Q-value vector of each objective using parallel tracks, and combines the outputs of each track to generate a multi-objective Q-value vector.

[0018] In an example embodiment, in the training phase, the preference space is divided into multiple subspaces, a preference vector is sampled in each round, and the sampling probability of the subspace is dynamically adjusted according to the distribution of non-dominated solutions.

[0019] The second aspect of the present disclosure provides a multi-agent reinforcement learning device for space constraint oriented job scheduling, which comprises: a job scheduling environment construction module configured to construct a space constraint oriented job scheduling environment including the operation sequence of jobs, the resource configuration and spatial position of machines, and the movement speed of jobs; a Markov decision process modeling module configured to model the space constraint oriented job scheduling problem as a multi-objective partially observable Markov decision process, including a state space, a joint action space, an observation space of an agent, a state transition function, and a reward function; a multi-objective reward function design module configured to design a multi-objective reward function including a total completion time reward, a movement cost reward, and a space density reward; and a centralized training and decentralized execution training module configured to train parameters capable of obtaining the Pareto optimal solution of the space constraint oriented job scheduling problem by the conditional agent network and the multi-objective hybrid network interacting with the job scheduling environment through the centralized training and decentralized execution framework.

[0020] A third aspect of this disclosure also provides a multi-agent reinforcement learning method for spatially constrained job scheduling. This method includes: introducing spatial constraints into the job scheduling problem to construct a spatially constrained flexible job shop scheduling environment; modeling the spatially constrained flexible job shop scheduling problem as a multi-objective partially observable Markov decision process, defining the state of the job scheduling environment, the observations of the agents, and their actions; shaping rewards for the spatially constrained flexible job shop scheduling problem by designing multi-objective rewards; and designing a multi-objective multi-agent reinforcement learning method to train a Pareto optimal solution for the spatially constrained flexible job shop scheduling problem through interaction with the job scheduling environment.

[0021] As a preferred embodiment of this disclosure, constructing a spatially constrained flexible job shop scheduling environment includes: setting N jobs J = {J1, J2, ..., J...} waiting to be processed. N Each job has a predetermined sequence of operations O = {O1, O2, ..., O}. P Set up L machines M = {M1, M2, ..., M} L The machine is equipped with the resources required for various operations to handle the job. The resource distribution of the machine is set as follows: Total | C l |Class Resources(|C) l Set the machine's location in the workshop and its movement speed (|≤P).

[0022] As a preferred embodiment of this disclosure, the spatially constrained flexible job shop scheduling environment is characterized by: (1) temporal diversity: the time for a job to complete an operation includes waiting time, movement time, and processing time; (2) spatial distribution characteristics: the machines are neatly and evenly arranged in the shop, and when a job is processed on a certain machine, the local density of the job can be calculated based on the occupancy of the surrounding machines (e.g., occupancy rate), which can measure the density of job movement in the shop and the risk of collision; and (3) resource randomness: the resource distribution of machines in the environment can change when processing different batches of jobs.

[0023] Based on the above three environmental characteristics, the scheduling objective for completing a batch of tasks is to maximize the completion time T of all tasks while reducing the density of local tasks. total The minimum can be expressed by the following formula:

[0024]

[0025] In the above formula, Indicates assignment J i In processing operation O p The previous waiting time, Indicates assignment J i From machine M used to process the previous operation i,p-1 Move the position to the machine M selected in the current operation. i,p The location movement time is calculated using the distance between the two machines and the operation movement speed. Indicates operation O p Processing time.

[0026] As a preferred embodiment of this disclosure, modeling a multi-objective partially observable Markov decision process includes: defining N agents, where each task is considered an agent; defining a state space S∈S, containing state information of the agents and the environment; and defining a joint action space A={A1,…,A...} n}, where A i This represents the action that agent i can take, a. i The set of states; define the state transition function P, denoted as P(s′|s,a): S×A×S→[0,1], which describes the probability of transitioning to state s′ after the agent takes a joint action a in state s; define the reward function R(s,a): S×A→R, which describes the reward obtained by the agent after taking an action. In a space-constrained job scheduling environment, due to the existence of multiple objectives, the reward is in the form of a vector of multiple objectives; define the agent's observation space O={O1,…,O n}; and define the discount factor γ∈[0,1].

[0027] As a preferred embodiment of this disclosure, the reward shaping for the spatially constrained flexible job shop scheduling problem designs different rewards for multiple objectives, including total completion time reward, movement cost reward and spatial density reward.

[0028] As a preferred embodiment of this disclosure, the multi-objective multi-agent reinforcement learning method employs a centralized training and distributed execution framework, which is divided into a conditional agent network and a multi-objective hybrid network. The training steps include:

[0029] The Conditional Agent Network consists of several multilayer perceptron layers and a gated recursive unit layer. It takes the agent's current observation, historical actions (e.g., the previous action), and preference vector as input to the network, and trains the model to generate appropriate policies based on the corresponding preferences.

[0030] By using a masking layer to restrict the agent network from selecting available actions, the Q-value vector of the currently available actions can be obtained.

[0031]

[0032] In the above formula, the Q-value vector represents the value of the selected action under the current observation and the preferred target, ω is a preference vector representing the preference weight for each target in the schedule, τ i represents the historical information of the agent i, represents the effectiveness of the action, after multiplying it with the Q value vector, the Q value of the effective action remains unchanged, and the Q value of the ineffective action will become negative infinity, so that the ineffective action will not be selected in the action selection;

[0033] The agent uses an ε-greedy algorithm to select the next action:

[0034]

[0035] The Q value of the action generated by each conditional agent network for each target is recombined, and the Q value of all agents for each target is combined into a vector, which is input into parallel tracks in the multi-target hybrid network, each track processes the Q value of one target and inputs the global state, and the weights and biases of the network are generated. Finally, the outputs of each track are connected together to generate a Q value vector Q composed of multiple targets tot , which represents the total action state value of all agents under different targets:

[0036]

[0037] In the update phase, the trajectories in the replay buffer are sampled, and the preference vector is also sampled from the preference space, and the TD (time difference) target is calculated:

[0038]

[0039] The target network and the online network are updated using the idea of DQN (deep Q network).

[0040] As a preferred embodiment of the present disclosure, the multi-target multi-agent reinforcement learning method, in the training phase, samples a preference vector as input every round, divides the entire preference space into multiple subspaces, and adjusts the sampling probability of different subspaces according to the current distribution of non-dominated solutions in the target space.

[0041] The fourth aspect of the present disclosure provides an electronic device comprising a processor and a memory, the memory being used to store a program, when the program is executed by the processor, the method as described above is executed.

[0042] The fifth aspect of the present disclosure provides a computer readable medium having a program stored thereon, when the program is executed by the processor, the method as described above is executed.

[0043] The technical scheme provided by the embodiment of the present disclosure has the beneficial effects that the multi-agent reinforcement learning method for space-constrained job scheduling can fully consider the entity distribution in the space, the job transportation between machines, the finiteness of resources, and uncertain factors in the environment, and can improve the scheduling efficiency and / or safety in the space-constrained flexible job shop scheduling problem.

[0044] The above technical scheme of the present disclosure realizes one of the aforementioned effects, and does not require each technical scheme to realize all the technical effects mentioned above.

[0045] In addition, the effects of the present disclosure not only include the effects set forth herein, but also include other effects that are obvious to those skilled in the art of the present disclosure by referring to the claims, the specification, and the drawings accompanying the specification. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical scheme in the embodiments of the present disclosure, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0047] Figure 1 is a flowchart of a multi-agent reinforcement learning method for space-constrained job scheduling according to an embodiment of the present disclosure;

[0048] Figure 2 is a schematic diagram of a flexible job shop scheduling environment for space constraints according to an embodiment of the present disclosure;

[0049] Figure 3 is a framework diagram of a multi-objective multi-agent reinforcement learning network according to an embodiment of the present disclosure;

[0050] Figure 4 is a block diagram of a multi-agent reinforcement learning device for space-constrained job scheduling according to an embodiment of the present disclosure;

[0051] Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure; and

[0052] Figure 6 is a block diagram of a computer readable medium according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0053] The present inventors have found the above problems after in-depth research on the existing flexible job shop scheduling method. Research has found that existing methods usually focus on optimizing processing time, machine utilization and other objectives, while spatial constraints are often ignored. However, in the job shop, space is also an important factor affecting the efficiency of job scheduling. First of all, the transportation problem of jobs in the workshop, in actual production, the transportation time between processes often accounts for a large proportion of the entire processing time, and the transportation time is different when different machines are selected. Most methods assume that the transportation time is included in the job processing time, or introduce a time lag due to transportation. Although this method can reduce the difficulty of solving, ignoring this factor may affect the overall scheduling efficiency of the transportation time. In addition, when the job density in the workshop becomes larger, the possibility of congestion and collision during transportation will also increase, and when the job moves in the workshop, it needs to wait or change the path due to the risk of collision, which will reduce the efficiency of scheduling.

[0054] It should be noted that the defects of the above prior art solutions are the result of the inventors' practice and careful research, therefore, the discovery process of the above problems and the technical solutions proposed by the embodiments of the present disclosure to solve the above problems are the contributions of the inventors to the prior art.

[0055] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in combination with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all. Generally, the components of the embodiments of the present disclosure described and shown in the drawings can be arranged and designed in various different configurations.

[0056] It should be noted that similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present disclosure, the terms "first", "second", "third", "fourth" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0057] After the above in-depth analysis, the embodiment of the present disclosure proposes a multi-agent reinforcement learning method for space-constrained job scheduling. First, the space constraint is introduced into the job scheduling problem, and a space-constrained flexible job shop scheduling environment is constructed. The space-constrained flexible job shop scheduling problem is modeled as a multi-objective partially observable Markov decision process, and the state of the job scheduling environment, the observation and action of the agent are defined. The reward shaping for the space-constrained flexible job shop scheduling problem is designed, and a multi-objective reward is designed. A multi-objective multi-agent reinforcement learning method is designed, which trains the Pareto optimal solution of the space-constrained flexible job shop scheduling problem by interacting with the job scheduling environment. The present disclosure regards the job as an agent, uses a multi-objective multi-agent reinforcement learning method, obtains the Pareto optimal solution of the multi-objective job scheduling problem, and improves the efficiency and safety of the space-constrained job scheduling.

[0058] In the following, the embodiments will be described with reference to the drawings.

[0059] Figure 1 is a flowchart of the multi-agent reinforcement learning method for space-constrained job scheduling according to the embodiment of the present disclosure. Figure 2 is a schematic diagram of the space-constrained flexible job shop scheduling environment according to the embodiment of the present disclosure. Figure 3 is a framework diagram of the multi-objective multi-agent reinforcement learning network according to the embodiment of the present disclosure.

[0060] The first aspect of the present disclosure provides a multi-agent reinforcement learning method for space-constrained job scheduling. As shown in Figure 1 The multi-agent reinforcement learning method for space-constrained job scheduling according to the embodiment of the present disclosure includes: constructing a space-constrained job scheduling environment including the operation sequence of the job, the resource configuration and the spatial position of the machine, and the moving speed of the job (S1); modeling the space-constrained job scheduling problem as a multi-objective partially observable Markov decision process, which includes a state space, a joint action space, an observation space of the agent, a state transition function, and a reward function (S2); designing a multi-objective reward function including a total completion time reward, a moving cost reward, and a space density reward (S3); and using a centralized training and decentralized execution framework, interacting with the job scheduling environment through a conditional agent network and a multi-objective hybrid network, and training parameters capable of obtaining the Pareto optimal solution of the space-constrained job scheduling problem (S4).

[0061] In step S1, the space constraint is introduced into the job scheduling problem, and a space-constrained flexible job shop scheduling environment is constructed. By modeling the spatial position of the machine and the moving speed of the job, the moving time between machines and the local space density can be considered, so as to improve the efficiency of scheduling and avoid collision.

[0062] In an example embodiment, a job shop scheduling environment that constructs a sequence of operations for jobs, resource configurations and spatial locations of machines, and a speed of job movement can include: setting a set of N jobs to be processed, each job having a sequence of P operations, where N and P are natural numbers greater than 1; setting a set of L machines, setting a resource configuration and spatial location of each machine, where L is a natural number greater than 1; and setting a speed of job movement.

[0063] As shown in Figure 2 , a flexible job shop scheduling environment with spatial constraints has N jobs J = {J1, J2,..., JN} waiting to be processed, each job having a determined sequence of operations O = {01, 02,..., OP} that must be processed in order. Uniformly scattered in the shop are machines M = {M1, M2,..., ML} that process various operations, each machine being equipped with resources R = {R1, R2,..., RM} needed to process the operations of the jobs. N P Each machine has a different set of resources, each resource having a certain amount, and each operation of a job consumes one of the resources. The resource distribution of a machine can be represented as L l l p When a resource is available at an idle machine, the job selects the machine and moves to its location to be processed. The job is assumed to be mobile, such as a robot in a logistics warehouse or a drone that needs to complete a security task, and its speed of movement is assumed to be v. All machines are non-preemptive, meaning that a machine processing an operation cannot be interrupted to process another operation even if it has the resources available.

[0064] To consider spatial constraints and randomness, the environment has the following characteristics: (1) temporal diversity: the time for a job to complete an operation includes waiting time, movement time, and processing time; (2) spatial distribution characteristics: the machines are arranged neatly and uniformly in the shop, and when a job is being processed at a machine, the local density of the job can be calculated based on the occupancy of the machines around it, which can measure the density and collision risk of the job moving in the shop; and (3) resource randomness: the resource distribution of the machines in the environment can change when processing different batches of jobs.

[0065] In an example embodiment, the scheduling goal of completing jobs is to minimize the maximum completion time T total ​​​​​​​​:

[0066]

[0067] In formula (1), represents the waiting time of the i-th job before processing the p-th operation, represents the moving time of the i-th job from the position of the machine M i,p-1 processing the (p-1)-th operation to the position of the machine M i,p processing the p-th operation, the moving time is obtained by calculating the quotient of the distance between the two machines and the moving speed of the job, represents the processing time of the p-th operation, and i is a natural number greater than or equal to 1 and less than or equal to N, and p is a natural number greater than or equal to 1 and less than or equal to P. For example, based on the above three environmental characteristics, the scheduling goal of completing a batch of jobs is to minimize the maximum completion time T total of all jobs.

[0068] In addition, the environment follows several assumptions: (1) all jobs of the same batch enter the waiting area of the workshop at the same time at the beginning and can all be processed; (2) when the job moves, the path planning is not considered, but the shortest straight line path from point to point is moved; (3) each job must process the operation in the predetermined order; (4) each machine can only process one operation of one job at the same time at most; and (5) each resource is only used for one operation of one job, and will be consumed after the operation is completed.

[0069] In step S2, the flexible job shop scheduling problem with space constraints is modeled as a multi-objective partially observable Markov decision process, and the state space of the job scheduling environment, the observation space of the agent and the joint action space are defined.

[0070] In an exemplary embodiment, the state space includes the occupancy state of the machine, the remaining working time of the machine, the set of processable operations of the machine, the number of remaining operations of the job, the identifier of the next operation of the job and the processing time of the next operation of the job; the joint action space includes the selection action of the machine, the waiting action, the busy action and the completion action; and the observation space of the agent includes the moving time of the agent from the current position to the machine and the set of other jobs within the observation range of the agent.

[0071] In step S2, the spatially constrained flexible job shop scheduling problem is modeled as a multi-objective Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Each agent can only observe a portion of the environmental state, and the decision-making process is decentralized, meaning each agent makes decisions independently without relying on a central controller. The multi-objective Dec-POMDP is represented as a tuple.<N,S,A,P,R,O,Ω,γ> N represents the number of agents (N>1), where each task is considered as one agent; s∈S represents the state space, containing the state information of the agents and the environment; A={A1,…,A n} represents the joint action space, where A i This represents the action that agent i can take, a. i The set of states; P represents the state transition function P(s′|s,a): S×A×S→[0,1], which describes the transition from state s to state s after the agent takes a joint action a in state s. ′ The probability; R represents the vector reward function R(s,a): S×A→R, describing the reward obtained by the agent after taking an action. In the case of complete cooperation, all agents share the reward function; O={O1,…,O n} represents the joint observation space composed of the observations o of each agent; Ω represents the observation function Ω(s,a): S×A→O, due to partial observability, each agent can only obtain its own observation value from the observation function; γ∈[0,1] is the discount factor.

[0072] The task agent and the scheduling environment interact at each time step t. Given a global state s... t In this process, each agent i acquires its own local observations. And take an action based on the observation. All agents form a joint action A t This causes the environment to transition to state s. t+1 And obtain a shared reward vector r t+1 The agent's goal is to learn the policy π(a). i |o i ):O i ×A i →[0,1], to maximize the discount return: In other words, the ultimate goal of scheduling tasks is to minimize the completion time of all jobs under space constraints.

[0073] The global state s∈S represents the global information of the current scheduling environment, including the state information of machines and jobs. The global state s takes the following form:

[0074]

[0075] In formula (2), represents the occupation state of each machine at the current time, that is, which job is occupied or idle, represents the remaining working time of the machine, represents the operation set that the machine can process, represents the number of remaining operations of each job at the current time, and represent the next operation identifier (ID) and the required time of the job to be processed, respectively. The global state is used to assist the agent training during training, and the agent cannot access the global state during execution.

[0076] The action space A = {1, 2, …, L, L+1, L+2, L+3}, a ∈ A represents the action to be taken by the agent in the next step. The meanings of different action types are shown in Table 1, which defines four types of actions: selection action, waiting action, busy action and completion action, to represent different behaviors of the agent.

[0077] Table 1

[0078]

[0079] The observation o ∈ O of the agent is the input of each agent network, which represents the information of each agent itself and the local information of the environment that it can observe, and each agent selects a suitable action according to its own observation at each time step. The local observation of the agent can be represented as:

[0080]

[0081] In formula (3), represents the time required for the agent i (or the i-th agent) to move from the current position to the machine l. obs i represents the observation range of the agent i, since the machines are evenly and uniformly distributed in the environment, where M i represents the set of adjacent machines around the current position of the agent (i.e. the machine selected in the last step), represents the jobs being processed on these machines.

[0082] In step S3, the reward shaping for the space-constrained flexible job shop scheduling problem is performed, and a multi-objective reward is designed. In the exemplary embodiment, a total completion time reward is inversely proportional to the maximum remaining time of all jobs during the scheduling process, and inversely proportional to the total completion time of all jobs after the scheduling is completed; a movement cost reward is inversely proportional to the movement time of the agent from the current position to the target machine; and a space density reward is inversely proportional to the occupancy rate of the neighboring machines around the machine where the agent is currently located.

[0083] In step S3, due to the introduction of the space constraint, multiple objectives are added to the flexible job scheduling problem, including the total completion time, the movement cost, and the space density, and a reward function needs to be designed for each objective.

[0084] The total completion time objective is to minimize the maximum completion time T total of all jobs, which can be obtained after all jobs are completed, so this time is designed as the reward after the scheduling is completed, and at each time step during the scheduling process, the reward is defined as the maximum remaining time of all jobs:

[0085]

[0086] In equation (4), represents the total time of the remaining operations of job i (or the i-th job), and K1 and K2 are hyperparameters. During the scheduling process, the shorter the total time of the remaining operations, the greater the reward, and after the scheduling is completed, the shorter the total time used, the greater the reward. This reward encourages the job agent to quickly complete all operations.

[0087] The movement cost objective is to minimize the longest movement time required by all jobs after the current selected machine, and the reward is defined as:

[0088]

[0089] In equation (5), represents the movement time of agent i from the position of the machine M i,p-1 selected in the previous step to the position of the machine M i,p selected in this step, which is obtained by calculating the quotient of the distance between the two machines and the speed: The agent will try to choose a closer available machine to obtain a higher movement reward.

[0090] The space density objective is to minimize the local job density of the job shop. The local job density is used to measure the degree of job concentration in the environment, and the greater the local density, the more intensive the jobs, and the greater the possibility of congestion and collision. The local density of agent i is where |obs i | and |M i| are the number of agents and machines in the observation range of agent i, respectively. The reward is defined as the average local density in the environment:

[0091]

[0092] Equation (6) obtains a greater reward by encouraging agents to choose machines more dispersedly, thereby reducing the local density.

[0093] Through the rewards of the above three objectives, the agent learns to balance the three objectives, complete the scheduling in a shorter time under the conditions of reducing the moving cost and the local density, and improve the efficiency and safety of the scheduling.

[0094] Step S4, a multi-objective multi-agent reinforcement learning method is designed to train a Pareto optimal solution for the flexible job shop scheduling problem facing spatial constraints by interacting with the job scheduling environment.

[0095] In an example embodiment, a centralized training and decentralized execution framework is used to train a parameter capable of obtaining a Pareto optimal solution for the job scheduling problem facing spatial constraints by interacting with the job scheduling environment through a conditioned agent network and a multi-objective mixing network. The parameter includes: a conditioned agent network that outputs a Q-value vector including objectives related to total completion time, moving cost, and spatial density based on the observation of the agent, historical actions, and preference vector; and a multi-objective mixing network that processes the Q-value vector of each objective in parallel tracks and combines the outputs of each track to generate a multi-objective Q-value vector.

[0096] In step S4, due to the existence of multiple constraint conditions, including time urgency and spatial safety, and these objectives are mutually conflicting, traditional single-objective optimization methods often have difficulty finding a global optimal solution. Therefore, a distributed multi-objective multi-agent reinforcement learning scheduling algorithm is used to deal with this complex problem. The specific method implementation adopts a centralized training and decentralized execution (CTDE) framework of multi-agent reinforcement learning. The network framework is divided into two parts: a conditioned agent network (CAN) and a multi-objective mixing network (MOMN), as shown in Figure 3 Step S4 can further include the following steps.

[0097] First, the design condition agent network CAN estimates the vector Q function of each agent, including the value function of all objectives. In this step, the CAN consists of several multi-layer perceptron (MLP) layers and a gated recurrent unit (GRU) layer. The current observation of the agent, the history of actions (e.g., the last step action), and the preference vector are taken as the input of the network, and the model is trained to generate appropriate strategies based on the corresponding preferences. The CAN generates the Q value vector for each objective based on the observation history of the agent which represents the value of the selected action under the current observation and the preferred objective, where ω is the preference vector representing the preference weight for each objective in the schedule, τ i represents the historical information of agent i. The purpose of inputting the preference vector is to make the neural network estimate the vector Q function with specific weights.

[0098] Second, use the mask layer to limit the action selection of the agent network. In this step, not all actions selected from the policy model are reasonable due to the constraints of the scheduling environment. Therefore, a mask layer is added to limit the selection of available actions by the policy model, and each agent maintains an available action vector which represents the effectiveness of the action. After multiplying it with the Q value vector, the Q value of the effective action remains unchanged, and the Q value of the ineffective action will become negative infinity. The agent can obtain the Q value vector of the current available action in each step from the scheduling environment

[0099]

[0100] Third, the agent uses the ε-greedy algorithm to select the next action. In this step, the agent generates Q values through partial observation, and uses the ε-greedy algorithm to select the next action in the training phase:

[0101]

[0102] Then use the hybrid network to generate the estimate of the joint action value function, which is used to calculate the TD error to train the CAN. When executed in a decentralized manner, the agent only selects the optimal action.

[0103] Fourth, use the hybrid network MOMN to generate the Q value vector Q tot of multiple objectives. In this step, the Q values of each action generated by each CAN for each objective are recombined, and the Q values of all agents for each objective are combined into a vector, which is input into the parallel tracks in the MOMN, each track contains two neural network layers, which process the Q values of one objective. Finally, the outputs of each track are connected together to generate the Q value vector Q tot which represents the total action state value of all agents under different objectives:

[0104]

[0105] In this step, multiple super networks are used to generate weights and biases for the MLP (Multi-Layer Perceptron) layers of the hybrid network. Each track has four super networks, two of which are used to generate weights and two of which are used to generate biases. Each super network that produces weights consists of a linear layer and uses an absolute value activation function to ensure that the output is non-negative. This ensures that the output of the hybrid network satisfies the monotonicity constraint. The super networks that produce biases do not require an absolute value activation function because biases do not have a non-negative constraint. For the biases of the last layer of each track, a two-layer super network and a ReLU (Rectified Linear Unit) activation function are used. All super networks take the global state as input, ensuring that the hybrid network can utilize global information.

[0106] Fifth, the TD target update network is calculated. In this step, a trajectory is sampled from the replay buffer, and another preference vector is also sampled from the preference space. The TD target is calculated using the target network:

[0107]

[0108] The evaluation network is updated using equation (11) below, and its parameters θ are periodically copied to the parameters θ of the target network - :

[0109] L(θ) = E τ,a,ω [|ω T y-ω T Q tot (τ, a, ω, θ) |] (11)

[0110] In an exemplary embodiment, in the training phase, the preference space is divided into multiple subspaces, a preference vector is sampled in each round, and the sampling probability of the subspaces is dynamically adjusted according to the distribution of non-dominated solutions. In this step, in the training phase, a preference vector is sampled as input in each round, the entire preference space is divided into multiple subspaces, and the sampling probability of different subspaces is adjusted according to the current distribution of non-dominated solutions in the target space.

[0111] After training, the agent model will give the optimal action under different preference weights according to the scheduling environment, and obtain the Pareto optimal solution set for different targets (e.g., total completion time, movement cost, and space density) preferences, i.e., the optimal scheduling scheme.

[0112] After training by the algorithm in the above embodiment, the trained model is saved. When different scheduling problems need to be solved, for example, the number of jobs is different, the operation of the job is different, or the machine resource is different, etc., the trained intelligent agent model can be called to obtain the solution scheme under different target weights, for example, the solution scheme under the condition of preferring shorter time or smaller space density. The user can customize the weights of the time and space targets (each weight is between 0 and 1, and the sum of all weights is 1), and thus the optimal solution scheme under the corresponding target condition can be obtained. Compared with the traditional single-target job scheduling method, the method is more flexible, has a wider range of application, and can effectively solve the job scheduling problem under the space constraint.

[0113] To verify the technical feasibility and superiority of the multi-agent reinforcement learning method for space-constrained job scheduling proposed in the present disclosure, the present embodiment constructs a simulated flexible job shop scheduling problem scene. The scene strictly considers the space movement constraint of jobs between machines, specifically simulates a scheduling problem containing 10 jobs, each job containing 10 operations, and the processing time of the operations is uniformly randomly generated within 10 time units. The movement speed of the jobs between the machines is set to 5 distance units / time unit, the movement cost is directly proportional to the Euclidean distance passed, and the jobs can only move along a straight line or on a set path grid. The scene sets 21 machines, each machine is assigned a unique two-dimensional coordinate position and the machines have different capabilities, and certain operations can only be processed by a specific subset of machines.

[0114] The comparative method selects a rule-based scheduling method (i.e., selecting the nearest available machine), a heuristic method genetic algorithm, and a single-objective multi-agent reinforcement learning method. The evaluation indexes include total completion time, average movement time, and local density.

[0115] The present embodiment successfully trains and learns in the simulated space-constrained job scheduling environment and generates a feasible scheduling scheme (all jobs are finally completed), which verifies the technical feasibility. Compared with the rule-based scheduling method, the method according to the present embodiment can achieve comparable total completion time and movement time and lower local density, which can effectively reduce congestion and collision risk. Compared with the genetic algorithm, the method according to the present embodiment can achieve lower total completion time, movement time, and local density. Compared with the single-objective multi-agent reinforcement learning method, the method according to the present embodiment can flexibly adjust the weights of multiple targets and autonomously select whether to expect shorter time or lower local density.

[0116] The second aspect of the present disclosure provides a multi-agent reinforcement learning device 400 for space-constrained job scheduling, comprising: a job scheduling environment construction module 401 configured to construct a space-constrained job scheduling environment comprising an operation sequence of a job, resource configuration and spatial position of a machine, and a job moving speed; a Markov decision process modeling module 402 configured to model the space-constrained job scheduling problem as a multi-objective partially observable Markov decision process comprising a state space, a joint action space, an observation space of an agent, a state transition function, and a reward function; a multi-objective reward function design module 403 configured to design a multi-objective reward function comprising a total completion time reward, a moving cost reward, and a local density reward; and a centralized training and decentralized execution training module 404 configured to train parameters capable of obtaining a Pareto optimal solution to the space-constrained job scheduling problem by interacting with the job scheduling environment through a conditional agent network and a multi-objective hybrid network using a centralized training and decentralized execution framework.

[0117] In the description of the first aspect of the present disclosure, the various method steps involved in the technical solution of the present disclosure have been described in detail, therefore, the above description can be applicable to the multi-agent reinforcement learning device 400 for space-constrained job scheduling of the second aspect of the present disclosure, accordingly, no repeated description is given here.

[0118] The present disclosure also provides an electronic device comprising a memory and a processor, the memory storing a program, and the processor being configured to acquire the program and execute the above-described multi-agent reinforcement learning method for space-constrained job scheduling when executing the program.

[0119] Figure 5 is a block diagram of an electronic device implementing the multi-agent reinforcement learning method for space-constrained job scheduling according to some embodiments of the present disclosure. As Figure 5 shown, the multi-agent reinforcement learning method for space-constrained job scheduling in the above-described embodiments can be implemented by the electronic device as Figure 5 shown, the electronic device comprises at least one processor, a memory, and at least one I / O interface.

[0120] The processor can be a general-purpose central processing unit (CPU) and a graphics processing unit (GPU), or can be an application-specific integrated circuit (ASIC). The memory can include at least one of a volatile memory and a non-volatile memory. The memory can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions; can be a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions; can be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage, a magnetic disk storage, or other magnetic storage devices, or any other medium capable of storing instructions or data in the form of programs and accessible by a computer, but is not limited thereto. The memory can exist independently, and be connected to the processor through an address bus, a data bus, and a control bus. The memory can also be integrated with the processor.

[0121] The memory is used to store programs for implementing the schemes of the present disclosure, and is controlled by the processor to perform. The processor is used to execute the programs stored in the memory. The programs can include one or more software modules. The multi-agent reinforcement learning method for spatial constraint job scheduling in the above embodiments can be implemented by one or more software modules of the processor and the programs in the memory. However, the present disclosure is not limited thereto. The multi-agent reinforcement learning method for spatial constraint job scheduling in the above embodiments can also be implemented by a circuit.

[0122] The I / O interface is connected with input devices such as a mouse, a microphone, a keyboard, a touch screen, and the like, and output devices such as a speaker, a printer, a display, and the like. The I / O interface can also use any transceiver-like device for communicating with other devices or communication networks such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), and the like.

[0123] As an example embodiment, the electronic device can include a plurality of processors, each of which can be a single-core processor or a multi-core processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0124] The electronic device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation manner, the computer device can be a desktop computer, a notebook computer, a network server, a palmtop computer, a mobile phone, a tablet computer, a wireless terminal device, a communication device (e.g., an access point, a router, a gateway, and the like), or an embedded device. The embodiments of the present disclosure do not limit the type of computer device as long as it has a processor and a memory.

[0125] It should be understood that Figure 5 The electronic device shown is merely one example of the present disclosure, and the electronic device of the present disclosure can also include elements or components not shown in the above examples. For example, some electronic devices also include a display unit such as a display screen, some electronic devices also include human-computer interaction elements such as buttons, keyboards, etc., and some electronic devices also include various sensors such as gesture sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, acceleration sensors, grip sensors, proximity sensors, color sensors, infrared (IR) sensors, biometric sensors, temperature sensors, humidity sensors, illuminance sensors, etc. As long as the electronic device can execute a computer-readable program in the memory to implement the method or at least part of the steps in the method described in the present disclosure, it can be considered as an electronic device covered by the present disclosure.

[0126] Through the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, can be implemented by hardware, or can be implemented by a combination of software and necessary hardware. Therefore, as Figure 6 indicated, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-transitory computer-readable storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of commands to make a computing device (which can be a personal computer, a server, or a network device, etc.) execute the above-mentioned method according to the embodiments of the present disclosure.

[0127] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of readable storage media include, but are not limited to, an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0128] The present disclosure also provides a computer-readable storage medium, which stores a program, and the program is executed by a processor to implement the above-mentioned multi-agent reinforcement learning method for spatial constraint job scheduling. Figure 6 is a block diagram showing a computer-readable medium according to an embodiment of the present disclosure

[0129] The computer readable storage medium can include a computer-readable storage medium comprising a computer-readable storage medium that stores the program code. Such computer-readable storage medium or media can be non-transitory. The above general description of computer-readable storage medium has been provided for the purpose of clarity and illustration, and is not intended to be limiting of the computer-readable storage medium. Computer-readable storage medium can include any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable storage medium can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired computer program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable storage medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or twisted pair, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable storage medium.

[0130] In one or more embodiments, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over a computer-readable medium as one or more instructions or code. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. In this manner, a computer-readable medium can take many forms, including but not limited to, a tangible storage medium (e.g., random access memory (RAM), read only memory (ROM), a magnetic readable medium such as a hard drive, floppy diskette, or magnetic tape) or transmission medium (e.g., electrical, optical, microwave, laser, infrared, or any combination thereof). Combinations of the above should also be included within the scope of computer-readable media.

[0131] The above computer readable medium stores one or more programs (e.g., computer executable instructions), which when executed by one or more devices, cause the computer readable medium to implement the methods of the present disclosure.

[0132] The above description is only preferred embodiments of the present disclosure and the best technical principles used, and is not intended to limit the scope of the present disclosure, but only represents the preferred embodiments of the present disclosure. Those skilled in the art should understand that the scope of the application involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present disclosure.

Claims

1. A multi-agent reinforcement learning method for spatially constrained job scheduling, characterized in that, include: Construct a job scheduling environment that includes the operation sequence of the job, the resource configuration and spatial location of the machine, and the spatial constraints of the job's movement speed; The spatially constrained job scheduling problem is modeled as a multi-objective partially observable Markov decision process, which includes a state space, a joint action space, an agent's observation space, a state transition function, and a reward function. The design includes a multi-objective reward function that incorporates total completion time incentives, mobility cost incentives, and spatial density incentives. as well as By utilizing a centralized training and distributed execution framework, and through interaction with the job scheduling environment via a conditional agent network and a multi-objective hybrid network, parameters capable of obtaining Pareto optimal solutions to spatially constrained job scheduling problems are trained.

2. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 1, characterized in that, Constructing a job scheduling environment that includes spatial constraints such as the operation sequence of the job, machine resource allocation and spatial location, and job movement speed includes: Set a set of N jobs to be processed, each job having an operation sequence of P operations, where N and P are natural numbers greater than 1; Define a set of L machines, specifying the resource configuration and location for each machine, where L is a natural number greater than 1; and Set the movement speed of the operation.

3. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 2, characterized in that, The goal of job scheduling is to minimize the maximum completion time T of all jobs. total : in, This represents the waiting time for the i-th job before processing the p-th operation. This indicates that the i-th job starts from processing the p-th job. - One operating machine M i,p-1 The position is moved to the machine M that processes the p-th operation. i,p The time taken to move to a new location is obtained by calculating the quotient of the distance between the two machines and the speed at which the operation moves. Indicates the processing time of the p-th operation, and i is a natural number greater than or equal to 1 and less than or equal to N, and p is a natural number greater than or equal to 1 and less than or equal to P.

4. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 2, characterized in that, The state space includes the machine's occupancy status, the machine's remaining working time, the set of operations that the machine can process, the number of remaining operations for the job, the identifier of the next operation for the job, and the processing time of the next operation for the job. The combined action space includes the machine's selected action, waiting action, busy action, and completed action; as well as The agent's observation space includes the agent's current position, the machine's movement time, and the set of other tasks within the agent's observation range.

5. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 4, characterized in that, The total completion time bonus is inversely proportional to the maximum remaining time of all jobs during scheduling, and inversely proportional to the total completion time of all jobs after scheduling ends; The mobility cost reward is inversely proportional to the travel time from the agent's current location to the target machine; and The spatial density reward is inversely proportional to the occupancy rate of neighboring machines around the agent's current machine.

6. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 1, characterized in that, By utilizing a centralized training and distributed execution framework, and interacting with the job scheduling environment through a conditional agent network and a multi-objective hybrid network, the parameters trained to obtain a Pareto optimal solution to the space-constrained job scheduling problem include: The conditional agent network outputs a Q-value vector based on the agent's observations, historical actions, and preference vectors, including objectives such as total completion time, movement cost, and spatial density; and The multi-objective hybrid network uses parallel tracks to process the Q-value vectors of each object and combines the outputs of each track to generate a multi-objective Q-value vector.

7. The multi-agent reinforcement learning method for space-constrained job scheduling according to claim 6, characterized in that, During the training phase, the preference space is divided into multiple subspaces, a preference vector is sampled in each round, and the sampling probability of the subspace is dynamically adjusted according to the distribution of non-dominated solutions.

8. A multi-agent reinforcement learning device for spatially constrained job scheduling, characterized in that, include: The job scheduling environment construction module is configured to build a job scheduling environment that includes the operation sequence of the job, the resource configuration and spatial location of the machine, and the spatial constraints of the job's movement speed. The Markov decision process modeling module is configured to model the spatially constrained job scheduling problem as a multi-objective partially observable Markov decision process, wherein the multi-objective partially observable Markov decision process includes a state space, a joint action space, an agent's observation space, a state transition function, and a reward function. The multi-objective reward function design module is configured to design a multi-objective reward function that includes total completion time reward, movement cost reward, and spatial density reward. as well as The centralized training and distributed execution training module is configured to utilize the centralized training and distributed execution framework to interact with the job scheduling environment through a conditional agent network and a multi-objective hybrid network, thereby training parameters capable of obtaining Pareto optimal solutions to spatially constrained job scheduling problems.

9. An electronic device, comprising: processor; And a memory for storing a program that, when executed by the processor, performs the method as described in any one of claims 1-7.

10. A computer-readable medium storing a program that, when executed by a processor, performs the method as described in any one of claims 1-7.