Vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning

By employing a multi-agent deep reinforcement learning approach, a vehicle manufacturing painting resource scheduling model was constructed, which solved the efficiency and stability issues of painting resource scheduling in a cloud environment, and achieved efficient production in personalized body orders and multi-platform collaborative manufacturing.

CN120975515BActive Publication Date: 2026-02-06CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511485978.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-06
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently schedule vehicle manufacturing painting resources in a cloud environment, especially in personalized body orders and multi-platform collaborative manufacturing scenarios, where painting resource scheduling methods cannot meet the demands for high-smoothness and high-stability production.

Method used

A multi-agent deep reinforcement learning-based approach is adopted to construct a vehicle manufacturing painting resource scheduling model. The agents are trained using the KAN-MAPPO model, and the scheduling scheme is optimized by combining multi-stage resource scheduling and dynamic interaction mechanisms to minimize production energy consumption and completion time.

Benefits of technology

It enables efficient scheduling of vehicle manufacturing painting resources in a cloud environment, optimizes the production process, reduces production energy consumption and completion time, and improves production stability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975515B_ABST
    Figure CN120975515B_ABST
Patent Text Reader

Abstract

The application provides a kind of vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning, belongs to the technical field of artificial intelligence and intelligent manufacturing technology field.Its characterized in that the method first analyzes the characteristics of vehicle manufacturing coating production and dynamic resource attributes in cloud environment, and constructs a three-stage coating resource scheduling model considering manufacturing energy consumption and completion time;On this basis, through the dynamic interaction mechanism of multi-agent in cloud platform, the multi-agent proximal policy optimization (MAPPO) algorithm is used to solve the vehicle manufacturing coating resource scheduling problem in three stages;At the same time, in order to enhance the explainability of the scheduling strategy and ensure the convergence of the algorithm, KAN (Kolmogorov-Arnold Networks) is used as the neural network structure of the agent.The application is widely used in vehicle manufacturing coating production enterprises, and the proposed model and method can obtain a scheduling scheme that meets the energy consumption and completion time requirements, and can efficiently and independently handle resource maintenance, body back line and other dynamic events.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence and intelligent manufacturing technology, and particularly relates to a whole vehicle manufacturing coating resource intelligent scheduling method based on multi-agent deep reinforcement learning. BACKGROUND

[0002] Consumers' demand for personalized customized cars is growing, and the shortening of product life cycle has become the norm. At the same time, under the guiding ideology of networked collaboration, digital transformation and intelligent change, a new service-oriented cross-enterprise and multi-platform collaborative manufacturing mode based on industrial internet, cloud manufacturing and other technologies has gradually formed. This service-oriented manufacturing mode integrates the manufacturing resources of enterprises and collects the needs of consumers; at the same time, it also puts forward higher requirements for the production process management, coordination of specialized operation and optimization of resource allocation in the automotive industry, among which the coating resource scheduling in the automotive coating service is one of the main challenges to improve operational management. In particular, how to effectively utilize the huge resource pool of whole vehicle coating resources in the cloud environment to coordinate the heterogeneous resources of the coating production process to complete small-batch and multi-type body orders. As a typical heterogeneous manufacturing resource collaborative manufacturing scene under the industrial cloud environment, efficient and intelligent scheduling of coating resources through the cloud platform is the key to realizing high-flow and strong-stable whole vehicle manufacturing production. Compared with other processes, the coating process has strict sequence requirements, is easily affected by the previous treatment process, and has high energy consumption and heavy pollution in coating treatment. The scheduling process usually studies the coating resource scheduling method from the aspects of the number of color switching times, production energy consumption, pollution emission, etc. Therefore, full consideration is given to the characteristics of multiple production stages, multiple resource selection and multiple task types involved in the whole process of whole vehicle manufacturing coating production under the cloud environment, the whole vehicle manufacturing coating resource scheduling process is accurately described, and an intelligent and efficient scheduling scheme solving algorithm is designed in combination with the power distribution. SUMMARY

[0003] In order to solve the problems of the prior art, that is, the prior art is difficult to adapt to the individualized scene of whole vehicle manufacturing coating resources under the cloud environment, the present application provides a whole vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning, which fully considers the multi-stage characteristics and multi-agent interaction and cooperation requirements of the scheduling problem.

[0004] The technical scheme of the present application is:

[0005] Step S10, the whole vehicle manufacturing coating resource scheduling is induced into a hybrid flow shop scheduling problem, which involves first stage "pretreatment and electrophoretic coating" resource scheduling, second stage "spray coating process" resource scheduling, and third stage "final inspection and defect repair" resource grouping, and based on this, a hypothetical condition is designed;

[0006] Step S20, according to the service demand uploaded to the cloud platform by the resource demand side and the resource provider, a coating task model under the cloud environment is constructed, and the resources involved in modeling, such as "pretreatment and electrophoretic coating", "spray process", "final inspection and defect repair" and logistics transportation, are synchronized;

[0007] Step S30, based on the resource and task model constructed in S20, the target function of optimization scheduling is determined, which involves minimizing the maximum completion time , minimizing the production energy consumption ;

[0008] Step S40, an optimization scheduling KAN-MAPPO deep reinforcement learning model is established, KAN is used as the neural network structure of the agent, the representation method of the state space and the reward function are determined, and the KAN-MAPPO model is trained through the multi-agent dynamic interaction mechanism;

[0009] Step S50, based on the designed hypothesis conditions, target function and resource model, the trained KAN-MAPPO model is used to optimize the scheduling of the whole vehicle manufacturing coating resources.

[0010] Further, in step S10, the designed hypothesis conditions are as follows:

[0011] 1) All resources are in an available state at the initial time t=0;

[0012] 2) The first stage resource can process one vehicle body order at a time, the second stage resource processes a single vehicle body at a time, and the third stage resource can process multiple vehicle bodies at the same time, and the processing process of all stages has non-preemptive characteristics;

[0013] 3) The scheduling process only includes the above three stages, and the process differences caused by different vehicle body structures are not considered;

[0014] 4) Each vehicle body order contains multiple vehicle body spraying operations, and the vehicle body is the smallest non-disassembled unit of the order;

[0015] 5) The scheduling model only considers the energy consumption and time consumption of long-distance logistics transportation, and the in-plant logistics time and energy consumption are included in the processing time and energy consumption of each stage;

[0016] 6) The cloud platform scheduling resources are standardized virtual resources that can adapt to the spraying needs of all vehicle models on the platform, but the number of resources owned by each enterprise in different stages is different;

[0017] 7) According to the characteristics of each coating layer, it is assumed that the vehicle body that has completed spraying does not have the ability to transport long distances across enterprises;

[0018] 8) The logistics transportation process can be operated all day long, and there is no difference in day and night transportation efficiency;

[0019] 9) The second and third stage resources both contain equipment and buffer, and cannot be subdivided;

[0020] 10) The occurrence of paint defects in the painting process follows a uniform probability distribution, and the repair time of a single defect is a fixed value;

[0021] 11) The state of resource unavailability caused by unexpected factors follows a specific probability distribution.

[0022] Further, in step S20, the painting task model construction method in the cloud environment is as follows:

[0023] According to the characteristics of the vehicle manufacturing industry, the vehicle body order p in the cloud environment is modeled as a set , and the total number of vehicle body orders is , that is, ; wherein, , is the vehicle body c involved in the vehicle body order p, and the specific elements are as follows:

[0024]

[0025] Wherein, p is the vehicle body order identifier; is the initial position of the order; c is the vehicle body identifier; is the final delivery time required by the vehicle body order; is the color requirement required for vehicle body painting; is the current geographical position of the vehicle body.

[0026] Further, in step S20, the "pretreatment and electrophoretic coating" resource, the "spraying process" resource, the "final inspection and defect repair" resource, and the logistics transportation resource involved, the model construction method is as follows:

[0027] Modeling of the first stage "pretreatment and electrophoretic coating" resource :

[0028]

[0029] Wherein, i is the resource identifier; represents the speed of the resource in completing the order task, that is, the number of vehicle bodies completed by the first stage resource per unit time; represents the geographical position of the resource; represents the energy consumption data of the first stage resource when providing services, wherein , , respectively represent the energy consumption per unit time when the resource provides service, the energy consumption per unit time when the resource is idle, and the loading energy consumption of one car body; represents the scheduled car body sequence, and the car body sequence is arranged according to the processing order; is the probability of resource unavailability;

[0030] the second stage “spraying process” resource modeling:

[0031]

[0032] wherein j is the resource identifier; represents the time for the resource to complete the task, i.e. the time required for completing spraying of one car body; represents the preparation time for color change of the spraying resource; represents the resource identifier required for subsequent final inspection of the resource; represents the geographical location where the resource is located; represents the energy consumption data of the resource when providing service in the second stage, wherein respectively represent the energy consumption per unit time when the resource provides service, the energy consumption per unit time when the resource is idle, the spraying head cleaning energy consumption, and the loading energy consumption of one car body; represents the scheduled car body sequence, and the car body sequence is arranged according to the processing order; is the car body sequence in the buffer area, and the car body sequence is arranged according to the buffer area exit order; is the probability of resource unavailability;

[0033] the third stage “final inspection and defect repair” resource modeling:

[0034]

[0035] wherein k is the resource identifier; is the number of car bodies that can be processed by one batch of “final inspection and defect repair” resources; represents the time for the resource to complete the task, i.e. the time required for completing final inspection of one batch of car bodies; represents the time required for the resource to complete defect repair of one car body; represents the energy consumption data of the resource when providing service in the second stage, wherein respectively represent the energy consumption per unit time when the resource provides service, the energy consumption per unit time when the resource is idle, the paint surface repair energy consumption, and the loading energy consumption of one batch of car bodies; represents the scheduled car body sequence, and the car body sequence is arranged according to the processing order; is the car body sequence in the buffer area, and the car body sequence is arranged according to the buffer area exit order; is the probability of resource unavailability;

[0036] for the logistics transportation resource model , including a logistics distance matrix , transportation equipment speed , and unit distance logistics transportation energy consumption :

[0037]

[0038] wherein, ; u represents the number of geographical locations involved in the scheduling process; represents the distance from the a-th geographical location to the b-th geographical location; represents the time from the a-th geographical location to the b-th geographical location; it is worth noting that the distance and time back and forth between two geographical locations may not be the same, i.e. .

[0039] Further, in step S40, the KAN-MAPPO-based deep reinforcement learning model adopts Kolmogorov-Arnold Networks (KAN) as the neural network structure of the agent, and the specific neural network structure is as follows:

[0040] The neural network of each agent includes two types: Policy network and Value network. The Policy network predicts the next action according to the current state, and the Value network assists in training the Policy network. Assuming that the current state is , the data flow in the Policy network is as follows:

[0041]

[0042]

[0043] wherein, is an intermediate variable; is a SoftMax activation function; is a vector representing the availability of actions, with the same number of bits as the action space; represents the selection probability of the action, based on the probability distribution formed by to sample the action; and vary in size depending on different processing scenarios; represents a multi-layer KAN network, and takes the resource scheduling state as the input of the KAN network, which is specifically represented as:

[0044]

[0045] wherein, is a one-dimensional function matrix, i.e. ; are the input data size and output data size of the i-th layer of the KAN network, respectively; is composed of B-spline functions with trainable parameters, i.e. ; the structure of the Value network is similar to that of the Policy network, the only difference is that the Value network is used to evaluate the value obtained by taking action :

[0046]

[0047]

[0048] wherein, the last layer adopts a linear layer to ensure the accuracy of single variable output.

[0049] Further, in step S40, the state space is characterized as follows:

[0050] The state space describes the state of the painting resource scheduling environment through the global resource state , wherein, respectively represent the one-stage painting resource local state, the two-stage painting resource local state and the three-stage painting resource local state; wherein, ={ } includes the allocation matrix, the completion time matrix, and the energy consumption matrix of the one-stage painting resource; ={ } includes the allocation matrix, the completion time matrix, and the energy consumption matrix of the two-stage painting resource; ={ } includes the batch grouping matrix, the completion time matrix, and the energy consumption matrix of the three-stage painting resource. When the agent performs an action, the state representation is the parameter change of the corresponding resource-task (vehicle body) position.

[0051] Further, in step S40, the agent reward function is as follows:

[0052] The agent is trained by combining local rewards and global rewards. The local reward is the environmental feedback obtained by the agent after each decision is completed in the resource stage, and the global reward is the execution of all currently scheduled resources, which can directly reflect the global scheduling goal and facilitate the control of the training direction of the agent, wherein and are used to balance the influence of the two:

[0053]

[0054] 1) Global reward

[0055] The global reward is the opposite number of the total work completion time and the total energy consumption increment of all resources after the scheduling, and is weighted by the weight coefficient , The importance of the elements of the reward function is corrected:

[0056]

[0057] wherein the global work completion time increment ; the global production energy consumption increment ; and respectively represent the maximum work completion time and energy consumption of all resources after the agent performs an action at ; and respectively represent the maximum work completion time and energy consumption of all resources before the agent performs an action at ;

[0058] 2) Local reward

[0059] The local reward is the opposite number of the work completion time and the energy consumption increment of the resources of the production stage to which the agent belongs after the scheduling, and is weighted by the weight coefficient , The importance of the elements of the reward function is corrected:

[0060]

[0061] wherein i represents the production stage to which the agent belongs after the scheduling, i.e. , respectively represent the maximum work completion time and energy consumption increment of the i-stage resources after the scheduling at ; ; .

[0062] Further, in step S40, the multi-agent dynamic interaction mechanism is as follows:

[0063] Step S401: Enter the first stage "preprocessing and electrophoretic coating" resource scheduling: refresh the newly added car body order pool, select the orders one by one according to the priority, schedule the available one-stage resources for execution, and iterate through all car body orders until all car body orders are processed;

[0064] Step S402: enter the second stage "spraying process" resource scheduling: refresh the newly added car body task pool, select the tasks one by one according to the priority, schedule the available second-stage resources for execution, and iterate through all the car body tasks until all the car body tasks are completed.

[0065] Step S403: enter the third stage "final inspection and defect repair" resource group batch: select a three-stage resource, judge whether the number of car bodies in the buffer area of the resource meets the group batch condition, and if the condition is met, perform intelligent group batch until all the three-stage resources are iterated.

[0066] Step S404: refresh the buffer area state of all coating resources, and update the system time.

[0067] Step S405: judge whether all car body orders have been completed, if not, return to step S401 and execute in a loop. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 is a whole vehicle manufacturing coating resource scheduling framework based on multi-agent reinforcement learning;

[0069] Figure 2 is an example of a three-stage whole vehicle manufacturing coating resource scheduling environment in a cloud environment;

[0070] Figure 3 is a whole vehicle manufacturing coating resource scheduling flowchart based on a multi-agent dynamic interaction mechanism;

[0071] Figure 4 is a comparison of three-stage group batching with and without waiting time;

[0072] Figure 5 is an example of global state representation of the coating resource scheduling environment process;

[0073] Figure 6 is the convergence of the overall reward value and performance indicators under different neural networks;

[0074] Figure 7 is the scheduling optimization effect under different neural networks. DETAILED DESCRIPTION

[0075] The application will be described in further detail below with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application.

[0076] Example 1: The application proposes a whole vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning, which is used for cross-enterprise scheduling of whole vehicle manufacturing coating resources in a cloud environment, and the resource scheduling framework used is as shown in Figure 1 Under this framework, the specific steps of the method include:

[0077] Step S10: Based on the assumption condition, a scheduling environment is constructed, an example of which is shown in Fig. 1. Figure 2 As shown in the figure, the square icons represent work areas, denoted by M, and the circular icons represent buffer areas, denoted by B. It is assumed that each work area involves only one resource, and thus the problem can be described as: one car body passes through three stages in the same order; the first stage has one resource, i.e., a phosphating, electrophoresis, and other corrosion-resistant treatment operation; the second stage has m parallel production lines, each of which has one limited buffer area and one machine, for a total of m buffer areas and m parallel spraying devices for a spraying process; the third stage has one rework buffer area , one machine , and one buffer area ; it is worth noting that each product passes through all stages, including buffer areas and machines. In addition, the red arrows represent the spraying process after the redundant car bodies are transported to a different location; the brown arrows represent the treatment of car bodies to be sprayed after being transported to a different location due to factors such as damage to the spraying process equipment.

[0078] Step S20: Import certain whole vehicle manufacturing coating production example data into the car body order and each stage resource model. The case data set contains 20 items related to stages 1 and 3 resources and 65 items related to stage 2 resources; a total of 252 car bodies are coated for 21 orders.

[0079] Step S30: Based on steps S10 and S20, construct a multi-objective resource scheduling model, with the goal of minimizing the maximum completion time and minimizing production energy consumption The scheduling model is as follows:

[0080] ( 1 )

[0081] ( 2 )

[0082] S.t.

[0083] ( 3 )

[0084] ( 4 )

[0085] ( 5 )

[0086] ( 6 )

[0087] ( 7 )

[0088] ( 8 )

[0089] ( 9 )

[0090] ( 10 )

[0091] ( 11 )

[0092] ( 12 )

[0093] ( 13 )

[0094] wherein, is the completion time of the body order p; is the completion time of the first stage processing of the body order p; is the completion time of the second stage processing of the body c in the body order p; is the completion time of the third stage processing of the body c in the body order p; is the predecessor body order of the body order p is the completion time of the first stage processing; is the predecessor body of the body c in the body order p is the completion time of the second stage processing; is the predecessor body of the body c in the body order p is the completion time of the third stage processing; is the rework identifier of the body c in the body order p, ; is the spray gun color switch identifier of the body c in the body order p before spraying, ; is the production energy consumption required for the first stage processing of the body order p; is the production energy consumption required for the second stage processing of the body order p; is the production energy consumption required for the third stage processing of the body order p; is the number of processing resources that can be scheduled in the i-th stage; is the idle time of the manufacturing resource i in the j-th stage; is the logistics transportation energy consumption required for transporting the materials of the body order p in the first stage; is the logistics transportation energy consumption required for transporting the body c in the body order p in the second stage; is the first stage coating resource identifier required for completing the body order p; is the second stage coating resource identifier required for completing the body c in the body order p; To complete the third stage of painting resource identifiers required for the body c in the body order p. In addition, formulas (1) and (2) are the objective functions of the scheduling model: minimizing the maximum completion time, minimizing the energy consumption of painting production. Formulas (3)-(6) are used to calculate the maximum completion time of the scheduling scheme after the cloud platform scheduling; among them, formula (4) is the completion time of the body order p in the cloud platform, and it is limited that the body order p must be completed before the deadline ; formulas (5), (6) and (7) represent the completion time of the first, second and third stages of painting process of the body c of the body order p. Formula (8) is the total energy consumption of the scheduling scheme after the cloud platform scheduling, including the energy consumption of resource working time and the energy consumption of resource idle time; formulas (9), (10) and (11) represent the energy consumption of the three stages during the completion process of the body order p; in addition, the energy consumption of long-distance logistics transportation involved in the first and second stages is shown in formulas (12), (13).

[0095] Step S40: model training. Each training sample generation process is based on the constructed multi-agent dynamic interaction mechanism, as shown in Figure 3 . Among them, the three-stage batch problem needs to fully consider the waiting time of the three-stage resource processing body spraying final inspection, that is, the number of final inspection bodies of a batch after batching . Considering the waiting time in the dynamic environment can effectively improve the final inspection efficiency, as shown in Figure 4 . At time t, bodies 1 and 4 can be final inspected, but the subsequent arrival of bodies 3 causes the final inspection of bodies 1 and 4 to be interrupted, resulting in an increase in the delay time of bodies 2 and 5. Therefore, the time window triggering mechanism is introduced in the three-stage body batch problem, which can timely correct unreasonable waiting schemes, and the width of the time window is t0. The global resource state is constructed, and the representation method is shown in Figure 5 . In addition, the training process finds the optimal neural network parameters of each agent by constructing a strategy optimization model, so as to maximize the expected long-term return , which can be expressed as:

[0096] ( 14 )

[0097] Among them, is the state and action transition sequence of the scheduling decision in the th iteration period; represents the scheduling action selection strategy with the neural network parameters of agent k; represents the scheduling action selection strategy with the neural network parameters The action duration of the scheduling strategy. The PPO optimization model is composed of Actor and Critic neural networks, the Actor network receives the current state and outputs the corresponding decision, and the Critic network estimates the decision of the Actor network and feeds back the estimation result. The Actor and Critic networks perform training through policy gradient iteration to find the optimal parameters of each agent . In order to improve the adaptability and stability of PPO, the parameter update is limited within the confidence interval.

[0098] In the training process of updating the agent network model, each agent obtains the global resource state from the real-time scheduling model of coating resources to generate a random policy. Then, through the policy interaction between agents, the updated state and action transition sequence are obtained, and the Actor network and Critic network parameters of each agent are adjusted . Taking the neural network parameter as an example, the loss function Loss in the updating process of the Actor network parameter can be expressed as:

[0099] ( 15 )

[0100] Wherein, represents the truncation hyperparameter; represents the update amplitude of the neural network parameter, calculated by formula ( 16 ) ; represents the truncation function, responsible for limiting the update amplitude of the neural network within the interval to ensure convergence; is the advantage function at time , that is, under the state , the action adopted obtains the deviation of the reward value relative to the average value, calculated by formula ( 17 ).

[0101] ( 16 )

[0102] Wherein, and represent the updated and original Actor network parameters respectively. In addition, the advantage function at time is expressed as:

[0103] ( 17 )

[0104] Wherein, represents the single-step time difference error, which is calculated by formula ( 18 ) following the idea of time difference error; is a scaling factor of the advantage function.

[0105] ( 18 )

[0106] At the same time, the Critic network takes the mean square error of the estimated return as the loss function, and minimizes the loss function by updating the network parameters , so that the estimated return is more accurate. The Actor-Critic network uses an adaptive learning rate optimization algorithm for stochastic gradient iteration training to obtain the optimal network parameters . In addition, by deeply mining the parameter characteristics of the spline function obtained by training, the interpretability of the scheduling mechanism mined by the agent can be further improved. In order to verify the superiority of KAN-MAPPO in model training, it is compared with MAPPO using CNN, MLP network model, and the training effect is shown in Figure 6 . The convergence performance of KAN-MAPPO can quickly adapt to the multi-stage scheduling environment in the cloud environment, and the stability of the algorithm after training is significantly better, but there is a problem that the performance improvement speed slows down as the training deepens. Among them, the model training parameter settings are as follows: the number of iterations is 20000, the discount coefficient is 0.85; the learning rate of the Actor network is 10 -6 , the learning rate of the Critic network is 5x10 -6 ; the number of hidden layer neurons is 64, and the optimizer is Adam; the initial weight coefficient is set to =0.7, =0.3, and adjusted to =0.6, =0.0002 during training; the advantage function scaling factor is 0.95, and the PPO truncation coefficient is 0.2; the time window size is 2; the KAN model part settings include: the grid size is 3, the B-spline order is 3, the spline scale noise is 0.1, the truncation coefficient is 0.02, the grid parameter range is [-1, 1], and the activation function is SiLU.

[0107] Step S50: Perform optimized scheduling, and test the process using MAPPO with CNN and MLP network models. For each algorithm, 15 independent repeated experiments are performed, and the indicators of the solution and the reward value after scheduling are recorded, and the results are shown in Figure 7 . The results show that the intelligent scheduling method related to the present application can appropriately describe the state of resources and tasks from the model construction, and the proposed algorithm has good optimization performance and can well solve the scheduling scheme that meets the requirements.

Claims

1. A vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning, characterized in that, The method steps are as follows: Step S10, the vehicle manufacturing coating resource scheduling is summarized as a mixed flow shop scheduling problem, involving first stage "pretreatment and electrophoretic coating" resource scheduling, second stage "spray process" resource scheduling, third stage "final inspection and defect repair" resource grouping, based on this design assumption condition; Step S20, according to the service demand uploaded to the cloud platform by the resource demand side and the resource provider, a coating task model under the cloud environment is constructed, and the "pretreatment and electrophoretic coating" resource, "spray process" resource, "final inspection and defect repair" resource and logistics transportation resource involved in modeling are synchronized; Step S30, determining an objective function of the optimal scheduling based on the resource and task model constructed in S20, the objective function involving minimizing the maximum completion time , minimizing the production energy consumption ; Step S40, an optimization scheduling KAN-MAPPO deep reinforcement learning model is established, KAN is used as the neural network structure of the intelligent agent, the representation method of the state space and the reward function are determined, and the KAN-MAPPO model is trained through the multi-agent dynamic interaction mechanism; Step S50, based on the designed assumption condition, objective function and resource model, the trained KAN-MAPPO model is used to optimize and schedule the vehicle manufacturing coating resources; Further, in step S40, the multi-agent dynamic interaction mechanism is as follows: Step S401: enter the first stage "pretreatment and electrophoretic coating" resource scheduling: refresh the newly added body order pool, select the order one by one according to the priority, schedule the available one-stage resources for execution, and iterate through all body orders until all body orders are completed; Step S402: enter the second stage "spray process" resource scheduling: refresh the newly added body task pool, select the task one by one according to the priority, schedule the available two-stage resources for execution, and iterate through all body tasks until all body tasks are completed; Step S403: enter the third stage "final inspection and defect repair" resource grouping: select a three-stage resource, judge whether the number of body in the resource buffer meets the grouping condition, if yes, perform intelligent grouping, and iterate through all three-stage resources until all three-stage resources are completed; Step S404: refresh the buffer state of all coating resources and update the system time; Step S405: determine whether all body orders have been completed, if not, return to step S401 and execute in a loop.

2. The vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The assumption condition described in step S10 is as follows: 1) all resources are in an available state at the initial time t=0; 2) the first stage resource can process the whole body of one body order at a time, the second stage resource can process a single body at a time, the third stage resource can process multiple bodies at the same time, and the processing process of all stages has non-preemptive characteristics; 3) the scheduling process only includes the above three stages, and the process difference caused by different body structures is not considered; 4) each body order contains multiple spray jobs of bodies, and the body is the smallest non-disassembled unit of the order; 5) the scheduling model only considers the energy consumption and time consumption of long-distance logistics transportation, and the in-plant logistics time and energy consumption are included in the processing time and energy consumption of each stage; 6) the cloud platform scheduling resource is a standardized virtual resource that can adapt to the spray demand of all vehicle models of the platform, but the number of resources owned by each enterprise in different stages is different; 7) according to the characteristics of each coating, it is assumed that the body which has completed spraying does not have the ability to transport long distances across enterprises; 8) The logistics transportation process can realize all-weather operation, and there is no difference in transportation efficiency day and night; 9) The resources in the second and third stages include equipment and buffer zones and cannot be further subdivided; 10) The occurrence of paint defects in the coating process follows a uniform probability distribution, and the repair time of a single defect is a fixed value; 11) The state of resource unavailability caused by unexpected factors follows a specific probability distribution.

3. The vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The coating task model construction method in the cloud environment in step S20 is as follows: According to the characteristics of the whole vehicle manufacturing industry, the vehicle body order p in the cloud environment is modeled as a set , and the total number of vehicle body orders is , that is ; wherein , is the vehicle body c involved in the vehicle body order p, and the specific involved elements are as follows: ; wherein p is a body order identifier; is an initial location for the order; c is a body identifier; is a final delivery time required for the body order; is a color requirement needed for the body paint; is a current geographical location of the body; Further, in step S20, the model construction method of the resources involved in "pretreatment and electrophoretic coating", "spray process", "final inspection and defect repair" and logistics transportation is as follows: Resources for the first stage "pretreatment and e-coating" Modeling: ; wherein i is a resource identifier; represents the speed of the resource to complete the order task, i.e. the number of car bodies completed by the resource per unit time in a stage; represents the geographical location of the resource; represents the energy consumption data of the resource when providing services in a stage, wherein , , respectively represent the energy consumption per unit time when the resource provides services, the energy consumption per unit time when the resource is idle, and the loading energy consumption of one car body; represents the scheduled car body sequence, the car body sequence being arranged according to the processing order; is the probability that the resource is unavailable; The second stage "spray process" resource Modeling of: ; wherein j is a resource identifier; denotes the time for a resource to complete a task, i.e. the time for a resource to complete painting a car body; denotes the preparation time for a resource to change color for painting; denotes the resource identifier for a resource to be used for subsequent final inspection; denotes the geographical location of a resource; denotes the energy consumption data for a resource to provide service in the second stage, wherein denotes the energy consumption per unit time for a resource to provide service, the energy consumption per unit time for a resource to be idle, the energy consumption for cleaning a spray head, and the energy consumption for loading a car body, respectively; denotes the sequence of car bodies that have been scheduled, the sequence of car bodies being arranged according to the order of processing; denotes the sequence of car bodies in the buffer, the sequence of car bodies being arranged according to the order of leaving the buffer; denotes the probability of a resource being unavailable; Modeling of the resources for the third phase "Final Inspection and Defect Repair" : ; where k is a resource identifier; the number of car bodies that can be processed by a batch of "final inspection and defect repair" resources; the time for a resource to complete a task, i.e., the time required to complete a batch of car bodies final inspection; the time for a resource to complete a car body defect repair; the energy consumption data of the two-stage resource when providing services, wherein respectively represent the unit time energy consumption of the resource when providing services, the unit time energy consumption of the resource when idle, the paint repair energy consumption, and the loading energy consumption of a batch of car bodies; represents the sequence of car bodies that have been dispatched, and the sequence of car bodies is arranged according to the processing order; represents the sequence of car bodies in the buffer, and the sequence of car bodies is arranged according to the order of exiting the buffer; represents the probability that the resource is unavailable; For a logistics transportation resource model wherein a logistics distance matrix is included , a transportation equipment speed and a unit distance logistics transportation energy consumption : ; wherein, ; u represents the number of geographical locations involved in the scheduling process; represents the distance from the athgeographical location to the bthgeographical location; represents the time from the athgeographical location to the bthgeographical location; it is worth noting that the distance and the time back and forth between two geographical locations can not be the same, i.e. .

4. The vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, In step S40, the KAN-MAPPO-based deep reinforcement learning model uses KAN as the neural network structure of the agent, and the specific neural network structure is as follows: The neural network of each agent includes two kinds: a policy network and a value network. The policy network predicts the next action according to the current state, and the value network assists in training the policy network. Assuming that the current state is The data flow in the policy network is: ; ; where, is an intermediate variable; is a SoftMax activation function; is a vector representing the availability of actions, with the same number of bits as the action space; represents the selection probability of actions, based on the probability distribution formed by sampling actions; and The size of varies with different processing scenarios; represents a multi-layer KAN network, and the resource scheduling state is input to the KAN network, specifically represented as: ; wherein, is a one-dimensional function matrix, i.e. ; are the input data size and the output data size of the i-th layer of the KAN network, respectively; is composed of B-spline functions with trainable parameters, i.e. ; The structure of the Value network is similar to that of the Policy network, with the only difference being that the Value network is used to evaluate the value obtained by taking an action : ; ; The last layer uses a linear layer to ensure the accuracy of single variable output.

5. The vehicle manufacturing coating resource scheduling method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The agent reward function described in step S40 is as follows: The intelligent agent is trained by combining local rewards and global rewards. The local reward is the environmental feedback obtained by the intelligent agent at the stage of the resource after each decision is completed. The global reward is the execution of all the currently scheduled resources, which can directly reflect the global scheduling target and facilitate the control of the training direction of the intelligent agent. Among them And are used to balance the influence of both. ; 1) Global reward : The global reward is the opposite number of the total makespan and the incremental energy consumption of all resources after the scheduling, and is multiplied by a weight coefficient , Importance correction is made to the elements of the reward function: ; wherein the global makespan increment ; the global production energy consumption increment ; and respectively represent the maximum makespan and energy consumption of all resources after the agent performs an action at time ; and respectively represent the maximum makespan and energy consumption of all resources before the agent performs an action at time ; 2) Local rewards : The local reward is the inverse of the increase in completion time and production energy consumption generated by the resources of the production stage to which the agent belongs after this scheduling, and is determined by weighting coefficients. , Importance correction is applied to the elements of the reward function: ; Wherein, i represents the production stage to which the agent belongs after the scheduling, that is , are the maximum completion time and energy consumption increment of the i-stage resource after scheduling at time ; ; .

Citation Information

Patent Citations

  • TSN-5G train communication network asynchronous scheduling method based on multi-agent reinforcement learning

    CN119485218A

  • RIS-assisted multi-user communication beam forming optimization method based on KAN-SAC

    CN120528480A