Construction method and application of aviation assembly scheduling optimization model

By building assembly and disassembly drawings and training agents, and optimizing aviation assembly strategies, the existing aviation assembly scheduling problem is solved, and the aircraft assembly efficiency is improved and resource saving is achieved.

CN120471326APending Publication Date: 2025-08-12HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510412250.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing aviation assembly scheduling methods rely on on-site expert decisions, resulting in insulated scheduling efficiency. The existing machine learning methods are overloaded in computational burdens or difficult to converge quickly in complex environments, and cannot effectively improve production efficiency.

Method used

Build an air assembly scheduling optimization model, and by constructing assembly dissection graphs, extracting state feature vectors and inputting them into agents, training agents using reinforcement learning and reward shaping mechanisms, optimizing assembly strategies, and using policy neural networks and value neural networks to update parameters to minimize the number of idle air and total assembly time.

Benefits of technology

It improves aircraft assembly efficiency, reduces expert scheduling resources, promotes the efficiency and economicalization of the aviation assembly system, and solves the problem of slow learning convergence of agents in complex aviation assembly scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471326A_ABST
    Figure CN120471326A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of aeronautical manufacturing, and discloses a construction method and application of an aeronautical assembly scheduling optimization model, and the method comprises the steps: according to the number of aircrafts in a batch of production orders, the time span of each assembly link, the number of fixtures, the assembly sequence relation constraint of each aircraft, and the assembly sequence constraint of each aircraft on the fixtures, determining the number of the aircrafts in the batch of production orders; constructing an assembly disjunction graph of the batch production order aviation assembly environment; aviation assembly environment information is extracted from the assembly disjunction graph to form a state feature vector, the state feature vector is input into an intelligent agent, the action output by the intelligent agent is a scheduling strategy of each aircraft assembly link, and a reward value corresponding to each scheduling strategy is obtained; collecting a state feature vector, a scheduling strategy and a corresponding reward value of each airplane as a training data set; and training an intelligent agent by using the training data set, and taking the trained intelligent agent as an aviation assembly scheduling optimization model to output an optimal scheduling strategy of each batch of orders. The aircraft assembling efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of aviation manufacturing technology, and more specifically, relates to a construction method and application of an aviation assembly scheduling optimization model. Background Art

[0002] The modern aviation industry plays a vital role in driving technological innovation, global trade, and economic prosperity. Today, the aerospace manufacturing sector is increasingly pursuing production and assembly efficiency to steadily enhance national defense capabilities and production supply capacity while ensuring both quality and quantity. Aviation assembly, which involves the assembly and integration of complex aircraft components, accounts for 60%-70% of the total aviation manufacturing workload and is the core of aviation manufacturing, directly determining overall aircraft assembly efficiency and production effectiveness. However, due to its compactness, complexity, timeliness, and collaborative nature, most aviation assembly scheduling currently relies on on-site expert decision-making and scheduling. On-site scheduling and decision-making often require significant time, and the scarcity of experts is inconsistent with today's high-efficiency production. Therefore, designing aviation assembly scheduling optimization models is crucial to address the scarcity of scheduling resources and improve production and assembly efficiency.

[0003] Numerous papers have attempted to utilize machine learning methods to model and solve aviation assembly scenarios, such as GA and NSGA-II. GA focuses on optimizing a single objective, such as minimizing completion time, while NSGA-II simultaneously optimizes multiple objectives, such as balancing workstation load and cycle time. However, these algorithms are prone to becoming trapped in local optima due to their high parameter sensitivity and excessive computational overhead as the problem scale increases.

[0004] Another popular approach for finding optimal aviation assembly scheduling strategies is reinforcement learning. Reinforcement learning, which relies on data to determine optimal strategies without complex models, transforms assembly environment information into state feature vectors and feeds them into an intelligent agent. The agent then interacts with the environment to search for optimal decisions. However, existing methods typically rely on manual characterization of the assembly environment, which can lead to low assembly efficiency in complex environments with large production runs. Furthermore, existing reward mechanisms hinder rapid convergence of the intelligent agent in complex aviation assembly scenarios. Summary of the Invention

[0005] In response to the above-mentioned deficiencies or improvement needs of the prior art, the present invention provides a method for constructing an aviation assembly scheduling optimization model and its application, the purpose of which is to improve aircraft assembly efficiency.

[0006] To achieve the above objectives, the present invention provides a method for constructing an aviation assembly scheduling optimization model, comprising:

[0007] Based on the number of aircraft in a batch of production orders, the time span of each assembly link, the number of jigs, the assembly sequence constraints of each aircraft, and the assembly sequence constraints of each aircraft on the jig, an assembly disjunction diagram of the aviation assembly environment of the batch of production orders is constructed;

[0008] The aviation assembly environment information of the batch of production orders is extracted from the assembly disjunctive graph to form the state feature vector s t , and input it into the agent, the agent outputs action a t The scheduling strategy for each aircraft assembly link in this batch of production orders is obtained, and the reward value r corresponding to each scheduling strategy is obtained. t ; Among them, the reward function is to minimize the number of idle frames; the action a t Feedback to the aviation assembly environment to update the assembly disjunction graph and obtain the state feature vector s at the next moment t+1 , the tuple {s t ,a t ,r t ,s t+1}Store in the experience pool as a training data set;

[0009] The training data set is used to train the intelligent agent, the network parameters of the intelligent agent are reversely adjusted, and the trained intelligent agent is used as the aviation assembly scheduling optimization model.

[0010] Furthermore, the assembly disjunctive graph is: G = (V, C∪D), where the symbol ∪ represents taking a union;

[0011] Among them, the node V in the assembly disjunctive graph G is used to represent an assembly node for each aircraft. The attribute information of the assembly node includes: the serial number of the aircraft in the batch of production orders, the completion rate of the assembly link at the assembly node, the time span of the assembly link, and the link waiting for assembly; C is a directed connection arc formed by two adjacent assembly links of each aircraft, and D is an undirected connection arc formed by two adjacent assembly links of a jig.

[0012] Furthermore, the reward value is the reward value after reward plasticity, and the reward function after reward plasticity is for:

[0013]

[0014] Where r(s,a) represents the original reward function, and n is the number of jigs currently being assembled, F total is the total number of frames, s is the state feature vector of the input agent at different times, a is the action of the agent at different times; w φ (s) is the weight function, which is updated as follows:

[0015]

[0016] Where, I(w φ ) represents the updated weight function, the symbol := represents “defined as”; E[·] represents the expected operation, S represents the set of state feature vectors at all times; α is a constant related to the aviation assembly environment, τ T It is the trajectory obtained by the agent in this round before time t+1, and They are respectively in the trajectory τ T The total number of times the agent visits all states S and states s, and They are respectively in the trajectory τ T The total normalized reward for the agent to visit all states S and state s is calculated as:

[0017]

[0018] Where T is the end time of an episode of reinforcement learning; and Respectively represent the upper and lower limits of the reward value after reward shaping; s t is the state feature vector of the input agent at the current time t, a t is the action of the agent at the current moment t.

[0019] Furthermore, the intelligent agent includes: a policy neural network and a value neural network;

[0020] The strategy neural network is used to obtain the state feature vector s at the current moment based on the assembly disjunctive graph. t , select the corresponding action a from the action space t As output, and the action a t Feedback to the aviation assembly environment to update the current assembly disjunction graph and obtain the state feature vector s at the next moment t+1 ; wherein the action a t It is used to represent the scheduling strategy of the aircraft assembly link at the current moment. The strategy neural network selects the corresponding action a from the action space based on the current value function Q(s,a;θ) t As output;

[0021] The value neural network is used to determine the state feature vector s at the next moment based on the state feature vector s at the next moment. t+1 , output the estimate y of the value function Q(s,a;θ) t , to assist in updating the parameters θ of the policy neural network.

[0022] Furthermore, the action space distribution of the strategy neural network is:

[0023] A={LOR,MOR,SPT,LPT,LTPT,STPT,FIFO,LIFO}

[0024] Among them, LOR represents the maximum order ratio scheduling rule, MOR represents the minimum order ratio scheduling rule, SPT represents the shortest processing time scheduling rule, LPT represents the longest processing time scheduling rule, LTPT represents the longest total processing time scheduling rule, STPT represents the shortest total processing time scheduling rule, FIFO represents the first-in-first-out scheduling rule, and LIFO represents the first-in-last-out scheduling rule.

[0025] Furthermore, during the agent training process, the loss function of the policy neural network is the output y of the value neural network t The root mean square error between the loss function Q(s, a; θ) and the value function Q(s, a; θ); based on the loss function, the parameters θ of the policy neural network are updated using gradient back propagation;

[0026] The value of the neural network parameter θ - The update method is: θ - ←θ, which means that the policy neural network copies its parameters θ to the value neural network at every certain time step.

[0027] The present invention also provides an aviation assembly scheduling optimization method, comprising:

[0028] Based on the number of aircraft in a batch of production orders, the time span of each assembly link, the number of jigs, the assembly sequence constraints of each aircraft, and the assembly sequence constraints of each aircraft on the jig, an assembly disjunction diagram of the aviation assembly environment of the batch of production orders is constructed;

[0029] The aviation assembly environment information of the batch of production orders is extracted from the assembly disjunctive graph to form a state feature vector, and the state feature vector is input into the aviation assembly scheduling optimization model constructed by any of the above-mentioned aviation assembly scheduling optimization model construction methods. The aviation assembly scheduling optimization model outputs the scheduling strategy for each aircraft assembly link in the batch of production orders.

[0030] The present invention also provides an electronic device comprising a computer-readable storage medium and a processor;

[0031] The computer-readable storage medium is used to store executable instructions;

[0032] The processor is configured to read the executable instructions stored in the computer-readable storage medium to execute any one of the aforementioned methods for constructing an aviation assembly scheduling optimization model, or / and to execute the aforementioned aviation assembly scheduling optimization method.

[0033] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for constructing an aviation assembly scheduling optimization model as described in any one of the above items, or / and implements the aviation assembly scheduling optimization method as described above.

[0034] The present invention also provides a computer program product, comprising a computer program. When the computer program is run on a computer, the computer executes any one of the above-mentioned methods for constructing an aviation assembly scheduling optimization model, or / and executes the above-mentioned aviation assembly scheduling optimization method.

[0035] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0036] (1) The method of the present invention constructs an assembly disjunctive graph of the aviation assembly environment for a batch of production orders based on the number of aircraft in the production order, assembly time, number of jigs, assembly order constraints for each aircraft, and assembly order constraints for each aircraft on the jig. The aviation assembly environment information for the batch of production orders is directly extracted from the assembly disjunctive graph to form a state feature vector input by the intelligent agent. The intelligent agent interacts with the aviation assembly environment and outputs a scheduling strategy for the aircraft assembly process. Compared to the method of manually extracting the state features of the aviation assembly environment, the method of the present invention can quickly obtain the state features of the aviation assembly environment from the assembly disjunctive graph, and can improve assembly efficiency in complex environments with large production quantities. The intelligent agent designed by the present invention defines the scheduling strategy of the specific assembly link of the aircraft as an action function, and performs subsequent training on the intelligent agent based on the state feature vector, scheduling strategy and corresponding reward value obtained after multiple rounds of learning of the intelligent agent as a training data set. When the intelligent agent converges, since the intelligent agent uses minimizing the number of idle jigs as the reward function, it indirectly achieves the shortest total assembly time of the production order. When the obtained aviation assembly scheduling optimization model is used to make assembly scheduling strategy decisions, the production and assembly efficiency of the aircraft can be further improved, the resources of expert scheduling can be saved, and the efficiency and economy of the aviation assembly system can be promoted.

[0037] (2) As a preferred method, by introducing a reward shaping mechanism, the total number of times the agent visits all states S and states s under the reinforcement learning trajectory τ and And the total normalized reward of the agent visiting all states S and states s under trajectory τ and Update the weight function w φ(s), these parameters are closely related to different aviation assembly environment states, so that the weight function can distinguish the importance of different aviation assembly environment states and the degree of influence of each state s on key scheduling nodes, thereby guiding the intelligent agent to pay more attention to key scheduling nodes, reducing the number of rounds of intelligent agent training, and effectively solving the problem of slow learning convergence of intelligent agents in complex aviation assembly scenarios, further improving production and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of a method for constructing an aviation assembly scheduling optimization model provided in one embodiment of the present invention;

[0039] Figure 2 A schematic diagram of the interaction between an assembly environment and an intelligent agent provided in one embodiment of the present invention;

[0040] Figure 3 A structural diagram of an assembly environment and an intelligent agent model provided in one embodiment of the present invention;

[0041] Figure 4 This is a flowchart of an aviation assembly scheduling optimization model training provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0043] Example 1

[0044] like Figure 1 As shown, an embodiment of the present invention provides a method for constructing an aviation assembly scheduling optimization model, comprising:

[0045] Based on the number of aircraft in a batch of production orders, the time span of each assembly link, the number of jigs, the assembly sequence constraints of each aircraft, and the assembly sequence constraints of each aircraft on the jig, an assembly disjunctive graph G of the aviation assembly environment of this batch of production orders is constructed;

[0046] Extract the aviation assembly environment information of this batch of production orders from the assembly disjunctive graph G to form the state feature vector s t , and input it into the agent, the agent outputs action a t The scheduling strategy for each aircraft assembly link in this batch of production orders is obtained, and the reward value r corresponding to each scheduling strategy is obtained. t; Among them, the reward function is to minimize the number of idle frames; action a t Feedback to the aviation assembly environment to update the current assembly disjunction graph and obtain the state feature vector s at the next moment t+1 , the tuple {s t ,a t ,r t ,s t+1}Store in the experience pool as a training data set;

[0047] The training data set is used to train the intelligent agent, and the parameters of the intelligent agent are reversely adjusted. The trained intelligent agent is used as an aviation assembly scheduling optimization model.

[0048] In the embodiment of the present invention, a plurality of aircrafts in a batch of production orders P={P1, P2, ..., P P Each aircraft has several assembly links, such as fuselage structure assembly, wing assembly, fuel system assembly, and cabin interior assembly, and the time span of each assembly link is T ijk (i.e. the time k it takes to assemble aircraft i on jig j) is different. Specific assembly links are required to be carried out in a specific assembly environment, i.e., the number of jigs is generally expressed as F = {F1, F2, ..., F f}, therefore, the assembly sequence of an aircraft on the jig can be defined as P i ={P i1 ,P i2 ,...,P if}, P ij It means that aircraft i is assembled on jig j. Due to some other special requirements, there will be some subtypes among these aircraft. Some of their assembly links are different from the original aircraft, so the assembly time on the corresponding jig will also change.

[0049] Due to the existence of a fixed assembly sequence, the complete assembly of an aircraft must follow several rules to ensure the consistency and rationality of the sequence. The rules and constraints involved in assembly scheduling include:

[0050]

[0051] Among them, A ij Represents a certain aircraft P i The time when the jth (j∈1, 2, …, f) assembly link starts, T ij is the time when the jth assembly link is completed, x ijk Used to determine the assembly link P at time t ij Is it possible to j Assemble on, x ijk is a 0-1 variable, at time t, the assembly link P ij Can be in the frame Fj When the upper assembly is completed, it is 1, otherwise, it is 0; y iji′j′ Used to determine the link P at time t ij Is it possible to i′j′ Previously on the frame F j Assemble on, also a 0-1 variable; N represents an infinite positive number; F ij Indicates that the assembly step P can be executed ij The jig set of this process, for example, the last assembly process f of aircraft P1 and aircraft P2 can be assembled on jig F1 or F2, then F 1f ={F1,F2},F 2f ={F1,F2}. Therefore, the meanings of the four constraints above (the assembly order constraint for each aircraft and the assembly order constraint for each aircraft on the jig) are: an aircraft being assembled on a jig cannot be interrupted by the assembly tasks of other aircraft; multiple aircraft may be assembled on a given jig, and these tasks must be independent and indivisible; and the assembly order of an aircraft must be strictly maintained.

[0052] After obtaining the above-mentioned assembly information and constraint conditions, the environment in which the agent learns will construct an assembly disjunctive graph for this production order based on these assembly information and constraint conditions. The disjunctive graph G can be expressed as G = (V, C∪D), where the symbol ∪ represents a union. The node V in the disjunctive graph G is used to represent an assembly node for each aircraft. The attribute information of the node includes: the serial number of the aircraft, the completion rate of the assembly link at the assembly node, the time span of the assembly link, the link waiting for assembly, and other information; C is a directed connection arc formed by two adjacent assembly links of each aircraft, and D is an undirected connection arc formed by two adjacent assembly links of a jig. In an embodiment of the present invention, after initializing a disjunctive graph G, various attributes of the assembly process are initialized based on the existing information, and all the information of the node is returned, thereby obtaining a non-connected graph containing node feature information, namely a disjunctive graph.

[0053] Preferably, the agent in the embodiment of the present invention includes a policy neural network and a value neural network. When the agent starts learning, it extracts the aviation assembly environment information of the batch of production orders from the assembly disjunctive graph G constructed above, and combines them into state features for learning by the policy neural network and the value neural network. Figure 2 As shown, the agent obtains the state feature vector s at the current time t from the current assembly environment t After that, according to its own learning experience and the current value function Q(s,a;θ), it selects and executes an action a t(represents the scheduling strategy of the aircraft assembly link at the current moment), which is fed back to the current assembly environment, and the assembly disjunction graph G of the current aviation assembly environment is updated, thereby obtaining the state feature vector s at the next moment. t+1 .

[0054] In the embodiment of the present invention, the distribution of the action space of the policy neural network is:

[0055] A={LOR,MOR,SPT,LPT,LTPT,STPT,FIFO,LIFO}

[0056] According to the value of the current value function Q(s,a;θ), the action output by the policy neural network is one of the existing heuristic scheduling rules, namely the first-in-first-out rule (FIFO), the first-in-last-out rule (LIFO), the maximum order ratio rule (LOR), the minimum order ratio rule (MOR), the shortest processing time rule (SPT), the longest processing time rule (LPT), the shortest total processing time rule (STPT), and the longest total processing time rule (LTPT).

[0057] The aviation assembly environment feeds back the reward of this action to the agent (strategy neural network) for its subsequent learning. At the same time, the information of the aviation assembly environment changes to the next state s t+1 , tuple {s t ,a t ,r t ,s t+1} exists in the agent's experience pool.

[0058] The input of the value neural network is the state feature vector s of the next moment obtained from the experience pool t+1 , the output is an approximate estimate of the value function Q(s,a;θ) of the policy neural network, which is used to assist the update learning of the parameters θ of the policy neural network; where the value function Q(s,a;θ) is defined as:

[0059]

[0060] Where E[·] represents the expected operation; γ is the discount coefficient, and π represents the strategy of the agent (strategy neural network). represents the reward value obtained after the reward shaping mechanism is applied to the reward value; s is the state feature vector of the input agent at different times, s t is the state feature vector of the input agent at the current time t, a is the action of the agent at different times, a t is the action of the agent at the current moment t, T is the end time point of an episode, that is, from (t=1) to (t=T), and θ is the parameter of the policy neural network.

[0061] The data information in the batch production order is abstractly mapped to a disjunctive graph, and then the node information is extracted from the disjunctive graph to form the state features. The intelligent agent understands and analyzes this information and makes corresponding scheduling operations. The assembly environment then gives the reward value for this scheduling operation. The reward function r is defined as:

[0062]

[0063] Among them, n represents the number of jigs currently being assembled, F total Indicates the total number of jigs, i.e. the aforementioned f, F j Indicates the jig currently undergoing assembly. In this embodiment of the present invention, the number of aircraft in the production order, assembly time, and the number of existing jigs are used as information input and abstracted into state features for neural network learning. The specific assembly phase scheduling strategy for the aircraft is defined as an action function, and minimizing the number of idle jigs is used as a reward function. This indirectly minimizes the total assembly time and improves aircraft production and assembly efficiency.

[0064] On this basis, the reward plasticity mechanism is used to shape the existing reward function to accelerate the convergence speed of the agent during learning. In the embodiment of the present invention, the reward after shaping is designed. for:

[0065]

[0066] Among them, r(s,a) represents the original reward function, w φ (s) is the weight function, which is updated as follows:

[0067]

[0068] Among them, I(w φ ) represents the updated weight function, the symbol := represents “defined as”; E[·] represents the expected operation, S represents the set of state feature vectors at all times; α is a constant related to the aviation assembly environment, τ T It is the trajectory obtained by the agent in this episode before time t+1. The trajectory data records the entire process of the agent's interaction with the aviation assembly environment, reflecting the agent's "path" or "trajectory" within a certain period of time, specifically including the sequence of state, action, reward, and next state. and are the total number of visits to all states S and state s under the reinforcement learning trajectory τ, and are the total normalized rewards for visiting all states S and state s under trajectory τ, namely:

[0069]

[0070]

[0071] in, and They represent the upper and lower limits of the reward value after the reward shaping mechanism.

[0072] Finally, the state input data, scheduling strategy and reward value of the current batch are collected to obtain the training data set For subsequent training of the intelligent agent, the trained intelligent agent is used as an aviation assembly scheduling optimization model. The samples in the training process are the state feature vectors, scheduling strategies and corresponding reward values at different times. Among them, the intelligent agent can obtain the following information after each round of training: When applied, the aircraft and jig data of each batch of orders are input as states into the trained aviation assembly scheduling optimization model, so that it can output the optimal scheduling strategy for each batch of orders.

[0073] The training data set is used to train the intelligent agent. During the training process, the intelligent agents alternately update their respective network parameters through the primal-dual method.

[0074] In the embodiment of the present invention, a batch of production orders contains 10 aircraft, and each aircraft has 10 assembly link data as the training data set. Specifically, the model structure of the scheduling agent is as follows: Figure 3 As shown. The policy neural network uses the mean square error loss function (MSE), which is the approximate estimate y of the value function Q(s,a;θ) output by the value neural network. t The root mean square error loss function between the value function Q(s,a;θ) of the policy neural network, the parameters θ of the policy neural network are updated by gradient back propagation, and the update formula of the policy neural network parameters is:

[0075]

[0076] Among them, L(θ) is the loss function of the policy neural network, E[·] represents the expected operation, Q(s,a;θ) is the value function of the policy neural network, and y t is the output of the value network, which represents the approximate estimate of the value function by the value network. a′ is the action output by the agent at the next moment, s t+1 is the state feature vector of the next moment input, γ is the discount coefficient; s′ represents the state feature vector s at the next moment of s t+1 , is the gradient descent operation, η is the learning rate. The parameters of the value neural network are θ - , the formula for updating the value neural network parameters is: θ -←θ, where every c steps, the policy network copies its parameters to the value network.

[0077] The overall update training process is as follows Figure 4 As shown, it mainly includes the following three steps:

[0078] Step 1: Initialization. First, load data such as the number of aircraft, number of assembly links, and assembly time into the algorithm. Set the hyperparameters related to static training and randomly set the initial parameters of the policy neural network and the value neural network. The memory buffer is also initialized to store the transformation tuple {s t ,a t ,r t ,s t+1}.

[0079] Step 2: Interaction. In the embodiment of the present invention, the algorithm runs for a total of 800 rounds. In each time period t of each round, the aviation assembly environment combines its current state information to construct a disjunctive graph, extracts node information in the disjunctive graph to form state feature vectors, and converts these state feature vectors into state feature vectors. t Input into the agent's policy neural network, the policy neural network will output a specific scheduling plan a t , and feed it back to the aviation assembly environment. Then calculate the agent's reward value r t And get new reward values through the reward shaping mechanism The assembly environment moves to the next state s t+1 Finally, the tuple {s t ,a t ,r t ,s t+1} is stored in the memory buffer.

[0080] Step 3: Update. After an interaction is complete, the data in the buffer is used to calculate the loss function and update the policy and value neural networks via gradient backpropagation. All network parameters are updated over M episodes of a round (a complete interaction of the agent in an environment, from start to finish, i.e., starting from the initial state and reaching the final state after a series of actions). After the update is complete, the state is reset, the buffer is cleared, and the next round begins.

[0081] Example 2

[0082] An embodiment of the present invention provides an aviation assembly scheduling optimization method, comprising:

[0083] Based on the number of aircraft in each batch of production orders, the time span of each assembly link, the number of jigs, the sequential relation constraints of each aircraft assembly, and the assembly sequence of each aircraft on the jig, an assembly disjunctive graph G of the aviation assembly environment of this batch of production orders is constructed;

[0084] The aviation assembly environment information of the batch of production orders is extracted from the assembly disjunctive graph G to form a state feature vector, which is input into the aviation assembly scheduling optimization model to obtain the scheduling strategy for each aircraft assembly link in the batch of production orders; wherein, the aviation assembly scheduling optimization model is constructed by the construction method of the aviation assembly scheduling optimization model in Example 1.

[0085] The relevant technical solutions are the same as above and will not be repeated here.

[0086] The aviation assembly scheduling optimization method provided by the present invention was numerically simulated on a real data set of an assembly scheduling system consisting of ten aircraft. The method of the present invention (DQN-ESW) was compared with the GA and NSGA-II algorithms. The results are shown in Table 1 below.

[0087] Table 1 Experimental results of the proposed method and other algorithms on actual data sets

[0088]

[0089] Assuming an aviation assembly environment free of external interference, the proposed method achieves significantly lower mean and variance (Std) values than the other two methods. This demonstrates that the proposed method achieves superior scheduling results from small to large scheduling scales, particularly at large scales, achieving shorter total assembly times and exhibiting a relatively stable and smaller deviation (Std) in the resulting scheduling strategies. Therefore, under the same basic conditions as the other methods, this model achieves more efficient scheduling.

[0090] The aviation assembly scheduling optimization model constructed by the present invention inputs the data information of each batch of production orders as status information during the training phase, and only needs to input the status input information of the production order during the test deployment phase. Centralized training and decentralized execution improve the efficiency and effect of data training.

[0091] The method for constructing and applying an aviation assembly scheduling optimization model based on deep reinforcement learning of the present invention solves the technical problems of complex modeling of existing aviation assembly scenarios or difficulty in solving optimal scheduling strategies.

[0092] Example 3

[0093] An embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method for constructing the aviation assembly scheduling optimization model of the above-mentioned embodiment 1, or / and implements the steps of the aviation assembly scheduling optimization method in the above-mentioned embodiment 2.

[0094] The electronic device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory may be used to store computer programs and / or modules, and the processor may perform various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.

[0095] The relevant technical solutions are the same as above and will not be repeated here.

[0096] Example 4

[0097] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for constructing the aviation assembly scheduling optimization model in the above-mentioned embodiment 1 is implemented, or / and the steps of the aviation assembly scheduling optimization method in the above-mentioned embodiment 2 are implemented.

[0098] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0099] The relevant technical solutions are the same as above and will not be repeated here.

[0100] Example 5

[0101] An embodiment of the present application provides a computer program product, including a computer program. When the computer program is run on a computer, it enables the computer to execute the method for constructing the aviation assembly scheduling optimization model in the above-mentioned embodiment 1, or / and implement the steps of the aviation assembly scheduling optimization method in the above-mentioned embodiment 2.

[0102] The relevant technical solutions are the same as above and will not be repeated here.

[0103] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing an aviation assembly scheduling optimization model, characterized in that: include: Based on the number of aircraft in a batch of production orders, the time span of each assembly link, the number of jigs, the assembly sequence constraints of each aircraft, and the assembly sequence constraints of each aircraft on the jig, an assembly disjunction diagram of the aviation assembly environment of the batch of production orders is constructed; The aviation assembly environment information of the batch of production orders is extracted from the assembly disjunctive graph to form the state feature vector s t , and input it into the agent, the agent outputs action a t The scheduling strategy for each aircraft assembly link in this batch of production orders is obtained, and the reward value r corresponding to each scheduling strategy is obtained. t ; Among them, the reward function is to minimize the number of idle frames; the action a t Feedback to the aviation assembly environment to update the assembly disjunction graph and obtain the state feature vector s at the next moment t+1 , the tuple {s t ,a t ,r t ,s t+1 }Store in the experience pool as a training data set; The training data set is used to train the intelligent agent, the network parameters of the intelligent agent are reversely adjusted, and the trained intelligent agent is used as the aviation assembly scheduling optimization model.

2. The method for constructing an aviation assembly scheduling optimization model according to claim 1, characterized in that: The assembly disjunctive graph is: G = (V, C∪D), where the symbol ∪ represents a union; Among them, the node V in the assembly disjunctive graph G is used to represent an assembly node for each aircraft. The attribute information of the assembly node includes: the serial number of the aircraft in the batch of production orders, the completion rate of the assembly link at the assembly node, the time span of the assembly link, and the link waiting for assembly; C is a directed connection arc formed by two adjacent assembly links of each aircraft, and D is an undirected connection arc formed by two adjacent assembly links of a jig.

3. The method for constructing an aviation assembly scheduling optimization model according to claim 1 or 2, characterized in that: The reward value is the reward value after reward plasticity, and the reward function after reward plasticity is for: Where r(s,a) represents the original reward function, and n is the number of jigs currently being assembled, F total is the total number of frames, s is the state feature vector of the input agent at different times, a is the action of the agent at different times; w φ (s) is the weight function, which is updated as follows: Where, I(w φ ) represents the updated weight function, the symbol := represents "defined as"; E[·] represents the expected operation, S represents the set of state feature vectors at all times; α is a constant related to the aviation assembly environment, τ T It is the trajectory obtained by the agent in this round before time t+1, and They are respectively in the trajectory τ T The total number of times the agent visits all states S and states s, and They are respectively in the trajectory τ T The total normalized reward for the agent to visit all states S and state s is calculated as: Where T is the end time of an episode of reinforcement learning; and Respectively represent the upper and lower limits of the reward value after reward shaping; s t is the state feature vector of the input agent at the current time t, a t is the action of the agent at the current moment t.

4. The method for constructing an aviation assembly scheduling optimization model according to claim 3, characterized in that: The intelligent agent includes: a strategy neural network and a value neural network; The strategy neural network is used to obtain the state feature vector s at the current moment based on the assembly disjunctive graph. t , select the corresponding action a from the action space t As output, and the action a t Feedback to the aviation assembly environment to update the current assembly disjunction graph and obtain the state feature vector s at the next moment t+1 ; wherein the action a t It is used to represent the scheduling strategy of the aircraft assembly link at the current moment. The strategy neural network selects the corresponding action a from the action space based on the current value function Q(s,a;θ) t As output; The value neural network is used to determine the state feature vector s at the next moment based on the state feature vector s at the next moment. t+1 , output the estimate y of the value function Q(s,a;θ) t , to assist in updating the parameters θ of the policy neural network.

5. The method for constructing an aviation assembly scheduling optimization model according to claim 4, characterized in that: The action space distribution of the strategy neural network is: A={LOR,MOR,SPT,LPT,LTPT,STPT,FIFO,LIFO} Among them, LOR represents the maximum order ratio scheduling rule, MOR represents the minimum order ratio scheduling rule, SPT represents the shortest processing time scheduling rule, LPT represents the longest processing time scheduling rule, LTPT represents the longest total processing time scheduling rule, STPT represents the shortest total processing time scheduling rule, FIFO represents the first-in-first-out scheduling rule, and LIFO represents the first-in-last-out scheduling rule.

6. The method for constructing an aviation assembly scheduling optimization model according to claim 4 or 5, characterized in that: During the agent training process, the loss function of the policy neural network is the output y of the value neural network t The root mean square error between the loss function Q(s, a; θ) and the value function Q(s, a; θ); based on the loss function, the parameters θ of the policy neural network are updated using gradient back propagation; The value of the neural network parameter θ - The update method is: θ - ←θ, which means that the policy neural network copies its parameters θ to the value neural network at every certain time step.

7. An aviation assembly scheduling optimization method, characterized in that: include: Based on the number of aircraft in a batch of production orders, the time span of each assembly link, the number of jigs, the assembly sequence constraints of each aircraft, and the assembly sequence constraints of each aircraft on the jig, an assembly disjunction diagram of the aviation assembly environment of the batch of production orders is constructed; The aviation assembly environment information of the batch of production orders is extracted from the assembly disjunctive graph to form a state feature vector, and the state feature vector is input into the aviation assembly scheduling optimization model constructed by the aviation assembly scheduling optimization model construction method described in any one of claims 1 to 6. The aviation assembly scheduling optimization model outputs the scheduling strategy for each aircraft assembly link in the batch of production orders.

8. An electronic device, characterized in that: comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read the executable instructions stored in the computer-readable storage medium to execute the method for constructing an aviation assembly scheduling optimization model according to any one of claims 1 to 6, or / and to execute the aviation assembly scheduling optimization method according to claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for constructing an aviation assembly scheduling optimization model according to any one of claims 1 to 6 is implemented, or / and the aviation assembly scheduling optimization method according to claim 7 is implemented.

10. A computer program product, characterized in that The method comprises a computer program, which, when executed on a computer, enables the computer to execute the method for constructing an aviation assembly scheduling optimization model according to any one of claims 1 to 6, or / and execute the aviation assembly scheduling optimization method according to claim 7.