Method and device for scheduling a suspension system

By using a weighted optimization network to dynamically adjust the scoring and penalty weights in the hanging system, the scheduling path of the clothes hangers is optimized, which solves the problem of low overall processing efficiency of the hanging system and improves production efficiency and scene adaptability.

CN122133994APending Publication Date: 2026-06-02ZHEJIANG YIKEDA INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG YIKEDA INTELLIGENT TECH CO LTD
Filing Date
2026-02-14
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The overall processing efficiency of hanging systems in the light manufacturing industry is low, and they are difficult to adapt to production scenarios with multiple styles, multiple orders, and dynamic adjustments.

Method used

By determining the candidate locations of the hangers to be scheduled on the hanging system, along with their scoring and penalty factors, a weight optimization strategy is generated using a trained weight optimization network. This strategy dynamically adjusts the scoring and penalty weights to optimize the hanger scheduling path.

Benefits of technology

It improves the overall production efficiency of the hanging system, achieves workstation load balance, improves garment flow efficiency and maximizes equipment utilization, shortens order delivery cycle, and adapts to the needs of multi-style, small-batch, fast-response production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133994A_ABST
    Figure CN122133994A_ABST
Patent Text Reader

Abstract

This application relates to a scheduling method and apparatus for a hanging system. The scheduling method includes: determining multiple candidate positions for hangers to be scheduled, and a scoring factor, penalty factor, current scoring weight, and current penalty weight corresponding to each candidate position; generating a state space based on the current scoring weight, current penalty weight, and real-time global operating state information of the hanging system; inputting the state space into a trained weight optimization network to obtain a weight optimization strategy, and adjusting the current scoring weight and current penalty weight according to the weight optimization strategy to obtain a target scoring weight and target penalty weight; determining the score of each candidate position based on the target scoring weight, target penalty weight, scoring factor, and penalty factor; determining the target position for the hangers to be scheduled based on the scores of each candidate position, and scheduling the hangers to be scheduled to the target position. This application solves the problem of low overall processing efficiency in hanging systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of suspension system technology, and in particular to a method and apparatus for scheduling suspension systems. Background Technology

[0002] In light manufacturing industries such as apparel and home textiles, hanging systems serve as core production equipment, handling the transport of hangers, process flow, and station allocation, and are crucial for achieving large-scale, automated production. This system binds hangers to production schedules through hanging stations. Hangers then sequentially pass through multiple sewing stations along the hanging line to complete different processing steps. The core process includes card reading upon entry, start-up within the station, and exit upon completion, during which the system logic allocates the next target station. However, related technologies using fixed-route allocation methods suffer from low overall processing efficiency, making it difficult to adapt to the dynamic production scenarios of factories with multiple styles, orders, and adjustments.

[0003] Currently, no effective solution has been proposed to address the issue of low overall processing efficiency of suspension systems in related technologies. Summary of the Invention

[0004] This application provides a scheduling method and apparatus for a hanging system, which at least solves the problem of low overall processing efficiency of hanging systems in related technologies.

[0005] In a first aspect, embodiments of this application provide a scheduling method for a suspension system, the method comprising:

[0006] Determine multiple candidate locations for the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate location;

[0007] Based on the current scoring weight, the current penalty weight, and the real-time global operating status information of the suspension system, a corresponding state space is generated;

[0008] The state space is input into the trained weight optimization network to obtain a weight optimization strategy. Based on the weight optimization strategy, the current rating weight and the current penalty weight are adjusted to obtain the target rating weight and the target penalty weight, respectively.

[0009] Based on the target score weight, the target penalty weight, the score factor, and the penalty factor, the score corresponding to each candidate station is determined;

[0010] Based on the scores corresponding to each candidate station, the target station for the hanger to be scheduled is determined, and the hanger to be scheduled is then scheduled to the target station.

[0011] In some embodiments, the training process of the weight optimization network includes multiple training rounds; each training round includes multiple time steps.

[0012] The training process for optimizing the network's weights for each training epoch includes:

[0013] For each time step, generate the current state transition tuple and store the state transition tuple in a preset experience replay buffer;

[0014] Randomly sample multiple state transition tuples from the experience replay buffer;

[0015] For each state transition tuple, the current state space and the current weight adjustment action in the state transition tuple are input into the initial weight optimization network to obtain the predicted action value and determine the target action value corresponding to the state transition tuple.

[0016] Based on the predicted action value and the target action value, the corresponding loss function result is determined, and the gradient of the loss function result is backpropagated to the initial weight optimization network to obtain the updated weight optimization network.

[0017] In some embodiments, the weight optimization network includes an evaluation network and a target network;

[0018] The step of backpropagating the gradient of the loss function result to the initial weight optimization network to obtain the updated weight optimization network includes:

[0019] The gradient of the loss function result is backpropagated to the initial evaluation network of the initial weight optimization network to obtain the updated evaluation network. After a preset number of time steps, the current parameters of the evaluation network are copied to the target network to obtain the updated target network.

[0020] Based on the updated evaluation network and the updated target network, the updated weight optimization network is determined.

[0021] In some embodiments, determining the target action value corresponding to the state transition tuple includes:

[0022] When the current state transition tuple is the last state transition tuple among multiple randomly sampled state transition tuples, the reward result in the current state transition tuple is used as the target action value;

[0023] When the current state transition tuple is not the last state transition tuple among multiple randomly sampled state transition tuples, the new state space in the current state transition tuple is input into the target network to predict the maximum action value, and the target action value is generated based on the reward result in the current state transition tuple and the maximum action value.

[0024] In some embodiments, generating the current state transition tuple includes:

[0025] Select a weight adjustment action from the preset action space, and adjust the current scoring weight and penalty weight based on the weight adjustment action to obtain the first scoring weight and the first penalty weight;

[0026] The first scoring weight and the first penalty weight are applied to the suspension system, and after the suspension system has been running for a preset period of time, the production data of the suspension system are statistically obtained.

[0027] Determine the reward result corresponding to the production data, and the new state space corresponding to the production data;

[0028] Based on the current state space, the weight adjustment action, the reward result, and the new state space, generate the current state transition tuple.

[0029] In some embodiments, the preset action space includes multiple discrete actions for adjusting weights; the sum of the first preset probability and the second preset probability is 1;

[0030] The step of selecting weight adjustment actions from a preset action space includes:

[0031] Based on the first preset probability, a discrete action is randomly selected from the preset action space as the weight adjustment action; or,

[0032] Using the second preset probability, the discrete action that maximizes the action value of the evaluation network output in the weight optimization network is selected from the preset action space as the weight adjustment action.

[0033] In some embodiments, determining the reward result corresponding to the production data includes:

[0034] Identify multiple sub-production data sets within the production data; each sub-production data set includes at least one category of sub-production data.

[0035] For each sub-production data set, determine the target reward corresponding to the sub-production data set;

[0036] Based on the target reward corresponding to each of the sub-production data sets, the reward result corresponding to the production data is determined.

[0037] In some embodiments, adjusting the current scoring weight and penalty weight based on the weight adjustment action to obtain a first scoring weight and a first penalty weight includes:

[0038] Based on the weight adjustment action, the current scoring weight and penalty weight are adjusted to obtain the second scoring weight and the second penalty weight;

[0039] The second scoring weight and the second penalty weight are normalized to obtain the first scoring weight and the first penalty weight.

[0040] In some embodiments, before adjusting the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight, the method further includes:

[0041] Obtain the current production time and find the target time period to which the current production time belongs from multiple preset time periods;

[0042] Based on the target time period and the correspondence between each preset time period and the weighting strategy, the target weighting strategy corresponding to the target time period is determined.

[0043] The target weight strategy is determined as the weight optimization strategy.

[0044] Secondly, embodiments of this application provide a scheduling device for a suspension system, the device comprising:

[0045] The determination module is used to determine multiple candidate positions of the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate position.

[0046] The state space generation module is used to generate a corresponding state space based on the current scoring weight, the current penalty weight, and the real-time global operating state information of the suspension system.

[0047] The weight optimization module is used to input the state space into the trained weight optimization network to obtain the weight optimization strategy, and adjust the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight respectively.

[0048] The scoring calculation module determines the score corresponding to each of the candidate stations based on the target score weight, the target penalty weight, the score factor, and the penalty factor.

[0049] The hanger scheduling module is used to determine the target station of the hanger to be scheduled based on the score corresponding to each candidate station, and to schedule the hanger to be scheduled to the target station.

[0050] Compared to related technologies, the scheduling method and apparatus for the hanging system provided in this application solves the problem of low overall processing efficiency of the hanging system by determining multiple candidate positions of the hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate position; generating a corresponding state space based on the current scoring weight, current penalty weight, and real-time global operating status information of the hanging system; inputting the state space into a trained weight optimization network to obtain a weight optimization strategy, and adjusting the current scoring weight and current penalty weight according to the weight optimization strategy to obtain the target scoring weight and target penalty weight respectively; determining the score corresponding to each candidate position based on the target scoring weight, target penalty weight, scoring factor, and penalty factor; determining the target position of the hanger to be scheduled based on the score corresponding to each candidate position, and scheduling the hanger to be scheduled to the target position.

[0051] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0052] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0053] Figure 1 This is a hardware structure block diagram of a terminal for a scheduling method of a suspension system according to an embodiment of this application;

[0054] Figure 2 This is a flowchart of a scheduling method for a suspension system according to an embodiment of this application;

[0055] Figure 3 This is a schematic diagram of the weighted optimization network hanging integration process according to an embodiment of this application;

[0056] Figure 4 This is a schematic diagram of the overall clothes hanger target location selection scoring and overall scheduling process according to an embodiment of this application;

[0057] Figure 5This is a structural block diagram of the scheduling device of the hanging system according to an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0059] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0060] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0061] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure block diagram of a terminal for a scheduling method of a suspension system according to an embodiment of this application. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0062] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the scheduling method of the hanging system in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0063] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0064] The hanging system includes style (style number, color, size), order (binding a style, there will be order details corresponding to different colors and sizes), order process (determining the sewing process), order process route map (determining the actual sewing station corresponding to the process), and production schedule (after the production schedule is generated, workers can hang the garment pieces on the hanger at the hanging station, and after the hanging is completed, the hanger is bound to the corresponding production schedule).

[0065] In the hanging system production process, hangers are read by a card upon entry: after reading the card, the hanger enters the sewing station; within the station, another card is read; after card reading is completed, the hanger begins processing, indicating that the worker has started processing. Upon completion, the worker presses the "complete" button, the hanger is finished and exits the station. During this process, the system calculates bridging logic and assigns the next target station and the next station (the next station is used for bridging calculations and involves cross-line transfers). A hanging line has N stations, and each station can perform multiple processes. A single hanger (with accessories and cut pieces) will undergo multiple processes, processed at multiple stations. The hanging line operates in a counter-clockwise loop.

[0066] After the hangers are hung at the hanging station (usually station number 1), the hangers and production schedules are linked in the system.

[0067] After the hangers in the hanging system are completed, they are assigned to target stations according to the user-defined route map. If there are multiple sewing stations within a hanging line that can process the next process, the station allocation method set by the user will not take into account factors such as actual operating conditions (uneven load of sewing stations, hanger routing efficiency, worker skills and efficiency factors, hanger cross-line costs, and urgency priority of hanger processes), which will reduce the overall processing efficiency of the system.

[0068] To address the aforementioned problems, this embodiment provides a scheduling method for a suspension system. Figure 2 This is a flowchart of a scheduling method for a suspension system according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0069] Step S201: Determine multiple candidate positions of the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate position.

[0070] Specifically, after the hanging system completes the current process and the hanger leaves the station, the system first filters out all stations within the entire factory that have the processing capability for that process based on the requirements of the hanger's next process, forming a candidate station list for hangers to be scheduled. During the filtering process, invalid stations that are not workstations, have equipment malfunctions, have not completed personnel login, or have stations that do not match the next process must be excluded to ensure that all candidate stations meet the basic processing conditions.

[0071] Subsequently, for each selected candidate station, the system extracts corresponding scoring factors and penalty factors. The scoring factors, used to increase the station's score, mainly cover hanger path efficiency (the time cost of historical hangers traveling from the current station to the target station based on statistics), station processing efficiency (calculating efficiency based on historical data of all stations performing the current process), and station style matching (style, color, and size). The penalty factors, used to decrease the station's score, mainly include station load (number of hangers currently queuing in the station / total hanger capacity in the station), order urgency (delivery date of the production schedule in the system), cross-line penalty (hangers need to be transferred across hanging lines), and station health (station equipment utilization rate).

[0072] Simultaneously, the system will retrieve the currently effective weight configuration scheme to obtain the current scoring weights corresponding to the aforementioned scoring factors, as well as the current penalty weights corresponding to the penalty factors. These weights will be dynamically adjusted according to preset strategies or algorithms, and will always maintain the sum of the scoring weights and penalty weights at 1, providing a basis for subsequent station scoring calculations.

[0073] Step S202: Based on the current scoring weight, the current penalty weight, and the real-time global operating status information of the suspension system, generate the corresponding state space;

[0074] The global operational status information includes hanger status, station status, system status, and path status. Hanger status data specifically includes the current station location of the hanger to be scheduled, the list of completed processes, the remaining process sequence, and the urgency level of the order (remaining days). Station status mainly includes the real-time load of each candidate station, the historical average processing efficiency for the current hanger style / process, and the station equipment health (availability). System status involves the overall load of each hanging line, the overall load standard deviation, and the total work-in-process (WIP) volume (number of hangers started at each station) across the entire system. Path status data is the estimated travel time from the current hanger location to each candidate station (derived from historical data statistics).

[0075] The system integrates and quantifies the aforementioned global operating status information and the previous weight configuration data (the current scoring weight corresponding to the scoring factor and the current penalty weight corresponding to the penalty factor), and finally generates a state space that can comprehensively characterize the operating status of the suspension system at a certain moment, providing a complete and accurate input data foundation for the subsequent decision analysis of the weight optimization network.

[0076] Step S203: Input the state space into the trained weight optimization network to obtain the weight optimization strategy, and adjust the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight respectively.

[0077] Specifically, after constructing the state space, the system integrates the current weight configuration and the state space data of the overall operating state of the suspension system, and inputs it into a weight optimization network that has undergone offline training and online fine-tuning. This weight optimization network can be built on a Deep Q-Network (DQN) architecture, with a built-in dual-network structure of evaluation network and target network. Through simulation training based on historical production data from the factory and continuous iteration with real-time production data, it has the ability to optimize weights to adapt to dynamic production scenarios.

[0078] After receiving the state space input, the weight optimization network combines historical experience data stored within the network to calculate and output the optimal weight optimization strategy for the current production state. This strategy clarifies the specific adjustment direction and magnitude of the weights of each scoring factor and penalty factor, and the adjustment actions cover two types of discrete operations: small increases / decreases (±0.01) and large increases / decreases (±0.1).

[0079] When executing the weight optimization strategy, the system first adjusts the current scoring weight and current penalty weight according to the strategy instructions to obtain the adjusted transition weight. Subsequently, to ensure that the sum of the scoring weight and the penalty weight is always 1, L1 normalization is performed on the transition weight. That is, the adjusted scoring weight and penalty weight are divided by the sum of their respective weight sets, and finally the target scoring weight and target penalty weight that meet the normalization requirements are obtained, providing an accurate weight basis for the subsequent comprehensive score calculation of candidate stations.

[0080] Step S204: Determine the score corresponding to each candidate station based on the target score weight, target penalty weight, score factor, and penalty factor;

[0081] Specifically, after completing the dynamic optimization configuration of the target scoring weight and target penalty weight, the system will extract the values ​​of various scoring factors and penalty factors corresponding to each valid candidate station. Among them, the scoring factors include three types of positive impact indicators: the historical path efficiency of the hanger from the current station to the candidate station, the historical average efficiency of the candidate station in processing the current process, and the color and size matching degree between the candidate station and the style of the hanger; the penalty factors include four types of negative impact indicators: the real-time load rate of the candidate station, the order urgency of the production schedule corresponding to the hanger, the cost of transferring the hanger across lines to the station, and the equipment utilization rate of the candidate station.

[0082] Subsequently, the system strictly follows the preset scoring calculation formula, multiplying each scoring factor by its corresponding target scoring weight and summing the results to obtain the positive score value for the station; then, multiplying each penalty factor by its corresponding target penalty weight and summing the results to obtain the negative score value for the station. Finally, the positive and negative score values ​​are accumulated to obtain the comprehensive score for the candidate station, with the comprehensive score controlled within the range of [0,1]. The specific calculation formula is as follows:

[0083] ;

[0084] in, The overall score for each candidate station; F i For rating factors; w i For the corresponding rating factor F i The scoring weights are used to adjust the degree of influence of different scoring factors on the overall score; P j As a penalty factor; w p,j The penalty weight is the penalty factor; and The scoring weight and penalty weight are summed to 1. The score calculated in this way can comprehensively and accurately reflect the suitability of the candidate station for undertaking the hanger process under the current production conditions, providing a scientific and quantitative basis for subsequent target station decisions.

[0085] Step S205: Based on the scores corresponding to each candidate station, determine the target station for the hanger to be scheduled, and schedule the hanger to be scheduled to the target station.

[0086] Specifically, after calculating the comprehensive score of all candidate stations, the system will select the target station based on the score, generate a scheduling instruction containing the unique identifier of the hanger and the target station, and schedule the hanger to the target station.

[0087] The above steps first determine the scoring / penalty factors and their current weights, then construct a comprehensive system state space based on the global operating status. The network dynamically outputs weight strategies adapted to the current production scenario based on the trained weights. After precisely adjusting the scoring and penalty weights, the suitability of each candidate station is quantified. Finally, the optimal target station is selected based on the score to complete the scheduling. This process comprehensively considers core production elements such as path efficiency, station load, and order urgency. It solves problems such as unbalanced load, high cross-line costs, and slow response to urgent orders in traditional fixed allocation. Through dynamic weight optimization and intelligent decision-making, it achieves balanced workstation load, improved garment flow efficiency, and maximized equipment utilization. Simultaneously, it ensures shorter order delivery cycles, adapts to the needs of multi-style, small-batch, fast-response production, and significantly improves the overall production efficiency and scenario adaptability of the hanging system.

[0088] In some embodiments, the training process of the weight optimization network includes multiple training rounds; each training round includes multiple time steps.

[0089] The training process for optimizing the network weights for each training epoch includes:

[0090] For each time step, generate the current state transition tuple and store the state transition tuple in a preset experience replay buffer;

[0091] Randomly sample multiple state transition tuples from the experience replay buffer;

[0092] For each state transition tuple, the current state space and the current weight adjustment action in the state transition tuple are input into the initial weight optimization network to obtain the predicted action value and determine the target action value corresponding to the state transition tuple.

[0093] Based on the predicted action value and the target action value, the corresponding loss function result is determined, and the gradient of the loss function result is backpropagated to the initial weight optimization network to obtain the updated weight optimization network.

[0094] The training process of the weight optimization network adopts a multi-round iterative mode. Each training round consists of multiple consecutive time steps. Through state interactions, experience accumulation, and network iteration within each time step, the weight adjustment strategy is gradually optimized. The specific execution flow for training the weight optimization network in each training round is as follows:

[0095] Within each time step t, the system will base its current state space S on... t The weight adjustment action performed is a t The reward result R obtained after the action is performed t And the new state space S obtained after the action is executed and updated. t+1 Generate a complete state transition tuple (S t ,a t ,R t ,S t+1 A training epoch has multiple time steps, which generates multiple state transition tuples. The generated state transition tuples are stored in real time in a pre-defined experience replay buffer. This buffer is used to store a large amount of interactive experience generated during training, providing data support for subsequent network training.

[0096] To avoid the impact of temporal correlations in the training data on network convergence, the system randomly samples a certain number of state transition tuples from the experience replay buffer to form the training data. This random sampling method disrupts the temporal order of the data, making network training more stable.

[0097] For each sampled state transition tuple, the system will store the current state space S in the tuple. j and the current weight adjustment action a t The input is fed into the initial weight optimization network. Based on the input data, the initial weight optimization network outputs the predicted action value Q(S) for the current state. j ,a j ;θ), where θ is the network parameter. Simultaneously, the system combines the reward result from the state transition tuple with the new state space to calculate the target action value y corresponding to the action. j This provides a reference baseline for subsequent network parameter updates.

[0098] Next, the system will construct a corresponding loss function and calculate the loss function result based on the difference between the predicted action value and the target action value. The loss function formula is as follows:

[0099] ;

[0100] in, θ represents the loss function value, characterizing the prediction error of the network under the current parameters; the smaller the value, the more accurate the network prediction. θ represents the network parameters, which are the core objects that need to be iteratively optimized during training and directly determine the output accuracy of the network. M represents the number of state transition tuples randomly sampled from the experience replay buffer, reducing the impact of temporal correlation on training through batch sampling. j The target action value corresponding to the j-th sampled tuple is a true reference value based on environmental feedback and future state value; Q(S j ,a j ;θ) is used to evaluate the value of the network's predicted actions, i.e., based on the current parameters θ, the state S of the j-th tuple is input. j and action a j Then, the network outputs the action value prediction result. This formula quantifies the prediction bias of a single sample by calculating the squared difference between the target action value and the predicted action value of each sampled tuple; by averaging the squared errors of M sampled tuples, the average error of the batch of samples is obtained, which reduces the interference of outliers of individual samples on training and improves the stability of parameter updates.

[0101] Subsequently, the gradient of the loss function result is backpropagated to the initial weight optimization network. The network parameters are iteratively updated using the gradient descent algorithm, so that the network's predicted action value gradually approaches the real target action value, thereby improving the network's evaluation accuracy of the weight adjustment action value, and finally obtaining the optimized weight optimization network.

[0102] In the above steps, through iterative training with multiple training rounds and time steps, generation and storage of state transition tuples and random sampling, calculation and comparison of predicted action value and target action value, and network parameter update based on backpropagation of loss function gradient, the weight optimization network is accurately adapted to the dynamic production state of the hanging system, effectively improving the stability and convergence speed of network training. The final output weight optimization strategy can significantly optimize the rationality of hanger station scheduling and improve the overall production efficiency and load balance level of the hanging system.

[0103] In some embodiments, the weight optimization network includes an evaluation network and a target network;

[0104] The gradient of the loss function result is backpropagated to the initial weight optimization network to obtain the updated weight optimization network, including:

[0105] The gradient of the loss function result is backpropagated to the initial evaluation network of the initial weight optimization network to obtain the updated evaluation network. After a preset number of time steps, the current parameters of the evaluation network are copied to the target network to obtain the updated target network.

[0106] Based on the updated evaluation network and the updated target network, the updated weight optimization network is determined.

[0107] The gradient of the loss function result is backpropagated to the initial evaluation network of the initial weight optimization network. The core parameter of this initial evaluation network is θ. The gradient backpropagation process iteratively adjusts parameter θ using the gradient descent algorithm based on the difference between the predicted action value and the target action value, thereby minimizing the loss function and ultimately obtaining the updated evaluation network. The updated evaluation network can more accurately output the action value corresponding to the current state and the weight adjustment action, providing a more reliable decision-making basis for subsequent weight optimization strategy generation.

[0108] After updating the evaluation network parameters, the system initiates a parameter synchronization mechanism for the target network. Specifically, the system copies the updated parameters θ of the evaluation network to the target network at preset time step intervals (e.g., every C time steps), replacing the original parameters θ' of the target network, thus obtaining the updated target network. The target network maintains fixed parameters during the interval between two parameter updates. This serves to provide a stable benchmark for calculating the target action value, avoiding training oscillations caused by real-time changes in the evaluation network parameters, and ensuring the stability and convergence of the weight optimization network training process.

[0109] Finally, based on the updated evaluation network and the updated target network, the updated weight optimization network is determined. In subsequent training, the evaluation network is responsible for outputting action values ​​in real time and participating in parameter iteration, while the target network is responsible for providing a stable reference for target action values. The two work together to improve the adaptability of the weight optimization network to the dynamic production scenarios of the suspension system.

[0110] The above steps achieve iterative parameter updates by backpropagating the gradient of the loss function only to the evaluation network. At the same time, the parameters of the evaluation network are synchronized to the target network according to a preset time step and their parameters are kept stable. The training method of dual-network division of labor and cooperation effectively avoids the problem of target value oscillation caused by real-time parameter changes during single-network training. It improves the stability and convergence speed of the weight optimization network training, and ensures that the trained weight optimization network can output a more accurate weight adjustment strategy, thereby optimizing the rationality of the hanging system's clothes hanger position scheduling and overall production efficiency.

[0111] In some embodiments, the training process of the weight optimization network also includes an initialization phase. This initialization phase, as a fundamental preparatory step for network training, requires the completion of three core operations: First, the dual-network structure of the agent (weight optimization network) is initialized, namely, a Q-network (evaluation network) for real-time evaluation of action value and a target Q-network for providing stable target values. Initial parameters θ (evaluation network parameters) and θ' (target network parameters) are assigned to the two networks respectively, laying the foundation for subsequent network iterations and parameter updates. Second, the experience replay buffer D is initialized, with a preset fixed capacity N. This buffer stores state transition tuples generated during training, providing a data storage medium for subsequent random sampling and reducing data temporal correlation. Finally, the state space S required during training is explicitly defined. t The state space needs to comprehensively cover the global operation information of the hanging system, specifically including the hanger state, position state, system state, path state, and weight configuration information consisting of the current scoring weight and penalty weight, to ensure that the state space can completely represent the system operation status at time t, and provide comprehensive input basis for agent decision-making and network training.

[0112] In some embodiments, determining the target action value corresponding to the state transition tuple includes:

[0113] When the current state transition tuple is the last state transition tuple among multiple randomly sampled state transition tuples, the reward result in the current state transition tuple is used as the target action value.

[0114] When the current state transition tuple is not the last state transition tuple among multiple randomly sampled state transition tuples, the new state space in the current state transition tuple is input into the target network to predict the maximum action value. Based on the reward result in the current state transition tuple and the maximum action value, the target action value is generated.

[0115] Specifically, when the current state transition tuple is the last of multiple randomly sampled state transition tuples, it means that the training scenario corresponding to this tuple has entered the termination phase, and there is no value continuation for subsequent states. At this time, the reward result R in the current state transition tuple is... j Directly used as the value of the target action y j , that is, y j =R j The core basis of this design is that the value of the action in the termination phase is determined solely by the immediate reward from the environment after the current action is performed, without the need to add the potential value of subsequent states. This ensures that the value of the target action is perfectly matched with the actual training scenario in that phase, avoiding calculation errors caused by introducing non-existent subsequent state values.

[0116] When the current state transition tuple is not the last of multiple randomly sampled state transition tuples, the target action value needs to be generated by combining the immediate reward and the potential value of subsequent states. Specifically, firstly, the new state space recorded in the current state transition tuple is extracted and input into the target network of the weight optimization network. The target network predicts the action value corresponding to all possible weight adjustment actions in this new state space based on a fixed parameter θ' (not updated in real time with the current training batch), and selects the maximum action value to represent the optimal potential value in the new state. Subsequently, the reward result in the current state transition tuple is weighted and fused with the aforementioned maximum action value (usually with a discount factor γ to adjust the influence of subsequent values) to finally generate the target action value corresponding to the current state transition tuple. The specific formula is as follows:

[0117] ;

[0118] Among them, y j R represents the target action value corresponding to the j-th state transition tuple; j In the j-th state transition tuple, γ represents the reward result of the suspension system within a preset time after performing the weight adjustment action; γ is a discount factor used to adjust the degree of influence of the potential value of the next state on the value of the current target action, usually set according to the factory production cycle; S j+1 This is the next state space relative to the current state, i.e., the new state space in the state transition tuple; The candidate action for the next state, i.e., in S j+1Below, the weight optimization network can perform all discrete actions; To select the action with the maximum value, i.e., for S j+1 All candidate actions The corresponding Q value is taken at its maximum value, which is the value of selecting the potential optimal action in the next state; The action value output by the target network is calculated by the target network in the weight optimization network, and the next state S is output. j+1 Next action The estimated value. Where θ' is the parameter of the target network (periodically copied from the evaluation network).

[0119] The above steps employ differentiated computational logic. When the tuple is the last one (corresponding to the training termination scenario), the target action value is directly defined using the immediate reward, avoiding invalid predictions of potential value without subsequent states and ensuring the accuracy of value calculation in the termination scenario. Conversely, when the tuple is not the last one (corresponding to the regular training scenario), the target action value is generated by combining the maximum action value of the new state output by the target network with the immediate reward, fully considering the long-term potential benefits of subsequent states, and avoiding training oscillations based on the characteristics of the target network's fixed parameters. Ultimately, the target action value accurately reflects the actual value of the action in different training scenarios, providing a stable and reliable benchmark for evaluating network parameter updates, effectively ensuring the convergence of the weight optimization network training, and thus ensuring that the trained network can output the optimal weight strategy adapted to the dynamic production scenario of the hanging system, improving the rationality of hanger scheduling and the overall production efficiency of the system.

[0120] In some embodiments, generating the current state transition tuple includes:

[0121] Select a weight adjustment action from the preset action space, and adjust the current score weight and penalty weight based on the weight adjustment action to obtain the first score weight and the first penalty weight;

[0122] The first scoring weight and the first penalty weight are applied to the suspension system, and the production data of the suspension system are collected after the suspension system has been running for a preset time.

[0123] Determine the reward results corresponding to the production data, and the new state space corresponding to the production data;

[0124] Based on the current state space, weight adjustment actions, reward results, and the new state space, generate the current state transition tuple.

[0125] Specifically, firstly, a weight adjustment action a that is suitable for the current production state is selected from the preset action space. tAfter selecting the weight adjustment action, the system first adjusts the current scoring weights (clothes hanger path efficiency weight, station processing efficiency weight, and station style matching weight) and the current penalty weights (station load weight, order urgency weight, cross-line penalty weight, and station health weight) according to the action instructions to obtain the adjusted weights.

[0126] Secondly, the first scoring weight and the first penalty weight are applied to the actual scheduling process of the hanging system. After the system runs for a preset duration according to these weights (the specific duration can be customized in the system configuration file according to the factory's production rhythm), the system collects production data of the hanging system through a preset data acquisition interface. This production data covers all dimensions of information related to weight optimization. During the data acquisition process, the system can use SQL statistical statements to query and calculate production data information such as output rewards, smoothness rewards, load balancing between workstations, and scheduling delays in order delivery caused by negative rewards from the intelligent hanging system database.

[0127] Next, based on the statistically obtained production data, the corresponding reward result R is calculated respectively. t With the new state space S t+1 In the reward calculation phase, the system strictly adheres to the preset total reward function. It first decomposes the production data into three positive sub-production data sets: output, smoothness, and load balancing; and three negative sub-production data sets: order delay, station congestion, and cross-line transfer. For positive data, it calculates output rewards, smoothness rewards, and load balancing rewards; for negative data, it calculates order delay penalties, station congestion penalties, and cross-line transfer penalties. Then, it uses hyperparameters to weight and fuse these positive rewards and negative penalties to obtain the final reward result, quantifying the actual production value of the current weight adjustment action. In the new state space generation phase, the system updates the hanger state, station state, system state, and path state in the production data to real-time data after the weight adjustment action is executed. It also integrates the current application's scoring weights and penalty weights to form a new state space that comprehensively represents the current operating status of the hanging system, completing the information iteration from the current state to the next state.

[0128] Finally, based on the current state space S t 1. Weight adjustment actions that have been executed t The calculated reward result R t and the updated new state space S t+1 According to (S) t ,a t ,R t ,S t+1The current state transition tuple is generated in a structured format. This tuple fully records the entire chain of information from "system state at a certain moment - weight adjustment action - action feedback reward - system state after action". It can be directly stored in the experience replay buffer to provide standardized experience data support for experience sampling and parameter updates in subsequent network training. This ensures that the weight optimization network can continuously iterate based on real production interaction data and gradually improve its adaptability to dynamic production scenarios.

[0129] The above steps ensure the standardization and adaptability of weight adjustments by selecting weight adjustment actions from the preset action space and adjusting the scoring and penalty weights. The adjusted weights are then applied to the hanging system, and production data is collected to achieve deep interaction between the weight strategy and the actual production scenario. The reward result and new state space are determined by combining the production data, accurately quantifying the actual value of the weight adjustment and updating the system state information. Finally, the current state space, weight adjustment actions, reward results, and new state space are integrated to generate a state transition tuple, fully recording the entire production interaction experience from "state-action-feedback-new state." This process ensures the completeness and timeliness of the experience data required for training the weight optimization network, and accurately captures the production state of the hanging system through dynamic interaction and data feedback. This provides a reliable basis for subsequent network parameter updates and weight strategy optimization, effectively improving the adaptability of the weight optimization network to dynamic production scenarios, thereby ensuring the rationality of the hanging system's hanger scheduling and the stable improvement of overall production efficiency.

[0130] In some embodiments, the preset action space includes multiple discrete actions for adjusting weights; the sum of the first preset probability and the second preset probability is 1;

[0131] Select weight adjustment actions from the preset action space, including:

[0132] With a first preset probability, a discrete action is randomly selected from a preset action space as the weight adjustment action; or,

[0133] Using a second preset probability, select discrete actions from the preset action space that maximize the action value of the evaluation network output in the weight optimization network as weight adjustment actions.

[0134] The preset action space is a set of discrete actions specifically designed for precise adjustment of scoring and penalty weights. This action space is designed to fully cover various scenarios for weight adjustment, containing 28 discrete actions, specifically divided into four categories: 7 actions for small increases (Δw=+0.01), 7 actions for small decreases (Δw=-0.01), 7 actions for large increases (Δw=+0.1), and 7 actions for large decreases (Δw=-0.1). Each action corresponds to a clearly defined weight adjustment target and adjustment range, ensuring that the weight configuration is accurately changed after the action is executed.

[0135] The action selection process employs a dual-probability collaborative strategy, where the sum of the first and second preset probabilities is 1. The specific ratio of these probabilities can be flexibly configured within the system based on the needs of the factory production scenario and the network training phase. On one hand, the system randomly selects a discrete action from the preset action space using the first preset probability ε as the weight adjustment action. The core function of this random selection mechanism is to ensure the network's exploratory ability and prevent it from getting trapped in local optima. By randomly trying different weight adjustment actions, the system can explore more potential optimization directions, especially in the early stages of training or when significant changes occur in the production scenario. It can quickly traverse various weight configuration combinations, accumulating rich interaction experience for the network and laying the foundation for subsequent strategy optimization. On the other hand, the system selects a discrete action from the preset action space using the second preset probability (1-ε) that maximizes the action value output by the evaluation network in the weight optimization network as the weight adjustment action. Specifically, this action must satisfy the condition that maximizes the action value output by the evaluation network in the weight optimization network: the evaluation network will evaluate the system based on the current time step's state space S. t Using its own network parameters θ, the action value is calculated for each of the 28 discrete actions in the action space, and the action value Q(S) corresponding to each action is output. t The system generates a 28-dimensional action value vector (a, θ). It then compares and filters these 28 action values, selecting the discrete action with the highest value as the weight adjustment action for the current time step. t =argmax a Q(S t The final weight adjustment action a is determined by (a, θ). tThe core purpose of this selection mechanism is to fully utilize the network's learned experience and prioritize weight adjustments that have been validated by historical data and are more likely to bring optimal production benefits. This ensures the effectiveness and stability of the weight optimization strategy, and significantly improves the accuracy of weight adjustments and the optimization effect on production efficiency, especially in the later stages of network training. It should be noted that the first preset probability (i.e., the exploration probability ε) gradually decreases in a linear decay manner, correspondingly increasing the second preset probability (i.e., the utilization probability 1-ε). This gradually increases the strategy's utilization tendency and reduces random exploration behavior as the training progresses.

[0136] Through the aforementioned dual-probability action selection mechanism, a dynamic balance between the network's exploration capability and utilization efficiency is achieved. This ensures that the network can continuously explore new optimization directions while fully utilizing existing experience to output efficient weight adjustment actions. Furthermore, through ε linear decay, the exploration tendency gradually decreases and the utilization tendency gradually increases during the training process. This guarantees that the network can fully explore diverse weight configurations in the early stages of training to accumulate comprehensive experience and avoid local optima, while also ensuring that the network can fully utilize the learned optimal experience in the later stages of training to improve the accuracy and stability of weight adjustment. This effectively balances the network's exploration capability and utilization efficiency, improves the convergence and adaptability of the weight optimization network training, and ultimately enables the trained network to output the optimal weight strategy that adapts to the dynamic production scenario of the hanging system, significantly optimizing the rationality of the hanger station scheduling and the overall production efficiency of the system.

[0137] In some embodiments, determining the reward result corresponding to the production data includes:

[0138] Identify multiple sub-sets of production data within the production data; each sub-set of production data includes sub-production data of at least one category;

[0139] For each sub-production data set, determine the target reward corresponding to that sub-production data set;

[0140] Based on the target reward corresponding to each sub-production data set, the reward result corresponding to the production data is determined.

[0141] Specifically, firstly, multiple sub-production data sets are identified within the production data. Each sub-production data set corresponds to a core indicator related to production efficiency and contains at least one category of sub-production data. Based on the production characteristics of the hanging system, the sub-production data sets are specifically divided into six categories: First, the output-related sub-production data set, which includes data such as the total number of completed hangers within a time step and the standard working hours for each completed hanger, directly reflecting the effect of weight adjustments on production output; second, the smoothness-related sub-production data set, which covers data such as the actual turnover time of hangers between stations and the historical average turnover time, used to measure the improvement in garment flow efficiency; third, the load balancing-related sub-production data set, including the real-time load rate of each station and the standard deviation and average of the load rate of all stations in the system. The data includes: 1) a data set representing the rationality of load distribution among workstations; 2) a data set representing order delivery, containing data such as the total number of active production orders, the estimated completion time and delivery date of each production order, and related to order delay risks; 3) a data set representing station congestion, covering data such as the total number of workstations in the entire system, the load rate of each station, and the preset load rate alarm threshold, used to determine whether there is an excessive queuing problem; and 4) a data set representing cross-line transfer, containing data such as the total number of hangers transferred across hanging lines and the fixed cost coefficient of a single cross-line transfer configured by the system, quantifying the additional costs of cross-line scheduling.

[0142] Secondly, for each sub-production data set, the corresponding target reward is determined based on the preset calculation rules. The target reward includes two categories: positive reward and negative penalty, which correspond to the improvement of production efficiency and loss, respectively.

[0143] For the production data set of output categories, output rewards can be obtained based on the standard working hours and reference standard working hours of the clothes hanger process in the production data. For example, a positive output reward can be obtained by calculating the ratio of the sum of the standard working hours of all completed clothes hangers within the time period to the system's reference standard working hours; the higher the ratio, the greater the reward. The specific formula is as follows:

[0144] ;

[0145] Among them, R production As a production bonus, it is based on the number of processes completed for all hangers at all stations in the entire factory during that period. N c In time step The total number of completed hangers in the entire system (based on historical hanger information retrieved from the intelligent hanging database). SAM i For the first The standard allowable minutes for each completed garment hanger process (obtained by aggregating historical information of individual hangers from the intelligent hanging system database). SAM stdIt serves as a reference standard working time for internal processes within the system (users set the standard working time for processes in the management system, which is stored in the database and read directly) for normalization.

[0146] For the production data set categorized as fluency, fluency rewards can be derived based on the average inter-station turnover time and historical average turnover time of the hangers in the production data. For example, a positive reward can be calculated based on the improvement rate between the current average turnover time and the historical average turnover time; a positive improvement rate earns a reward, while no improvement results in a reward of 0. The specific formula is as follows:

[0147] ;

[0148] Among them, R flow For the smoothness reward, the time spent by the clothes hanger waiting and transporting between stations is calculated as the recent historical average turnover time (e.g., the average of a sliding window) up to time step t-1, and the reward is the relative improvement rate; T avg_curr The average inter-station turnover time (time interval) of all clothes hangers within time step t. Internally, the system collects the routing information of the clothes hangers from the intelligent hanging system. The system records the time from when the clothes hanger is completed and leaves the previous station to when it starts processing at the next station, and calculates the time difference. avg_hist The historical average turnover time up to time step t-1 (statistics of all data before t-1, acquisition of hanger routing information from the intelligent hanging system, and the average time difference between the completion and exit time of each hanger from the previous station and the start time of processing at the next station).

[0149] For a load-balanced sub-production dataset, a load balancing reward can be derived based on the load rate of each workstation in the production data. For example, the positive reward can be calculated using the formula "1 - standard deviation of load rate / average load rate," with the reward value closer to 1 for a more balanced load. The specific formula is as follows:

[0150] ;

[0151] Among them, R balance For load balancing rewards, i.e., load balancing among workstations, to avoid assigning too many hangers to a single site, resulting in a backlog of hangers waiting to be started; σ load σ represents the standard deviation of the load rate; the more balanced the load, the better. load The smaller the value, the higher the reward; μ load At the end of time step t, the average load rate (number of queued coat racks / capacity) of all stations.

[0152] For the sub-production data set of order delivery categories, order delay penalties can be obtained based on the estimated completion time and delivery date of each production schedule in the production data. For example, a negative penalty for order delays can be obtained by calculating the sum of the deviation rates between the estimated completion time and delivery date of each production schedule; the higher the deviation rate, the more severe the penalty. The specific formula is as follows:

[0153] ;

[0154] Among them, P delay This refers to scheduling actions that incur order delay penalties, i.e., negative rewards that cause order delivery delays. N0 represents the total number of active production schedules within time step t; ETC O The estimated completion time for the production schedule is based on the current progress, calculated as "current time + (remaining task quantity / current production efficiency)". The remaining task quantity is the quantity of work not yet completed in the production schedule, and the current production efficiency is obtained by "completed task quantity / production schedule start time". (DDL) O The delivery date of the production schedule (users set the delivery date of the production schedule in the management system, which is stored in the database and can be read directly) is proportional to the degree of delay.

[0155] For the sub-production data set categorized as station congestion, the congestion penalty can be derived based on the load rate of each workstation and a preset alarm threshold. For example, the number of workstations exceeding the overload alarm threshold and the number of excess queue hangers can be counted and summed to obtain the negative congestion penalty. The specific formula is as follows:

[0156] ;

[0157] Among them, P congestion To address congestion in queueing, a negative reward system is implemented for excessive queuing to prevent situations where a coat hanger is consistently assigned to the same queue position; N s Total number of workstations in the entire system; load s τ represents the load factor of workstation s at the end of time step t. high The load rate alarm threshold configured for the system; I(·) is an indicator function, which takes the value 1 when the condition is true and 0 otherwise.

[0158] For sub-production datasets involving cross-line transfers, the cross-line transfer penalty can be calculated based on the total number of hangers transferred across different hanging lines in the production data. For example, the negative penalty can be calculated by multiplying the total number of hangers transferred across lines by the unit cost of transferring across lines; the more times the transfers occur across lines, the greater the penalty. The specific formula is as follows:

[0159] ;

[0160] Among them, P transfer Penalty for cross-line transfer; N xferLet C be the total number of hangers that are transferred across hanging lines within time step t. xfer Fixed cost coefficient for a single cross-line transfer configured for the system.

[0161] Finally, based on the target rewards corresponding to each sub-production data set, the final reward result for the production data can be determined through weighted fusion. The system pre-configures the weight hyperparameters for various target rewards, which can be flexibly adjusted according to the factory's production goals (such as prioritizing output improvement and reducing order delays). Specifically, one or more positive rewards from output, smoothness, and load balancing are multiplied by their corresponding hyperparameters and summed to obtain the total positive rewards; one or more negative penalties from order delivery, station congestion, and cross-line transfer are multiplied by their corresponding hyperparameters and summed to obtain the total negative penalties; the final reward result is the total positive rewards minus the total negative penalties. The specific formula can be:

[0162] ;

[0163] in, , , , , and Configure hyperparameters for each reward.

[0164] The above steps decompose production data into multiple sub-production data sets, quantify positive rewards and negative penalties respectively, and then obtain the final reward result through weighted fusion using flexibly configurable hyperparameters. This achieves precise quantification of the multi-dimensional production value of weight adjustment actions, such as short-term output improvement and long-term load balancing. At the same time, the use of detailed indicators and scientific formulas avoids the one-sidedness of single-dimensional evaluation, providing objective, comprehensive and actual production needs-based feedback for the weight optimization network. This effectively guides the network to iteratively optimize the weight strategy to adapt to the dynamic scenario of the hanging system, thereby improving the rationality of hanger scheduling, reducing production losses, and ensuring a dual improvement in the overall production efficiency and order delivery quality of the system.

[0165] In some embodiments, based on a weight adjustment action, the current scoring weight and penalty weight are adjusted to obtain a first scoring weight and a first penalty weight, including:

[0166] Based on the weight adjustment action, the current scoring weight and penalty weight are adjusted to obtain the second scoring weight and the second penalty weight;

[0167] The second scoring weight and the second penalty weight are normalized to obtain the first scoring weight and the first penalty weight.

[0168] The weight adjustment actions include two discrete operations: small-amplitude adjustments and large-amplitude adjustments. Small-amplitude adjustments involve a weight change of ±0.01, while large-amplitude adjustments involve a weight change of ±0.1. Specifically, if the selected weight adjustment action is "increase the weight of a certain scoring factor by 0.01," then the current weight value corresponding to that scoring factor is directly increased by 0.01, resulting in the transition weight for that scoring factor. If the selected weight adjustment action is "decrease the penalty weight of a certain penalty factor by 0.1," then the current penalty weight value corresponding to that penalty factor is directly decreased by 0.1, resulting in the transition weight for that penalty factor. The current weights corresponding to the scoring factors and the current penalty weights corresponding to the penalty factors are simultaneously adjusted according to the selected weight adjustment action, forming the second scoring weight and the second penalty weight. After the initial weight adjustment, the second scoring weight and the second penalty weight need to be normalized to ensure that the sum of the adjusted scoring weight and penalty weight is 1, ultimately yielding the first scoring weight and the first penalty weight that can be directly applied to the station scoring calculation. The normalization formula is as follows:

[0169] ;

[0170] ;

[0171] in, Weights for individual scoring factors (such as hanger path efficiency weight, station processing efficiency weight, and station style matching weight). The weights corresponding to individual penalty factors (such as station load weight, order urgency weight, cross-line penalty weight, and station health weight). It is the sum of all rating weights and all penalty weights.

[0172] In the above steps, the scoring and penalty weights are first matched in a directional manner through discrete adjustment actions. Then, normalization (dividing a single weight by the sum of all scoring and penalty weights) forces the total weight to be 1. This avoids the subsequent failure of the station scoring logic and the destruction of the comparability of different candidate station scores due to the superposition of values ​​after weight adjustment. It also strictly adheres to the constraint that the total weight is 1 in the scoring formula, while preserving the relative proportional relationship between each weight. This ensures that the increase or decrease of a factor's weight can truly reflect its impact on the station score, providing a scientific and controllable weight configuration basis for the subsequent accurate selection of the optimal target station and improvement of the scheduling efficiency of the hoisting system.

[0173] In some embodiments, before adjusting the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight, the method further includes:

[0174] Obtain the current production time and find the target time period to which the current production time belongs from multiple preset time periods;

[0175] Based on the target time period and the correspondence between each preset time period and the weighting strategy, determine the target weighting strategy corresponding to the target time period;

[0176] The target weight strategy is determined to be the weight optimization strategy.

[0177] The weight optimization strategy can also be determined based on the current production time. Specifically, the system first acquires the current production time in real time, which is synchronously provided by the clock module of the overhead crane system, ensuring the accuracy and real-time nature of the time acquisition. Simultaneously, the system pre-stores multiple preset time periods, which are divided based on the factory's production rhythm, historical production data, and expert experience. These preset time periods include the start of the morning shift, peak production periods, the period near the end of the shift, and the night shift. The time range of each period can be flexibly configured by the factory user through the system management interface to adapt to different factory work schedules and production plans. After acquiring the current production time and defining the preset time periods, the system compares the current production time with the time range of each preset time period one by one to find the target time period to which the current production time belongs. For example, if the factory's morning shift start period is set to 7:00-8:00, and the current production time is 7:30, then the target time period is determined to be the start of the morning shift; if the current production time is between 10:00-16:00, then the target time period is determined to be the peak production period, and so on to complete the time period matching.

[0178] Meanwhile, the system pre-establishes and stores the correspondence between each preset time period and the weight strategy, and this correspondence is determined based on the core production objectives of different time periods. The core objective at the start of the morning shift is to quickly start the production line and balance the load of each station. In the corresponding weighting strategy, the scoring factors are ranked as follows: station processing efficiency, hanger path efficiency, and station style matching. The penalty factors are ranked as follows: station load, station health, order urgency, and cross-line penalty. During peak production hours, the core objective is to maximize output and prioritize urgent orders. In the corresponding weighting strategy, the scoring factors remain unchanged, while the penalty factors are adjusted to: station load, order urgency, cross-line penalty, and station health. Near the end of the shift, the core objective is to clear the production queue and complete the day's production tasks. In the corresponding weighting strategy, the penalty factors are adjusted to: order urgency, station load, cross-line penalty, and station health. During the night shift, the core objective is to maintain production and extend equipment life. In the corresponding weighting strategy, the scoring factors are adjusted to: station style matching, station processing efficiency, and hanger path efficiency. The penalty factors are adjusted to: station health, station load, cross-line penalty, and order urgency.

[0179] After determining the target time period, the system extracts the target weight strategy corresponding to that time period based on the aforementioned preset correspondence. Subsequently, this target weight strategy is directly determined as the current weight optimization strategy used to adjust the scoring weight and penalty weight, providing a clear direction and basis for subsequent weight adjustments. This ensures that the weight configuration aligns with the production needs of the current time period, achieving precise implementation of production goals for different time periods.

[0180] Furthermore, the generation of weight optimization strategies can also be based on user settings. Users can customize the weight ratio of scoring factors and penalty factors, adjust the allocation of hanger target stations, and schedule garment production tasks according to different strategies.

[0181] Through the above steps, the weight optimization strategy is precisely aligned with the core production objectives of different time periods, avoiding the problem that fixed weight strategies are difficult to adapt to dynamic production scenarios. This ensures that weight adjustments always revolve around the core needs of the current time period, providing guidance for the accurate adjustment of subsequent target scoring weights and penalty weights in line with actual production, thereby improving the scheduling rationality and overall production efficiency of the hanging system at different production stages.

[0182] The integration and application of the hoisting system mainly consists of two core stages: initial / offline training and online application and fine-tuning. In the initial / offline training stage, the system utilizes historical production data from the factory or a simulation system to simulate hoisting production operation scenarios, driving an agent (i.e., a weight optimization network) based on the DQN algorithm to conduct large-scale interactive learning. Through continuous interaction between the agent and the simulated production environment, it gradually learns the weight configuration rules under different production states, initially constructing and training a weight scheduling strategy network adaptable to basic production scenarios. In the online application and fine-tuning stage, the trained weight optimization model is first deployed to the actual hoisting system in the factory. The hanging system continuously collects global operational status information, including hanger status, station status, system status, path status, and current weight configuration, during real-time production. The weight optimization network outputs the optimal weight adjustment action based on the current status, dynamically updating the scoring weight and penalty weight in the hanger scheduling scoring formula. At the same time, new experiences generated in actual production (i.e., new state transition tuples) are continuously stored in the experience replay buffer. Based on the new experiences in the buffer, the weight optimization model is fine-tuned online, enabling the weight scheduling strategy to continuously adapt to the dynamic changes in the factory's production mode and ensuring that the hanging system is in a state of efficient scheduling for a long time.

[0183] Figure 3This is a schematic diagram of the integration process of the weight optimization network and the hanging system according to an embodiment of this application. The flowchart shows the entire process of the integration and application of the weight optimization network and the hanging system. First, the simulation operation is triggered periodically using historical production data of the factory or a simulation system to collect the real-time status of the system and train the basic weight scheduling strategy. Then, the basic network is deployed to the actual hanging system of the factory. The network outputs the optimal weight adjustment action to dynamically adjust the scoring and penalty weights. At the same time, new experiences generated by the actual production of the factory will be continuously stored in the experience replay buffer. The system performs online fine-tuning of the weight optimization network based on the new experience to ensure that the network strategy can continuously adapt to changes in production mode and realize a dynamic closed loop between weight optimization and actual production.

[0184] Figure 4 This is a schematic diagram of the overall hanger target location selection, scoring, and overall scheduling process according to an embodiment of this application. The flowchart illustrates the complete process of overall hanger target location selection, scoring, and scheduling. Starting with the user's selection of a weight optimization strategy, three types of strategies are provided: a user-defined weight strategy (user-defined scoring / penalty factor weight ratios to adapt to personalized production needs), a time-based weight strategy (matching preset weights according to early shift start times, production peaks, etc., to match production targets for different time periods), and a weight optimization strategy based on a weight optimization network (dynamically adapting to the real-time system status through reinforcement learning). After the hangers are completed and leave the station, the system calls the scoring factors, scoring weights, penalty factors, and penalty weights corresponding to the selected strategy, substitutes them into the scoring formula to calculate the comprehensive score of each candidate station, then selects the station with the highest score as the target station, and finally executes the scheduling of hangers to that target station, realizing a closed loop from strategy selection to precise scheduling, ensuring that hanger allocation adapts to the production scenario and the dynamic state of the system.

[0185] Compared to the original fixed route allocation scheme, this application has three core advantages: First, it is more flexible. Factory users can deploy or import production tasks through the system interface according to the production plan. In the task setting and scheduling strategy configuration stage, the system provides three optional weight strategies. Users can choose according to actual production goals (such as quickly starting the production line and maximizing output) and management preferences to meet the production optimization needs in different scenarios. Second, the scheduling is more precise. It proposes a hanger target station scoring mechanism, which comprehensively considers key factors such as hanger path efficiency, station load, process urgency, station processing efficiency, and style process matching. The scoring formula quantifies the suitability of each candidate station to ensure that the hangers are scheduled to the optimal target station. Third, it is more adaptable. It adopts the DQN reinforcement learning algorithm to integrate the hanging system equipment information (such as station health) and production business information (such as order delivery time) in real time. It adaptively optimizes the scoring and penalty weight strategies, enabling the system to have long-term self-evolution capabilities and continuously adapt to dynamic production scenarios, ultimately improving the overall production efficiency of the entire hanging line. At the same time, it can also interface with third-party management systems through standardized API interfaces and SDKs to achieve cross-system collaborative adjustment of production strategies, further expanding the compatibility and scenario coverage of technology applications.

[0186] This embodiment also provides a scheduling device for a suspension system, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0187] Figure 5 This is a structural block diagram of the scheduling device of the suspension system according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0188] The determination module 51 is used to determine multiple candidate positions of the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate position.

[0189] The state space generation module 52 is used to generate the corresponding state space based on the current scoring weight, the current penalty weight, and the real-time global operating status information of the suspension system.

[0190] The weight optimization module 53 is used to input the state space into the trained weight optimization network to obtain the weight optimization strategy, and adjust the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight respectively.

[0191] The scoring calculation module 54 determines the score corresponding to each candidate station based on the target score weight, target penalty weight, score factor and penalty factor;

[0192] The hanger scheduling module 55 is used to determine the target station of the hanger to be scheduled based on the score corresponding to each candidate station, and to schedule the hanger to be scheduled to the target station.

[0193] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination. Specific examples in this embodiment can be found in the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0194] Furthermore, in conjunction with the scheduling method of the suspension system in the above embodiments, this application embodiment can provide a storage medium for implementation. The storage medium stores a computer program; when executed by a processor, the computer program implements any of the scheduling methods of the suspension system in the above embodiments.

[0195] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0196] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0197] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0198] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A scheduling method for a suspension system, characterized in that, include: Determine multiple candidate locations for the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate location; Based on the current scoring weight, the current penalty weight, and the real-time global operating status information of the suspension system, a corresponding state space is generated; The state space is input into the trained weight optimization network to obtain a weight optimization strategy. Based on the weight optimization strategy, the current rating weight and the current penalty weight are adjusted to obtain the target rating weight and the target penalty weight, respectively. Based on the target score weight, the target penalty weight, the score factor, and the penalty factor, the score corresponding to each candidate station is determined; Based on the scores corresponding to each candidate station, the target station for the hanger to be scheduled is determined, and the hanger to be scheduled is then scheduled to the target station.

2. The scheduling method for the suspension system according to claim 1, characterized in that, The training process of the weight optimization network includes multiple training rounds; Each training round includes multiple time steps; The training process for optimizing the network's weights for each training epoch includes: For each time step, generate the current state transition tuple and store the state transition tuple in a preset experience replay buffer; Randomly sample multiple state transition tuples from the experience replay buffer; For each state transition tuple, the current state space and the current weight adjustment action in the state transition tuple are input into the initial weight optimization network to obtain the predicted action value and determine the target action value corresponding to the state transition tuple. Based on the predicted action value and the target action value, the corresponding loss function result is determined, and the gradient of the loss function result is backpropagated to the initial weight optimization network to obtain the updated weight optimization network.

3. The scheduling method for the suspension system according to claim 2, characterized in that, The weight optimization network includes an evaluation network and a target network; The step of backpropagating the gradient of the loss function result to the initial weight optimization network to obtain the updated weight optimization network includes: The gradient of the loss function result is backpropagated to the initial evaluation network of the initial weight optimization network to obtain the updated evaluation network. After a preset number of time steps, the current parameters of the evaluation network are copied to the target network to obtain the updated target network. Based on the updated evaluation network and the updated target network, the updated weight optimization network is determined.

4. The scheduling method for the suspension system according to claim 3, characterized in that, Determining the target action value corresponding to the state transition tuple includes: When the current state transition tuple is the last state transition tuple among multiple randomly sampled state transition tuples, the reward result in the current state transition tuple is used as the target action value; When the current state transition tuple is not the last state transition tuple among multiple randomly sampled state transition tuples, the new state space in the current state transition tuple is input into the target network to predict the maximum action value, and the target action value is generated based on the reward result in the current state transition tuple and the maximum action value.

5. The scheduling method for the suspension system according to claim 2, characterized in that, The generation of the current state transition tuple includes: Select a weight adjustment action from the preset action space, and adjust the current scoring weight and penalty weight based on the weight adjustment action to obtain the first scoring weight and the first penalty weight; The first scoring weight and the first penalty weight are applied to the suspension system, and after the suspension system has been running for a preset period of time, the production data of the suspension system are statistically obtained. Determine the reward result corresponding to the production data, and the new state space corresponding to the production data; Based on the current state space, the weight adjustment action, the reward result, and the new state space, generate the current state transition tuple.

6. The scheduling method for the suspension system according to claim 5, characterized in that, The preset action space includes multiple discrete actions for adjusting weights; The sum of the first preset probability and the second preset probability is 1; The step of selecting weight adjustment actions from a preset action space includes: Based on the first preset probability, a discrete action is randomly selected from the preset action space as the weight adjustment action; or, Using the second preset probability, the discrete action that maximizes the action value of the evaluation network output in the weight optimization network is selected from the preset action space as the weight adjustment action.

7. The scheduling method for the suspension system according to claim 5, characterized in that, Determining the reward result corresponding to the production data includes: Identify multiple sub-production data sets within the production data; each sub-production data set includes at least one category of sub-production data. For each sub-production data set, determine the target reward corresponding to the sub-production data set; Based on the target reward corresponding to each of the sub-production data sets, the reward result corresponding to the production data is determined.

8. The scheduling method for the suspension system according to claim 5, characterized in that, The step of adjusting the current scoring weight and penalty weight based on the weight adjustment action to obtain the first scoring weight and the first penalty weight includes: Based on the weight adjustment action, the current scoring weight and penalty weight are adjusted to obtain the second scoring weight and the second penalty weight; The second scoring weight and the second penalty weight are normalized to obtain the first scoring weight and the first penalty weight.

9. The scheduling method for the suspension system according to claim 1, characterized in that, Before adjusting the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight, the method further includes: Obtain the current production time and find the target time period to which the current production time belongs from multiple preset time periods; Based on the target time period and the correspondence between each preset time period and the weighting strategy, the target weighting strategy corresponding to the target time period is determined. The target weight strategy is determined as the weight optimization strategy.

10. A scheduling device for a suspension system, characterized in that, The device includes: The determination module is used to determine multiple candidate positions of the clothes hangers to be scheduled on the hanging system, as well as the scoring factor, penalty factor, current scoring weight corresponding to the scoring factor, and current penalty weight corresponding to the penalty factor for each candidate position. The state space generation module is used to generate a corresponding state space based on the current scoring weight, the current penalty weight, and the real-time global operating state information of the suspension system. The weight optimization module is used to input the state space into the trained weight optimization network to obtain the weight optimization strategy, and adjust the current score weight and the current penalty weight according to the weight optimization strategy to obtain the target score weight and the target penalty weight respectively. The scoring calculation module determines the score corresponding to each of the candidate stations based on the target score weight, the target penalty weight, the score factor, and the penalty factor. The hanger scheduling module is used to determine the target station of the hanger to be scheduled based on the score corresponding to each candidate station, and to schedule the hanger to be scheduled to the target station.