Automatic logistics conveying line scheduling optimization method and device and related medium
By preprocessing multi-source data and making collaborative decisions with multiple agents, scheduling control messages are generated and verified, solving the problems of slow response and low resource utilization in traditional logistics conveyor scheduling systems under dynamic changes, and realizing real-time optimization and efficient scheduling.
Patent Information
- Application Number
- CN202511493400.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Traditional logistics conveyor scheduling systems struggle to cope with real-time changes in dynamic situations such as equipment failures and order surges, leading to conveyor congestion, low resource utilization, and slow response, lacking an end-to-end closed-loop adaptive scheduling mechanism.
By acquiring multi-source operational data for data preprocessing, constructing an environmental state vector sequence, calling a multi-agent collaborative decision engine to generate scheduling data, and finally generating scheduling control messages for the conveyor line controller through verification model validation and rule-based correction.
It improved the coordination of the logistics system, enabled real-time response and optimized scheduling to dynamic changes, reduced congestion on the conveyor lines, and improved resource utilization.
Smart Images

Figure CN120952283A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the application of large-scale model technology and the field of automated logistics equipment scheduling, and particularly to an automated logistics conveyor line scheduling optimization method, device and related media. Background Technology
[0002] With the rapid development of e-commerce and the increasing complexity of global supply chains, the logistics industry faces unprecedented challenges. Automated logistics conveyor lines, as the core infrastructure of modern warehousing and distribution centers, directly impact the throughput and cost of the overall logistics system. Traditional logistics conveyor line scheduling systems typically rely on preset rules or manual intervention, making it difficult to effectively cope with real-time changes in logistics demand, unexpected events (such as equipment failures and order surges), and complex conveyor path optimization problems. This leads to a series of pain points, including conveyor line congestion, low resource utilization, cargo backlog, and slow scheduling response.
[0003] While some existing technologies employ intelligent management systems, they struggle to uniformly absorb and make real-time decisions on multi-source heterogeneous operational data in dynamic situations such as equipment failures and order surges. This can easily lead to congestion on conveyor lines, low resource utilization, and slow response. Currently, there is a lack of end-to-end closed-loop adaptive scheduling mechanisms. Summary of the Invention
[0004] This invention provides an automated logistics conveyor scheduling optimization method, device, and related media, aiming to solve the problem of weak data perception in existing logistics systems, leading to poor coordination of logistics systems.
[0005] In a first aspect, embodiments of the present invention provide an automated logistics conveyor line scheduling optimization method, comprising: Acquire multi-source operational data and perform data preprocessing to obtain a standardized operational state sequence; The standardized operating state sequence is parameter-mapped to construct state features, resulting in an environmental state vector sequence. The multi-agent collaborative decision-making engine is invoked to perform policy proxy solution on the environmental state vector sequence to generate initial scheduling data; The initial scheduling data is input into the verification model for validation and then corrected according to rules to obtain the planned execution data; Based on the planned execution data, control instructions are parsed, equipment is mapped, and messages are encoded to generate a scheduling control message to be sent to the conveyor line controller. The scheduling control message to be sent is processed for transmission adaptation, and a scheduling control result set is output.
[0006] Secondly, embodiments of the present invention provide an automated logistics conveyor line scheduling optimization device, comprising: The data processing unit is used to acquire multi-source operational data and perform data preprocessing to obtain a standardized operational state sequence. The parameter mapping unit is used to perform parameter mapping on the standardized operating state sequence to construct state features and obtain an environmental state vector sequence. The strategy scheduling unit is used to call the multi-agent collaborative decision engine to perform strategy proxy solution on the environmental state vector sequence and generate initial scheduling data. The data verification unit is used to input the initial scheduling data into the verification model for verification and perform rule-based correction to obtain the planned execution data. The message generation unit is used to perform control instruction parsing, equipment mapping and message encoding according to the planned execution data, and generate a scheduling control message to be sent to the conveyor line controller. The scheduling output unit is used to perform transmission adaptation processing on the scheduling control message to be sent and output a scheduling control result set.
[0007] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the automated logistics conveyor scheduling optimization method of the first aspect.
[0008] Fourthly, embodiments of the present invention provide a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the automated logistics conveyor scheduling optimization method of the first aspect.
[0009] This invention provides an automated logistics conveyor line scheduling optimization method, including acquiring multi-source operational data and performing data preprocessing to obtain a standardized operational state sequence; mapping parameters to the standardized operational state sequence to construct state features, obtaining an environmental state vector sequence; calling a multi-agent collaborative decision engine to perform policy proxy solution on the environmental state vector sequence to generate initial scheduling data; inputting the initial scheduling data into a verification model for validation and performing rule-based correction to obtain planned execution data; performing control command parsing, equipment mapping, and message encoding based on the planned execution data to generate a scheduling control message to be sent to the conveyor line controller; performing transmission adaptation processing on the scheduling control message to be sent, and outputting a scheduling control result set. This invention processes the calculated planned execution data accordingly, outputs a scheduling control message to be sent, and performs transmission adaptation processing to obtain a scheduling control result set for controlling the logistics conveyor line, thus greatly improving the coordination of the logistics system.
[0010] This invention also provides an automated logistics conveyor scheduling optimization device, a computer device, and a storage medium, which have the same beneficial effects as described above. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating an automated logistics conveyor scheduling optimization method provided in an embodiment of the present invention; Figure 2 This is a schematic block diagram of an automated logistics conveyor line scheduling optimization device provided in an embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0015] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0016] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] Please see below. Figure 1 , Figure 1 A flowchart illustrating an automated logistics conveyor scheduling optimization method provided in this embodiment of the invention specifically includes steps S101~S106: S101. Acquire multi-source operational data and perform data preprocessing to obtain a standardized operational state sequence; S102. Perform parameter mapping on the standardized operating state sequence to construct state features and obtain an environmental state vector sequence; S103. Call the multi-agent collaborative decision-making engine to perform policy proxy solution on the environmental state vector sequence and generate initial scheduling data; S104. Input the initial scheduling data into the verification model for validation and perform rule-based correction to obtain the planned execution data; S105. Based on the planned execution data, perform control instruction parsing, equipment mapping and message encoding respectively to generate a scheduling control message to be sent to the conveyor line controller; S106. Perform transmission adaptation processing on the scheduling control message to be sent, and output the scheduling control result set.
[0018] This invention can perform end-to-end closed-loop scheduling for complex topologies formed by belt conveyors, roller conveyors, elevators, and diversion and merging nodes in warehousing and sorting centers.
[0019] In step S101, the system deploys RFID, photoelectric, vision, weighing, speed, encoder and vibration sensors at key nodes of the conveyor line to continuously collect raw data such as cargo position, speed, weight, size, line flow and equipment status; the sensing side can clean, fuse and extract features from the collected data, and obtain a standardized operating status sequence that can be directly consumed by upper-level decision-making after unifying the measurement and time axis.
[0020] In step S102, based on the conveyor topology and equipment list, the standardized time-series data is aligned to the "node-edge-workstation" structure (i.e., node-connection), and key quantities such as real-time flow, cargo distribution, queue length, equipment operating status and order priority of each segment are extracted. After time window statistics and missing imputation, the data is organized into an environmental state vector sequence according to the step time to meet the input requirements of subsequent strategy solving.
[0021] In one embodiment, step S102 includes: Based on the standardized operating state sequence, semantic annotation and classification are performed to obtain a type annotation sequence; Based on the conveyor line topology, each data item of the type-labeled sequence is mapped to a node, edge, and workstation identifier to obtain a topology-aligned data sequence. The topology-aligned data sequence is subjected to multimodal processing to obtain a set of multimodal state sub-vectors; Align and concatenate the set of multimodal state subvectors to obtain a time window statistical feature sequence; The time window statistical feature sequence is imputed for missing features and then continuous features are concatenated to obtain a fused feature tensor. The fused feature tensor is organized into a step input in chronological order, and then dimensional rearrangement and mask construction are performed to output an environment state vector sequence.
[0022] In this embodiment, a unified semantic and topological reference is first established for the multi-source data (including RFID, photoelectric, vision, weighing, speed sensor, encoder, and vibration sensor information) collected and preprocessed by the perception layer. This ensures that data from different modalities, sampling granularities, and equipment sources can be aligned on the same timeline and the same conveyor line topology, forming stable temporal feature inputs. The topology includes, but is not limited to, the physical layout and connection relationships of conveyor belts, roller conveyors, sorting machines, elevators, etc. The physical layout of the conveyor line can be planned in advance, and a physical location number for each conveyor line device can be provided (this number, also known as the device number in the scheduling system, is unique within the same project scenario). The aforementioned perception layer collection objects and preprocessing items include: deploying RFID, photoelectric, vision, weighing, speed / encoder, and vibration sensors at key nodes of the conveyor line to collect real-time data on cargo location, speed, weight, flow rate, and equipment status. Data cleaning, fusion, format conversion, outlier detection, and feature extraction are then performed to output a structured high-dimensional perception data stream, which serves as the basic input for this step. Nodes include: sorting points, merging points, and workstations (e.g., unmanned or manned workstations such as depalletizing, palletizing, picking, and inventory counting, which also include workstation attributes). In the semantic labeling and classification stage, the system identifies the type of standardized operating state sequences based on domain vocabulary and a pre-set dictionary, categorizing the original fields into semantic categories such as "location signal (e.g., RFID positioning, trajectory)," "transmission signal (photoelectric interruption counting, interval)," "appearance and volume characterization (vision, depth)," "load characterization (weighing)," "motion characterization (velocity sensor, encoder)," and "equipment health characterization (vibration)," resulting in a type-labeled sequence to clarify the role and constraint boundaries of each data item in subsequent feature construction. RFID is used for unique identification and trajectory recognition, photoelectric identification for counting and instantaneous speed estimation, vision for type, size, and anomaly detection, weighing for real-time load, velocity sensors and encoders for line speed and actual cargo movement speed, and vibration for key component status monitoring and fault early warning. The above semantic classifications correspond one-to-one with their sources.
[0023] During the topology alignment phase, the system assigns topological coordinates and equipment identifiers to semantically defined data items based on the conveyor line and the site topology (node-edge-workstation), achieving a three-dimensional mapping from "data source-physical location-functional unit": mapping RFID readers and their coverage areas to nodes and edges (i.e., the locations where each conveyor line connects to other conveyor lines or equipment, used to determine the direction of equipment execution); aligning photoelectric sensors to sorting ports, merging points, and cycle segments; aligning vision cameras to key identification and sorting workstations; binding weighing modules to specific weighing sections; associating speed sensors, encoders, and drive units; and binding vibration sensors to key components such as motors and bearings. After mapping, a unified topology index and clock stamp are written to all data items, outputting a topology-aligned data sequence to ensure that subsequent cross-modal combination and statistics are performed under the same topology-time reference.
[0024] In the multimodal processing stage, the system constructs state sub-vectors according to modality for the topology-aligned data sequence: displacement, dwell time, and repetition rate are extracted from RFID trajectories and state sequences; pass count, beat interval, and instantaneous velocity are extracted from photoelectric sequences; type encoding, three-dimensional size approximation, and abnormal event markers are extracted from visual sequences; net weight, weight fluctuation, and over-limit indicators are extracted from symmetrical weighted sequences; line rotation speed, acceleration / deceleration events, and slip indications are extracted from velocities and encoder sequences; and amplitude, spectral characteristics, and health scores are extracted from vibration sequences. After time synchronization and scale normalization, each modal sub-vector forms a multimodal state sub-vector set.
[0025] During the alignment and splicing stage, a sliding time window can be used to perform windowed statistics on the multimodal state sub-vector set. Statistical features such as mean, extreme values, variance, rate of change, quantiles, kurtosis, skewness, event count, and occupancy rate within the window are calculated and stacked in the order of nodes, edges, workstations, and time to obtain a time window statistical feature sequence. The window step size and length are configured according to the production cycle and on-site control cycle to ensure that the statistical features can cover the time scale of local congestion, fluctuations, and anomalies.
[0026] In the missing imputation and continuous feature splicing stage, the statistical feature sequence of the time window is subjected to missing test processing: for short-term missing data, forward or backward filling and linear interpolation are used; for long-term missing data, neighborhood topological regression or similar workstation migration estimation is used; and outliers are robustly truncated or quantile backoff is performed. Then, the imputed and corrected features are continuously spliced in the time dimension and topological dimension to generate a fusion feature tensor organized according to "time × topology × feature", which is used to carry the complete state expression before decision-making.
[0027] In the step input and mask construction phase, the fused feature tensor is arranged into step inputs in chronological order, and equal-length state slices are formed by combining the scheduling cycle. To adapt to the subsequent policy model, the tensor is dimensionally rearranged so that the time dimension, topological dimension, and feature dimension satisfy the model's desired reading order. Simultaneously, multiple types of masks are constructed, including: missing test masks (identifying the source of interpolation and the missing test confidence interval), shutdown masks (identifying nodes or edges under maintenance or inactive), no-cargo masks (identifying channels where no valid cargo passes through in the current window), and out-of-bounds masks (identifying specifications or overload limits). The step input sequence after dimensional rearrangement and mask injection is output as the environment state vector sequence, which is used by the multi-agent collaborative decision engine to solve the policy.
[0028] In step S103, a multi-agent collaborative decision-making engine (including a scheduling agent, an optimization agent, and a verification agent, combined with LLM for policy generation and correction) is used to solve the policy agent problem for the environment state vector. The decision layer can be achieved through the collaborative work of the scheduling agent, optimization agent, and verification agent: the scheduling agent estimates the current state and samples the policy based on a reinforcement learning model, producing candidate atomic actions (such as speed adjustment, traffic splitting, and priority adjustment); the optimization agent combines operations research, graph theory, and heuristic methods to solve and refine the candidate solutions for minimum delay routing, traffic splitting, and speed optimization; when necessary, a large language model (LLM) is used to perform semantic parsing of unstructured instructions and business constraints, providing supplementary information for policy shaping. After priority shaping, time window alignment, and continuity rearrangement, the cross-segment resource occupancy is merged into structured initial scheduling data.
[0029] Specifically, the multi-agent collaborative decision-making engine is based on the ReAct framework. The ReAct (Reasoning and Acting) framework is an AI agent design paradigm (not a specific framework project or finished product, but rather a design pattern or development philosophy). Its design philosophy is to achieve dynamic solutions to complex tasks through alternating "reasoning" and "acting" steps. The general execution flow of the framework is as follows: In each iteration, the ReAct agent first generates an internal reasoning trajectory based on Chain-of-Thought (CoT) to analyze the current state, evaluate alternatives, and formulate an action plan. Then, it executes specific actions, such as calling external APIs, querying knowledge bases, or operating environment interfaces (e.g., calling relevant algorithm models for decision-making, issuing task instructions to the PLC controller, etc.). The agent then observes the action results and injects feedback into the next round of reasoning, thus forming a closed-loop adaptive process.
[0030] In one embodiment, step S103 includes: A Markov decision process is constructed based on the environmental state vector sequence to obtain the Markov decision parameter set. The environmental state vector sequence is uniformly processed using the Markov decision parameter set to obtain the policy input tensor; The policy is input as a tensor into a scheduling agent network for estimation and policy sampling, and candidate atomic action sequences are output. Based on the candidate atomic action sequences, device quota constraints are applied to obtain candidate scheduling sequences; The candidate scheduling sequences are subjected to priority shaping, time window alignment, and continuity rearrangement to generate a set of structured scheduling segments; The structured scheduling fragment set is merged across resource usage segments to output initial scheduling data.
[0031] In this embodiment, the state space can be composed of elements such as real-time flow rate of each section of the conveyor line, cargo distribution, queue length of key nodes (such as sorting ports and merging points), equipment status, order priority, and historical congestion patterns; the action space includes discrete or continuous control actions such as adjusting the speed of specific sections (accelerating, decelerating), opening and closing diversion channels, changing cargo priority, and pausing and starting specific conveyor sections; the reward function can adopt "shortening transit time, improving conveyor line utilization, and ensuring on-time delivery" as positive objectives, and set penalty terms for congestion, cargo damage, and equipment failure to obtain a cumulative reward optimization objective (i.e., Markov decision parameter set) oriented towards maximizing throughput and minimizing delay.
[0032] The environmental state vector sequence is then standardized using a Markov decision parameter set to obtain the policy input tensor. The standardization process includes scaling, time alignment, and missing test mask injection for the state dimensions at different topological locations and time slices, and organizing the tensor into equal-length step segments according to the decision cycle. For unstructured special instructions from operators (e.g., temporary expedited processing, special pallet handling), intent recognition and semantic translation are performed using LLM to solidify them into computable priority prompts, prohibition or reachability constraints, or time window limits. These are then incorporated into the policy input tensor as additional contextual features to ensure that the policy solution is aware of business-side constraints.
[0033] Furthermore, the policy input tensor is input into the scheduling agent network (such as DQN) for estimation and policy sampling, outputting candidate atomic action sequences. The scheduling agent employs a deep reinforcement learning (DRL) algorithm, preferentially selecting a deep Q-network (DQN) or its variants (such as Double DQN, Dueling DQN, priority experience replay, etc.), calculating action-value estimates for each step segment and using trade-offs for policy sampling to obtain candidate atomic action sequences including speed adjustment, traffic splitting, priority adjustment, and segment-level start / stop control. In terms of training, offline pre-training is performed using historical and simulation data, followed by online small-batch fine-tuning based on real-time feedback to continuously adapt to on-site dynamics. Historical and simulation data include historical equipment feedback data, historical algorithm weight parameter information, and efficiency comparisons between historical feedback and post-tuning. Based on real-time operational data feedback, the weight parameter data can be automatically and adaptively fine-tuned, continuously approaching the optimal solution. Based on the candidate atomic action sequences, equipment quota constraints are applied to obtain candidate scheduling sequences. The constraints include at least: concurrent quotas, bandwidth and cycle time limits per device / channel / PLC cycle, safety intervals and occupancy conflict constraints for critical nodes, and derating operation boundaries triggered by equipment health status. For cross-segment coupled resource contention, the scheduling agent and optimization agent collaborate to implement capacity constraints and feasible region pruning on candidate atomic actions to obtain candidate scheduling sequences that meet physical and process constraints. The scheduling agent network draws inspiration from Deep Q-Network (DQN), with the specific formula as follows: ;
[0034] Where Q represents a function, and the reward R is... t =α·throughput increment - β·delay penalty, where α and β are weights, throughput increment represents the increase in the number of goods passing through the conveyor line (positive reward), and delay penalty represents the time it takes for goods to queue through the conveyor line or the deviation in transit time (negative penalty). The weight parameters are custom values and can be balanced through optimization.
[0035] Furthermore, the candidate scheduling sequences are subjected to priority shaping, time window alignment, and continuous rearrangement to generate a structured scheduling fragment set. Priority shaping is performed by hierarchically sorting orders based on urgency, estimated time of arrival (ETA), and business SLA; time window alignment is performed by boundary snapping and alignment according to control cycle and device response time to ensure consistency between message delivery and physical execution rhythm; continuous rearrangement is used to eliminate unnecessary switching and jitter, ensuring that control actions on the same topology path have the minimum number of switching operations and reasonable dwell time, thus obtaining a structured scheduling fragment set organized according to "time slice-topology segment-control command". The structured scheduling fragment set is then merged across segment resource usage to output initial scheduling data. During the merging process, the system performs interval merging and conflict resolution on resource usage of adjacent time slices and connected topology segments to ensure global consistency of actions across splitting / merging nodes under capacity and safety constraints; for fragments involving multi-device linkage, collaborative speed adjustment and path consistency verification are calculated to obtain initial scheduling data that can directly enter the subsequent verification stage, realizing a complete strategy generation closed loop from "estimation-sampling-pruning-shaping-merging".
[0036] In one embodiment, step S103 further includes: A multi-objective evaluation vector is constructed using the environmental state vector sequence to obtain a multi-objective optimization model; A weighted flow network is established based on the multi-objective optimization model, and the environmental state vector sequence is mapped to the corresponding path variables, flow splitting variables and velocity variables to obtain a graph model instance. Path set filtering and delay weight calculation are performed on the graph model instances respectively to obtain candidate path sets; The candidate path set and the multi-objective optimization model are combined to establish a linear programming problem to obtain the parameter-optimized solution. The optimized solution of the parameters is heuristically iteratively fine-tuned, and local recalculation is performed based on local congestion indications to obtain the iterative optimized solution; The iterative optimization solution is mapped to an optimized candidate scheduling sequence, and time-aligned with the environmental state vector sequence to output initial scheduling data.
[0037] In this embodiment, delay targets (transit time, queuing time), throughput targets (throughput per unit time), energy consumption targets (estimated impact of speed start-stop on energy consumption), and switching costs (penalties for frequent start-stop and traffic switching) are defined based on the business SLA and on-site operation strategy. Normalization and weight settings are applied to each target item (weights can be provided by historical statistics or strategy configuration) to obtain a multi-objective evaluation vector for solving the problem. Based on this vector, a multi-objective optimization model equivalent to "weighted summation" or "ε-constraint" is established. Then, a weighted flow network is built based on the multi-objective optimization model, and the environmental state vector sequence is mapped to corresponding path variables, traffic switching variables, and speed variables to obtain a graph model instance. Using the conveyor line topology as the graph basis, workstations as nodes, and conveyor sections as directed edges, attributes such as capacity, nominal speed, energy consumption coefficient, and safety interval are added to the edges. The demand and freight demand within the scheduling cycle are mapped to source-sink pairs. Path variables are defined to select candidate routes, flow distribution variables are defined to allocate traffic between parallel branches, and speed variables are defined to provide cooperative speed suggestions on adjustable speed edges. At the same time, capacity constraints, edge occupancy conflict constraints, and minimum or maximum speed boundaries are introduced to obtain a graph model instance with a physically feasible domain.
[0038] Furthermore, path set filtering and delay weight calculation are performed on the graph model instances to obtain a candidate path set. Preferably, the K-short circuit is calculated using a multi-objective variant of Dijkstra's or A* with a combined weight of "delay-energy consumption-switching cost". For each candidate path, the expected edge delay (including the influence of congestion coefficient and speed setting) and energy cost are estimated based on the environmental state vector sequence. The reachability and safety interval of the path across splitting or merging nodes are verified, filtering out infeasible or poor-quality paths and retaining the candidate path set that meets capacity or safety constraints.
[0039] Based on this, the candidate path set is combined with a multi-objective optimization model to establish a linear programming (or multi-objective linear programming) solution to obtain the parameter optimization solution. The objective function can be minimized by multi-objective weighted summation: Min J = Total delay + Energy consumption + • Switching cost− • Throughput; Constraints include: ① Flow conservation and source-sink satisfaction constraints; ② Edge capacity and concurrency quota constraints; ③ Equipment start-up / shutdown cycle time and safety interval constraints; ④ Coupling constraints between speed variables and path selection variables (such as linearized Big-M constraints). Solving yields optimized solutions for path selection, diversion ratio, and speed settings within the current time window. Heuristic iterative fine-tuning of the optimized solutions is then performed, and local recalculation is conducted based on local congestion indicators to obtain iterative optimized solutions. Specifically: Fine-grained perturbations of diversion ratio and speed settings are applied using a genetic algorithm or ant colony algorithm to evaluate the improvement magnitude of the target value and retain dominant individuals; when a local congestion indicator is detected (e.g., queue length on a certain edge or utilization exceeding the threshold), incremental recalculation of the affected subgraph is triggered, re-selecting K short-circuit paths and redistributing diversion ratios only within the local area, thereby achieving rapid congestion relief with lower computational cost. The above process is aligned with the scheduling cycle until the objective function converges or reaches the set number of iterations, at which point the iterative optimized solution is output. The specific formula for multi-objective linear programming is as follows: ; Where (x) represents the decision variable vector, including path variables, diversion variables, and velocity variables; f i f(x) represents the i-th objective function, f1(x) represents the delay, and f2(x) represents the energy consumption; w i Let A represent the weight coefficients, B represent the constraint matrix, and C represent the constraint right-hand vector (capacity constraint, equipment quota boundary). This formula can be used to transform multi-objective problems into single-objective linear programming problems.
[0040] Finally, the iterative optimization solution is mapped to the optimized candidate scheduling sequence and time-aligned with the environmental state vector sequence to output the initial scheduling data. The mapping rules include: translating path variables into specific traffic splitting, starting, stopping, and routing instructions; translating traffic splitting variables into the flow rate or release quota of each branch; translating speed variables into executable segment speed settings; performing boundary snapping and clock correction on adjacent time slices; merging continuous control segments on the same route to form a structured scheduling segment of "time slice-topology segment-control instruction"; and aligning it with the sensing-side clock stamp as initial scheduling data for processing by the verification agent and rule-making module.
[0041] In step S104, the verification agent performs parallel simulation evaluation of the initial scheme and implements multi-model voting, outputting consistency evaluation results for segments that may cause congestion or violations; combined with the safety process rule base, it applies rule-based correction to abnormal segments, and at the same time, the LLM performs semantic translation of temporary job requests or special tray processing requirements involving natural language and incorporates them into the constraint set, thereby performing constraint synthesis and updating on the initial scheduling data, and outputting plan execution data that can be directly issued.
[0042] In one embodiment, step S104 includes: The initial scheduling data is used to construct a verification input object, resulting in a verification input set. The validation input set is input into the corresponding model for parallel simulation to generate a multi-model evaluation result set. The consistency metric is calculated using the multi-model evaluation result set to obtain a consistency evaluation report; The consistency assessment report is input into the rule correction module for processing to obtain a set of correction rules. Based on the modified rule set, a large language model is invoked to perform semantic parsing and policy text translation on unstructured requests, and at the same time, it is merged with the modified rule set to obtain a modified constraint set; The initial scheduling data is constrained and synthesized using the modified constraint set, and the output plan execution data is updated.
[0043] In this embodiment, a verification input object is constructed using initial scheduling data to obtain a verification input set. Specifically, this includes: time windows, topology segments, control commands (such as segment speed settings, traffic diversion start / stop, priority adjustment), and their boundary conditions (capacity, concurrency quotas, safety intervals, equipment health status) for extracting scheduling segments; simultaneously, environmental state vector slices within the same window are loaded as scenario context to drive simulation and prediction. This verification input set fully expresses the ternary relationship of "scheme-environment-constraints," providing a unified input structure for subsequent parallel evaluation. The verification input set is then input into the corresponding models for parallel simulation, generating a multi-model evaluation result set. The verification agent maintains multiple complementary evaluators: discrete event simulation for replaying queue evolution and transit time; rule-based validators for rapid screening of strong constraints (such as hard boundaries for capacity, cycle time, and safety intervals); and prediction models for estimating congestion rate, delay distribution, and energy consumption trends. Each model runs independently and produces key indicators (e.g., average waiting time, node utilization, congestion probability, energy consumption estimation, and switching count), resulting in a multi-model evaluation result set. The consistency metric is then calculated using the multi-model evaluation result set to obtain a consistency evaluation report. The consistency of each evaluator's output can be measured according to the principle of "multi-model voting + threshold consistency": when core indicators such as congestion rate and latency are within safe thresholds and cross-model differences are within tolerable ranges, it is considered consistent; if significant deviations exist, the topological segment and time window causing the discrepancy are located, and the reasons for the inconsistency are marked. The report also lists the triggered hard constraints and their risk levels to guide subsequent rule-based corrections.
[0044] Furthermore, the consistency assessment report is input into the rule-based correction module for processing, resulting in a set of corrected rules. Based on compliance-related hard constraints and empirical soft constraints, the rule-based correction module generates structured correction suggestions for inconsistent or out-of-bounds scheduling segments, such as: reducing local segment speeds, delaying release windows, limiting concurrent quotas, and replacing paths via congested nodes. These suggestions are output as machine-readable rule entries (condition-action-scope-effective window), thus obtaining the corrected rule set. Subsequently, based on the corrected rule set, a large language model is invoked to perform semantic parsing and policy text translation on unstructured requests, and this is merged with the corrected rule set to obtain a set of corrected constraints. Unstructured requests can include: operator natural language instructions, such as contexts like "tray number XXX needs to be moved to the highest priority," and voice commands, such as using the open-source Whisper speech model to process operator voice input, converting it into natural language text. The natural language is then converted into unstructured text and sent to the large language model, thus continuing the closed loop. For temporary tasks, special cargo handling requirements, or abnormal handling instructions (natural language descriptions) from operators, LLM performs contextual understanding and intent recognition, transforming them into structured constraints such as priority adjustment, mandatory or prohibited paths, time-limited arrival, and maximum number of switching attempts. It then performs conflict detection and merging with the correction rule set, outputting a unified set of correction constraints to ensure that the business intent is accurately and executablely expressed at the scheduling end.
[0045] Finally, the initial scheduling data is constrained and synthesized using the modified constraint set, updating the output plan execution data. The constraint synthesis process checks and executes each initial scheduling segment, forcibly rewriting or replacing segments that do not meet hard constraints; fine-tuning the priority and time window of soft constraints using cost-based shaping; and reallocating execution capacity across segments and aligning with the clock cycle to ensure global feasibility and consistency with the clock cycle. After synthesis, plan execution data consistent with the field control cycle and directly coded into control messages is generated, serving as input for subsequent execution layer deployments.
[0046] In step S105, the system parses the planned execution data into instruction units for PLC channels, completes device identifier matching and address mapping; constructs the load structure and serializes and encodes it according to the control topic and interface specifications to obtain a set of control message frames; finally, it merges and groups them according to PLC channels to obtain scheduling control messages to be sent that are compatible with the field control network, thus realizing accurate mapping from strategy to physical execution interface.
[0047] In one embodiment, step S105 includes: Based on the control objectives corresponding to the planned execution data, and by performing instruction granularity segmentation, an instruction unit sequence is obtained; Based on the PLC controller channel, the instruction unit sequence is matched with device identifiers and mapped with address tags to obtain a device mapping instruction table; Perform topic mapping on the device mapping instruction table to obtain a protocol mapping table; The control payload is constructed according to the protocol mapping table, and the control payload is organized into a payload structure set according to the interface specification. The payload structure set is serialized and encoded to obtain a set of control message frames; The control message frame set is merged and grouped according to the PLC controller channel to output the dispatch control message to be sent to the conveyor line controller.
[0048] In this embodiment, the structured scheduling segments in the planned execution data are expanded one by one, and the topology segment identifier, control type, and target parameters are extracted. Abstract actions such as "segment speed setting," "diversion start / stop," "path switching," and "lifting / positioning" are refined into single, atomic control targets. Combining the field control cycle and equipment response time, adjacent control segments are snapped together and time-aligned to generate a sequence of instruction units that meets the field cycle requirements, providing clear control semantics and timing anchors for subsequent message generation. This parsing ensures that subsequent messages can drive the PLC to make fine-grained dynamic adjustments to speed, diversion, and path, thereby serving the goals of congestion avoidance and flow balancing. The channel configuration is then retrieved, and each instruction unit is assigned a unique device identifier and its address label in the corresponding PLC (such as register, coil, or data block index), and necessary safety and interlocking flags are loaded simultaneously. After planning is completed, the physical location of the equipment (e.g., from CAD drawings) will be marked with a unique number, which defaults to the unique identifier of the equipment. This data is an information set that can be obtained before system implementation. For segments involving multi-device linkage, occupancy conflict checks and capacity boundary verification are performed to ensure that the mapped instructions meet device concurrency quotas and safety interval constraints. After mapping, a device mapping instruction table of "device-address-action-parameter-time window" is formed, serving as the direct input for protocol layer encapsulation. Based on the field integration selection, matching standardized communication protocols and topic, node, and endpoint identifiers are selected for different PLC channels: when using message-based channels, mapping is done at the MQTT topic level; when using industrial interconnection channels, mapping is done at the OPC UA node path; when using HTTP or gateway channels, mapping is done at RESTful endpoints and resource paths. Topic mapping maintains a discriminable structure of "control object-action type-site line-time slice" and records transmission metadata such as service quality, timeout, and retry policies to obtain a protocol mapping table, supporting subsequent payload construction and orderly distribution.
[0049] Furthermore, control payloads are constructed according to the protocol mapping table, and then organized into payload structure sets according to the interface specifications. For speed setting actions, the payload must include at least the device address, target speed, effective time window, and out-of-bounds handling strategy; for traffic diversion and path switching actions, the payload must include the diversion identifier, switch status, target path, minimum hold duration, and switching cost identifier; for lifting and positioning actions, the payload must include the target level and position tolerance. All payloads are uniformly equipped with timestamps, serial numbers, and verification fields, and strictly adhere to the data model and field constraints of the selected protocol to form payload structure sets, ensuring that the instructions can be reliably executed by the PLC's built-in logic after arriving at the PLC. For MQTT / REST channels, serialization can be performed according to the agreed JSON or binary structure; for OPC UA channels, encoding and batch packaging are performed by writing to the model by node. During the encoding process, temporal order and idempotency are maintained to avoid duplicate execution; for multi-frame instructions linked across devices, batch markers are set to ensure arrival order and execution atomicity, ultimately generating a set of control message frames that can be directly transmitted over the network. Finally, the control message frame set is merged and grouped according to the PLC controller channel to output the dispatch control message to be sent to the conveyor line controller. The merging strategy uses "channel-station-control cycle" as the primary key, combining multiple frames of messages under the same cycle and the same PLC channel into a batch for distribution; for messages with sequential dependencies, a sequence relationship is added and an acknowledgment or retransmission strategy is configured. The merged dispatch control message is sent to the PLC through a standardized API. The PLC converts the digital instructions into electrical signals to drive the motors, splitters, combiners, and elevators, realizing dynamic and precise adjustment of the line parameters. After execution, it sends back feedback information such as actual speed, splitting status, and response time for subsequent closed-loop control and continuous learning.
[0050] In step S106, the message is transmitted and adapted before being sent to the PLC and actuator. Execution feedback generated by the execution layer, such as actual speed, diversion status, and equipment response time, is transmitted back to the perception and feedback layers and bound to the original scheduling records. The system writes key feedback elements into a vector database to establish a similarity index for subsequent retrieval and experience playback. Simultaneously, it calculates structured performance items (such as transit time, throughput, congestion rate, and energy consumption), associates them with the transmitted records, and outputs a scheduling control result set, closing the self-evolving loop of "perception-decision-execution-feedback." It is important to note that the actual controller controlling the logistics conveyor line is the PLC controller or other similar controllers, and the scheduling control result set is sent to the corresponding control units for execution.
[0051] In one embodiment, step S106 includes: Establish a transmission adapter object to identify and configure the message metadata in the scheduling control message to be sent, and generate a sending record for result binding; Based on the issued records, actual operational data is obtained and aggregated to obtain the original set of execution feedback. Extract the required interaction parameters from the original set of execution feedback to obtain a set of feedback elements; The feedback element set is written into a vector database and a similarity retrieval index is established to obtain a feedback vector index set; Based on the feedback vector index set, cargo parameter indicators are calculated to obtain structured result entries; The structured result entries are associated with the issued records to obtain the scheduling control result set.
[0052] In this embodiment, a transmission adapter object is established to identify and configure the metadata of the dispatch control messages to be sent, generating a dispatch record for result binding. For each batch of dispatched control messages, key information such as channel, station, control cycle, topology segment device mapping, instruction type, and timestamp is recorded to ensure a one-to-one correspondence with subsequent execution feedback, achieving a traceable binding relationship from "message-device-time window." This dispatch record serves as the anchor point for result aggregation and evaluation, connecting the execution layer's feedback layer's backhaul and the feedback layer's data entry process. Actual operational data is obtained based on the dispatch record, and the raw execution feedback set is aggregated. After the instruction is executed, the execution layer sends back key operational data such as actual speed, routing status, and device response time. The system extracts and aggregates data according to the time window and topology location indicated by the dispatch record to obtain the raw execution feedback set corresponding to the message batch, used for subsequent data cleaning, evaluation, and data entry. This backhaul mechanism ensures that the instruction execution effect can enter the feedback link in real time, supporting closed-loop control.
[0053] Furthermore, the necessary interaction parameters are extracted from the original set of execution feedback to obtain a set of feedback elements. The extracted elements include at least: key performance indicators (KPIs) closely related to scheduling quality, such as on-the-go time, throughput, congestion rate of each segment, equipment utilization, energy consumption, and abnormal fault markers, while retaining the context bound to the original message (decision fragment, action type, equipment topology index). This elementization process provides a standardized and structured representation for subsequent similarity retrieval and strategy review. The set of feedback elements is then written into a vector database, and a similarity retrieval index is established to obtain a set of feedback vector indexes. The system stores sensor-processed features, scheduling action records, execution results, reward signals, and necessary LLM interaction records in a high-dimensional vector format, and constructs a similarity index based on a nearest neighbor search structure. This index is used to quickly retrieve historically similar scenarios and their successful handling strategies when new congestion patterns or abnormal scenarios occur, providing data support for reinforcement learning and strategy optimization through "experience playback and case study learning."
[0054] Based on this, cargo parameter indicators are calculated using the feedback vector index set to obtain structured result entries. Cargo parameter indicators include cargo transit time deviation, congestion probability, etc. Combining similar historical entries indexed with current feedback elements, key parameters for single cargo, single time window, and single topology segment (such as cargo transit time and its deviation, path segment delay and congestion probability, unit throughput, energy consumption estimation, and equipment utilization) are calculated and stored, generating auditable structured result entries for direct consumption in operation and maintenance reports, alarm threshold assessment, and subsequent model training. Finally, the structured result entries are associated with the dispatch records to obtain the scheduling control result set. The system uses message batches and equipment topology indexes as keys to complete the closed-loop binding of "planning instructions - execution feedback - evaluation indicators," forming a result set output oriented towards multiple perspectives (time, equipment, path, order); this result set is also written back to the feedback layer, providing a continuous data source for the AI Agent's offline training, online fine-tuning, and model optimization, driving the system to achieve self-adaptation and self-evolution.
[0055] This invention relates to an automated logistics conveyor scheduling optimization system based on multi-agent systems, which operates according to a closed-loop mechanism of "perception-decision-execution-feedback-self-evolution." This corresponds one-to-one with steps S101 to S106 in the aforementioned technical solution, ensuring that the entire process from data acquisition to instruction implementation and effect feedback is calculable, verifiable, and traceable. The system first collects and preprocesses multi-source operational data at the perception layer, forming a standardized operational state sequence that can be directly consumed by the algorithm. Then, at the decision layer, this data is mapped into an environmental state vector sequence, driving multiple agents to collaboratively generate executable scheduling instructions. Next, the execution layer uses standardized APIs and PLC controllers to precisely convert abstract instructions into physical controls for the equipment. Finally, the feedback layer structurally stores the execution results and supports continuous learning, forming a stable cycle of "data-strategy-control-evaluation-optimization."
[0056] In the perception phase, IoT sensors (such as RFID, photoelectric, vision, and weighing sensors) collect key data in real time, including cargo location, speed, weight, dimensions, conveyor flow rate, and equipment status. After cleaning, fusion, and feature extraction, structured perception data is obtained. This data is aligned on the time axis and topological coordinates (nodes, edges, workstations) and input into the parameter mapping process to construct an environmental state vector sequence that fully represents the current on-site operating status, providing accurate and timely input for subsequent intelligent decision-making.
[0057] In the intelligent decision-making and collaboration phase, the scheduling agent uses a reinforcement learning model to evaluate the environmental state and sample policies, providing a preliminary scheduling scheme aimed at improving throughput and reducing latency. The optimization agent combines operations research, graph theory, and heuristic algorithms to globally refine and weigh path selection, traffic splitting ratios, and segment speed settings. The verification agent employs multi-model parallel simulation and a voting mechanism to ensure security, feasibility, and constraint consistency. When unstructured business requests or abnormal scenarios exist, LLM parses natural language intent and translates policies, injecting business-side constraints into the decision-making chain in a structured form. The final output is planned execution data that can be directly incorporated into the control chain.
[0058] In the command execution and control phase, planned execution data is transmitted to the PLC controller via standardized APIs (such as MQTT, OPC UA, and RESTful API), and is converted into specific electrical signals for actuators such as motors, shunts, combiners, and elevators, enabling dynamic and precise adjustments to speed, direction, and shunt switches. Information such as actual speed, shunt status, and equipment response time generated during execution is synchronously transmitted back, providing a basis for closed-loop control and rapid anomaly handling.
[0059] In the feedback and continuous learning phase, the system writes the execution results and their key performance indicators (such as travel time, throughput, congestion rate, equipment utilization, energy consumption, etc.), along with the raw perception data, policy decision records, and reward signals, into a vector database and establishes a similarity index. On the one hand, this enables the system to quickly retrieve similar historical scenarios and successful strategies as "experience replay" when encountering new congestion or abnormal patterns; on the other hand, it also provides a unified and reusable data foundation for offline training, online fine-tuning, and policy optimization of the model.
[0060] In the self-evolution phase, the AI Agent at the decision-making level triggers retraining periodically or when performance degrades or new patterns emerge: offline, it performs large-scale training using accumulated data to obtain a better global strategy; online, it uses incremental data for small-batch fine-tuning to quickly adapt to short-term disturbances and unexpected events. With the help of the above closed-loop and learning mechanisms, the system continuously improves scheduling efficiency and robustness in long-term operation, stably achieving autonomous and flexible production goals.
[0061] This invention's system has a wide range of applications. In automated logistics integration scenarios, addressing the complex conveyor topologies of large warehouses, smart factories, and sorting centers, the system can achieve seamless collaboration and global scheduling between different conveying equipment and sorting gear, optimizing the overall process and improving material turnover efficiency and throughput. In e-commerce warehousing and distribution center scenarios, the system can handle the business characteristics of large order volumes, significant peak and trough periods, and high timeliness, supporting rapid sorting, precise merging, and efficient outbound delivery, shortening fulfillment cycles and improving service quality. In manufacturing production line material distribution scenarios, the system can intelligently allocate materials based on production rhythm, real-time demand, and workstation inventory, ensuring "Just-In-Time" delivery of materials, reducing work-in-process inventory, and maintaining continuous and efficient production line operation.
[0062] In summary, this invention constructs a closed-loop system of "perception-decision-execution-feedback-self-evolution," deeply integrating multi-source high-precision perception data with multi-agent collaborative decision-making. This enables unmanned and autonomous scheduling in complex and dynamic transportation environments. Specifically: First, through the linkage of the AI Agent's ReAct framework and reinforcement learning model, the system can automatically plan paths, adjust traffic distribution and segment speeds under real-time data-driven conditions, continuously learn and iteratively optimize, significantly improving throughput, reducing in-transit and queuing latency, and reducing manual intervention and rule maintenance costs. Second, leveraging a collaborative architecture composed of "scheduling agent-optimization agent-verification agent-LLM assistance," operations research, graph theory, and heuristic algorithms are used to globally refine the initial plan. Combined with multi-model voting verification and security constraints, this makes decisions more comprehensive and reliable, avoiding local optima and erroneous deployments. Third, LLM is introduced to perform semantic parsing and policy translation of unstructured business requests and anomaly descriptions, enhancing policy interpretability and human-machine collaboration. The system possesses the capability to quickly and accurately implement solutions for scenarios such as urgent requests and special cargo handling; fourth, it ensures accurate and timely status representation through high-precision perception and standardized preprocessing using IoT sensors such as RFID, vision, and photoelectric sensors, enabling millisecond-level response and precise control to potential congestion; fifth, it accumulates end-to-end data on "perception-policy-execution-reward" using a vector database, supporting rapid retrieval of similar scenarios and experience playback, accelerating model training and online fine-tuning while ensuring continuous performance improvement over time; sixth, through deep integration of standardized APIs and PLCs, it stably maps abstract policies into executable control messages, flexibly adjusting speed, path, and diversion strategies to balance efficiency, energy consumption, and equipment safety.
[0063] Combination Figure 2 As shown, Figure 2 This is a schematic block diagram of an automated logistics conveyor line scheduling optimization device provided in an embodiment of the present invention. The automated logistics conveyor line scheduling optimization device 200 includes: Data processing unit 201 is used to acquire multi-source operating data and perform data preprocessing to obtain a standardized operating state sequence; Parameter mapping unit 202 is used to perform parameter mapping on the standardized operating state sequence to construct state features and obtain an environmental state vector sequence; The strategy scheduling unit 203 is used to call the multi-agent collaborative decision engine to perform strategy proxy solution on the environmental state vector sequence and generate initial scheduling data; Data verification unit 204 is used to input the initial scheduling data into the verification model for verification and perform rule-based correction to obtain planned execution data; The message generation unit 205 is used to perform control instruction parsing, equipment mapping and message encoding according to the planned execution data, and generate a scheduling control message to be sent to the conveyor line controller. The scheduling output unit 206 is used to perform transmission adaptation processing on the scheduling control message to be sent and output a scheduling control result set.
[0064] In this embodiment, the data processing unit 201 acquires multi-source operating data and performs data preprocessing to obtain a standardized operating state sequence; the parameter mapping unit 202 performs parameter mapping on the standardized operating state sequence to construct state features and obtain an environmental state vector sequence; the strategy scheduling unit 203 calls a multi-agent collaborative decision engine to perform strategy proxy solution on the environmental state vector sequence and generates initial scheduling data; the data verification unit 204 inputs the initial scheduling data into a verification model for verification and performs rule-based correction to obtain planned execution data; the message generation unit 205 performs control instruction parsing, equipment mapping, and message encoding according to the planned execution data to generate a scheduling control message to be sent to the conveyor line controller; and the scheduling output unit 206 performs transmission adaptation processing on the scheduling control message to be sent and outputs a scheduling control result set.
[0065] In one embodiment, the data processing unit 201 is specifically used for: Based on the standardized operating state sequence, semantic annotation and classification are performed to obtain a type annotation sequence; Based on the conveyor line topology, each data item of the type-labeled sequence is mapped to a node, edge, and workstation identifier to obtain a topology-aligned data sequence. The topology-aligned data sequence is subjected to multimodal processing to obtain a set of multimodal state sub-vectors; Align and concatenate the set of multimodal state subvectors to obtain a time window statistical feature sequence; The time window statistical feature sequence is imputed for missing features and then continuous features are concatenated to obtain a fused feature tensor. The fused feature tensor is organized into a step input in chronological order, and then dimensional rearrangement and mask construction are performed to output an environment state vector sequence.
[0066] In one embodiment, the policy scheduling unit 203 is specifically used for: A Markov decision process is constructed based on the environmental state vector sequence to obtain the Markov decision parameter set. The environmental state vector sequence is uniformly processed using the Markov decision parameter set to obtain the policy input tensor; The policy is input as a tensor into a scheduling agent network for estimation and policy sampling, and candidate atomic action sequences are output. Based on the candidate atomic action sequences, device quota constraints are applied to obtain candidate scheduling sequences; The candidate scheduling sequences are subjected to priority shaping, time window alignment, and continuity rearrangement to generate a set of structured scheduling segments; The structured scheduling fragment set is merged across resource usage segments to output initial scheduling data.
[0067] In one embodiment, the policy scheduling unit 203 is further specifically used for: A multi-objective evaluation vector is constructed using the environmental state vector sequence to obtain a multi-objective optimization model; A weighted flow network is established based on the multi-objective optimization model, and the environmental state vector sequence is mapped to the corresponding path variables, flow splitting variables and velocity variables to obtain a graph model instance. Path set filtering and delay weight calculation are performed on the graph model instances respectively to obtain candidate path sets; The candidate path set and the multi-objective optimization model are combined to establish a linear programming problem to obtain the parameter-optimized solution. The optimized solution of the parameters is heuristically iteratively fine-tuned, and local recalculation is performed based on local congestion indications to obtain the iterative optimized solution; The iterative optimization solution is mapped to an optimized candidate scheduling sequence, and time-aligned with the environmental state vector sequence to output initial scheduling data.
[0068] In one embodiment, the data verification unit 204 is specifically used for: The initial scheduling data is used to construct a verification input object, resulting in a verification input set. The validation input set is input into the corresponding model for parallel simulation to generate a multi-model evaluation result set. The consistency metric is calculated using the multi-model evaluation result set to obtain a consistency evaluation report; The consistency assessment report is input into the rule correction module for processing to obtain a set of correction rules. Based on the modified rule set, a large language model is invoked to perform semantic parsing and policy text translation on unstructured requests, and at the same time, it is merged with the modified rule set to obtain a modified constraint set; The initial scheduling data is constrained and synthesized using the modified constraint set, and the output plan execution data is updated.
[0069] In one embodiment, the message generation unit 205 is specifically used for: Based on the control objectives corresponding to the planned execution data, and by performing instruction granularity segmentation, an instruction unit sequence is obtained; Based on the PLC controller channel, the instruction unit sequence is matched with device identifiers and mapped with address tags to obtain a device mapping instruction table; Perform topic mapping on the device mapping instruction table to obtain a protocol mapping table; The control payload is constructed according to the protocol mapping table, and the control payload is organized into a payload structure set according to the interface specification. The payload structure set is serialized and encoded to obtain a set of control message frames; The control message frame set is merged and grouped according to the PLC controller channel to output the dispatch control message to be sent to the conveyor line controller.
[0070] In one embodiment, the scheduling output unit 206 is configured to: Establish a transmission adapter object to identify and configure the message metadata in the scheduling control message to be sent, and generate a sending record for result binding; Based on the issued records, actual operational data is obtained and aggregated to obtain the original set of execution feedback. Extract the required interaction parameters from the original set of execution feedback to obtain a set of feedback elements; The feedback element set is written into a vector database and a similarity retrieval index is established to obtain a feedback vector index set; Based on the feedback vector index set, cargo parameter indicators are calculated to obtain structured result entries; The structured result entries are associated with the issued records to obtain the scheduling control result set.
[0071] Since the embodiments of the apparatus and the embodiments of the method correspond to each other, please refer to the description of the embodiments of the method for the embodiments of the apparatus, which will not be repeated here.
[0072] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0073] This invention also provides a computer device, which may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the computer device may also include various network interfaces, a power supply, a graphics card, etc., to utilize the graphics card's performance to operate the model, such as for inference and training.
[0074] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0075] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for optimizing the scheduling of an automated logistics conveyor line, characterized in that, include: Acquire multi-source operational data and perform data preprocessing to obtain a standardized operational state sequence; The standardized operating state sequence is parameter-mapped to construct state features, resulting in an environmental state vector sequence. The multi-agent collaborative decision-making engine is invoked to perform policy proxy solution on the environmental state vector sequence to generate initial scheduling data; The initial scheduling data is input into the verification model for validation and then corrected according to rules to obtain the planned execution data; Based on the planned execution data, control instructions are parsed, equipment is mapped, and messages are encoded to generate a scheduling control message to be sent to the conveyor line controller. The scheduling control message to be sent is processed for transmission adaptation, and a scheduling control result set is output.
2. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The step of mapping parameters to the standardized operating state sequence to construct state features and obtain an environmental state vector sequence includes: Based on the standardized operating state sequence, semantic annotation and classification are performed to obtain a type annotation sequence; Based on the conveyor line topology, each data item of the type-labeled sequence is mapped to a node, edge, and workstation identifier to obtain a topology-aligned data sequence. The topology-aligned data sequence is subjected to multimodal processing to obtain a set of multimodal state sub-vectors; Align and concatenate the multimodal state subvector set to obtain a time window statistical feature sequence; The time window statistical feature sequence is imputed for missing features and then continuous features are concatenated to obtain a fused feature tensor. The fused feature tensor is organized into a step input in chronological order, and then dimensional rearrangement and mask construction are performed to output an environment state vector sequence.
3. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The step of invoking the multi-agent collaborative decision-making engine to perform policy proxy solution on the environmental state vector sequence and generate initial scheduling data includes: A Markov decision process is constructed based on the environmental state vector sequence to obtain the Markov decision parameter set. The environmental state vector sequence is uniformly processed using the Markov decision parameter set to obtain the policy input tensor; The policy is input as a tensor into a scheduling agent network for estimation and policy sampling, and candidate atomic action sequences are output. Based on the candidate atomic action sequences, device quota constraints are applied to obtain candidate scheduling sequences; The candidate scheduling sequence is subjected to priority shaping, time window alignment, and continuity rearrangement to generate a set of structured scheduling fragments; The structured scheduling fragment set is merged across resource usage segments to output initial scheduling data.
4. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The step of calling the multi-agent collaborative decision engine to perform policy proxy solution on the environmental state vector sequence and generate initial scheduling data also includes: A multi-objective evaluation vector is constructed using the environmental state vector sequence to obtain a multi-objective optimization model; A weighted flow network is established based on the multi-objective optimization model, and the environmental state vector sequence is mapped to the corresponding path variables, flow splitting variables and velocity variables to obtain a graph model instance. Path set filtering and delay weight calculation are performed on the graph model instances respectively to obtain candidate path sets; The candidate path set and the multi-objective optimization model are combined to establish a linear programming problem to obtain the parameter-optimized solution. The optimized solution of the parameters is heuristically iteratively fine-tuned, and local recalculation is performed based on local congestion indications to obtain the iterative optimized solution; The iterative optimization solution is mapped to an optimized candidate scheduling sequence, and time-aligned with the environmental state vector sequence to output initial scheduling data.
5. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The step of inputting the initial scheduling data into the verification model for validation and performing rule-based correction to obtain planned execution data includes: The initial scheduling data is used to construct a verification input object, resulting in a verification input set. The validation input set is input into the corresponding model for parallel simulation to generate a multi-model evaluation result set. The consistency metric is calculated using the multi-model evaluation result set to obtain a consistency evaluation report; The consistency assessment report is input into the rule correction module for processing to obtain a set of correction rules. Based on the modified rule set, a large language model is invoked to perform semantic parsing and policy text translation on unstructured requests, and at the same time, it is merged with the modified rule set to obtain a modified constraint set; The initial scheduling data is constrained and synthesized using the modified constraint set, and the output plan execution data is updated.
6. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The process of parsing control commands, mapping equipment, and encoding messages based on the planned execution data to generate a scheduling control message to be sent to the conveyor line controller includes: Based on the control objectives corresponding to the data parsing of the plan, and by performing instruction granularity segmentation, an instruction unit sequence is obtained; Based on the PLC controller channel, the instruction unit sequence is matched with device identifiers and mapped with address tags to obtain a device mapping instruction table; Perform topic mapping on the device mapping instruction table to obtain a protocol mapping table; The control payload is constructed according to the protocol mapping table, and the control payload is organized into a payload structure set according to the interface specification. The payload structure set is serialized and encoded to obtain a set of control message frames; The control message frame set is merged and grouped according to the PLC controller channel to output the dispatch control message to be sent to the conveyor line controller.
7. The automated logistics conveyor line scheduling optimization method according to claim 1, characterized in that, The transmission adaptation processing of the scheduling control message to be sent, and the output of the scheduling control result set, includes: Establish a transmission adapter object to identify and configure the message metadata in the scheduling control message to be sent, and generate a sending record for result binding; Based on the issued records, actual operational data is obtained and aggregated to obtain the original set of execution feedback. Extract the required interaction parameters from the original set of execution feedback to obtain a set of feedback elements; The feedback element set is written into a vector database and a similarity retrieval index is established to obtain a feedback vector index set; Based on the feedback vector index set, cargo parameter indicators are calculated to obtain structured result entries; The structured result entries are associated with the issued records to obtain the scheduling control result set.
8. An automated logistics conveyor line scheduling optimization device, characterized in that, include: The data processing unit is used to acquire multi-source operational data and perform data preprocessing to obtain a standardized operational state sequence. The parameter mapping unit is used to perform parameter mapping on the standardized operating state sequence to construct state features and obtain an environmental state vector sequence. The strategy scheduling unit is used to call the multi-agent collaborative decision engine to perform strategy proxy solution on the environmental state vector sequence and generate initial scheduling data. The data verification unit is used to input the initial scheduling data into the verification model for verification and perform rule-based correction to obtain the planned execution data. The message generation unit is used to perform control instruction parsing, equipment mapping and message encoding according to the planned execution data, and generate a scheduling control message to be sent to the conveyor line controller. The scheduling output unit is used to perform transmission adaptation processing on the scheduling control message to be sent and output a scheduling control result set.
9. A computer device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the automated logistics conveyor scheduling optimization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the automated logistics conveyor scheduling optimization method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Systems, methods, and apparatus for sharing tool manufacturing and design data
CN114879598A
Power wireless local area network multi-domain data processing method and system based on credible authentication
CN120568337A
Construction method of double-path collaborative decision network for multi-agent collaborative path optimization
CN120598148A
Systems and methods for supporting network slice services using transport devices
US20240163724A1
Cited By
Workshop logistics sorting virtual debugging system based on artificial intelligence and implementation method
CN121411374A
Cross-warehouse scheduling migration method and device, equipment and medium
CN122048248A
Migration method and device for cross-warehouse scheduling, equipment and medium
CN122048248B
Context awareness type intelligent production scheduling decision-making method and system and medium
CN122089011A