Intelligent goods location pre-allocation method and system based on four-layer closed-loop governance architecture
The intelligent location pre-allocation method based on a four-layer closed-loop governance architecture solves the problem of AI agents lacking physical constraints and closed-loop feedback in automated warehouses. It enables reliable and adaptive location allocation decisions and fully automated operation, improving the system's reliability and maintainability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN TODAY INT SOFTWARE TECH CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-24
AI Technical Summary
Existing AI agents lack physical constraints, closed-loop feedback correction capabilities, auditable decision-making processes, and fully automated operation guarantees in the allocation of storage locations in automated warehouses.
A smart storage location pre-allocation method based on a four-layer closed-loop governance architecture is adopted, including feedforward constraint injection, intelligent decision-making and confidence assessment, execution deviation feedback, hierarchical self-correction and automatic system rollback. By actively injecting the AI agent's decision context through a structured rule set, physical feasibility, adaptive optimization and fully automated operation are achieved.
To ensure the reliability and security of AI agent decision-making, achieve continuous adaptive optimization of allocation strategies, reduce system jitter, ensure the continuity and auditability of fully automated operation, and improve the maintainability and reliability of the system.
Smart Images

Figure CN122453330A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent scheduling technology for warehousing and logistics, and in particular to an intelligent pre-allocation method and system for cargo locations based on a four-layer closed-loop governance architecture. Background Technology
[0002] With the rapid development of e-commerce and intelligent manufacturing, automated storage and retrieval systems (AS / RS) have become the core infrastructure of modern warehousing and logistics systems. Location allocation is a key factor determining warehouse storage efficiency, inbound and outbound operation efficiency, and stacker crane energy consumption. Modern AS / RS have three structural configurations: pure single-depth, pure double-depth, and a hybrid of single-depth and double-depth, which introduces complexity to location allocation.
[0003] Currently, automated warehouse location allocation technologies are mainly divided into two categories: The first is based on traditional rules or operations research optimization methods. Rule-based methods (such as the ABC classification fixed area method) are simple to implement but difficult to adapt to dynamic changes; optimization algorithm-based methods (such as genetic algorithms) can theoretically obtain better solutions, but they have high computational complexity, are difficult to respond in real time, and lack the ability to handle unstructured physical constraints (such as topological occlusion in double-deep locations). The second is the AI agent method that has emerged in recent years. However, directly applying AI agents to location allocation has the following drawbacks: lack of guarantees for physical world constraints (such as accessibility of double-deep locations, stacker crane kinematics); lack of closed-loop correction capabilities based on execution feedback; the decision-making process is usually a black box and cannot be audited; when AI decisions are abnormal, there is a lack of automated protection mechanisms, either relying on manual intervention (reducing efficiency) or continuing to run automatically (lacking security). Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide an intelligent storage location pre-allocation method and system based on a four-layer closed-loop governance architecture. This aims to solve the technical problems that existing AI agents, when directly applied to storage location allocation in automated warehouses, lack physical world constraint guarantees, lack closed-loop feedback correction capabilities, lack of auditability of the decision-making process, and lack of fully automated operation guarantees.
[0005] Firstly, a method for intelligent warehouse location pre-allocation based on a four-layer closed-loop governance architecture is provided, comprising the following steps: Feedforward constraint injection: The physical world constraints of warehouse location allocation in an automated warehouse are encoded into a structured rule set through a feedforward guidance layer, and the corresponding rules of the structured rule set are injected into the decision context of the AI agent before each warehouse location pre-allocation decision. The structured rule set is divided into a subset of rigid constraint rules and a subset of flexible constraint rules; Intelligent decision-making and confidence assessment: Under the guidance of the decision context, the AI agent outputs a warehouse location pre-allocation scheme and a unified confidence level for the selected optimal warehouse location; Execution deviation feedback: The actual execution deviation of the warehouse location pre-allocation scheme output by the AI agent is detected by a feedback sensing layer, and a unified deviation magnitude is calculated; Hierarchical self-correction: The self-correction layer converts the unified deviation... The amplitude is compared with the first and second thresholds. Following a dual-threshold, three-mode mechanism, the system automatically selects one of three correction modes—parameter fine-tuning, strategy rollback, and complete replanning—and updates the strategy version used by the AI agent in the next pre-allocation decision. Automatic system rollback: When the unified confidence level falls below the confidence threshold, the upper-level system rollback gating layer automatically switches the cargo allocation control process from the AI agent decision mode to the upper-level system's default allocation mode, where the deterministic allocation algorithm built into the upper-level system executes the cargo allocation. Full-link observation: The entire process data—feedforward constraint injection, intelligent decision-making and confidence assessment, execution deviation feedback, hierarchical self-correction, and automatic system rollback—is recorded in a cross-sectional manner through the full-link observable layer. This full-link observable layer does not participate in decision-making or control.
[0006] Secondly, a smart warehouse location pre-allocation system based on a four-layer closed-loop governance architecture is provided, comprising: a feedforward guidance layer, used to encode the physical world constraints of warehouse location allocation into a structured rule set, and inject the corresponding rules of the structured rule set into the decision context of an AI agent before each warehouse location pre-allocation decision, wherein the structured rule set is divided into a subset of rigid constraint rules and a subset of flexible constraint rules; an AI agent, used to output a warehouse location pre-allocation scheme and a unified confidence level of the selected optimal warehouse location under the guidance of the decision context; a feedback sensing layer, used to detect the actual execution deviation of the warehouse location pre-allocation scheme output by the AI agent and calculate the unified deviation amplitude; and a self-correction layer, used to compare the unified deviation amplitude with a first threshold and a... A second threshold is compared, and a dual-threshold, three-mode mechanism is used to automatically select one of three correction modes—parameter fine-tuning, strategy rollback, and complete replanning—to execute, thereby updating the strategy version used by the AI agent in the next execution of the pre-allocation decision. A back-off gating layer in the upper-level system is used to automatically switch the cargo allocation control process from the AI agent's decision mode to the upper-level system's default allocation mode when the unified confidence level is lower than the confidence threshold, with the cargo allocation executed by the deterministic allocation algorithm built into the upper-level system. A full-link observable layer is used to record the entire process data of the feedforward guidance layer, feedback sensing layer, self-correction layer, upper-level system back-off gating layer, and AI agent in a cross-sectional manner. This full-link observable layer does not participate in decision-making or control.
[0007] The aforementioned intelligent storage location pre-allocation method and system based on a four-layer closed-loop governance architecture proactively injects a structured rule set separating rigid and flexible constraints into the AI agent's decision-making context through a feedforward guidance layer. This fundamentally solves the physical feasibility problem of the AI agent in physical execution scenarios, ensuring the reliability of the AI agent's decisions. The feedback sensing layer and self-correction layer work together to form a closed loop, enabling the system to automatically correct the allocation strategy based on actual execution deviations, achieving continuous adaptive optimization of the allocation strategy. The dual-threshold, three-mode hierarchical correction mechanism allows the system to automatically select the most economical repair path based on the severity of the deviation. Compared to a single full replanning approach, it significantly reduces system jitter; through the upper-level system rollback gating layer based on unified confidence, the system automatically rolls back, automatically switching to the upper-level system's default algorithm without manual intervention when the AI agent makes an abnormal decision, and automatically switching back after the AI agent recovers, ensuring the continuity and security of fully automated operation; through a cross-cutting, end-to-end observable layer, it completely records the entire chain of data from decision triggering to final execution, making the basis, process, result, and subsequent adjustments of each allocation decision traceable, achieving auditability of the complete decision chain, and greatly improving the system's maintainability and reliability. Attached Figure Description
[0008] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of the overall architecture of an intelligent cargo location pre-allocation system based on a four-layer closed-loop governance architecture according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating an intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] This invention provides an intelligent storage location pre-allocation method and system based on a four-layer closed-loop governance architecture, aiming to solve the technical problems of existing AI agents lacking physical world constraints, closed-loop feedback correction capabilities, unauditable decision-making processes, and fully automated operation guarantees when directly applied to storage location allocation in automated warehouses.
[0012] This invention is applicable to three types of warehouse structures: pure single-deep, pure double-deep, and mixed. The intelligent warehouse location pre-allocation system based on a four-layer closed-loop governance architecture is connected to the upper-level system of the automated warehouse and is used to control the stacker crane to perform the inbound allocation action of goods.
[0013] In this invention, the "supervisor system" refers to the collective term for the upper-level management and control system in an automated warehouse that sits above the AI intelligent agent and is responsible for overall warehousing operations and equipment control. This includes, but is not limited to, Warehouse Management System (WMS), Warehouse Execution System (WES), Warehouse Control System (WCS), and integrated warehouse management platforms, warehouse middleware, and warehouse scheduling systems with similar functions. The functions undertaken by the superior system include, but are not limited to, order management, inventory management, task scheduling, stacker crane path planning, and real-time motion control of stacker cranes and conveyor lines. In specific implementations, the above functions can be implemented by a single integrated system or by multiple subsystems such as WMS, WES, and WCS working together. This invention does not limit the specific implementation form of the superior system.
[0014] Figure 1 This is a schematic diagram of the overall architecture of an intelligent storage location pre-allocation system based on a four-layer closed-loop governance architecture, as provided in one embodiment of the present invention. Figure 1 As shown, the intelligent storage location pre-allocation system based on a four-layer closed-loop governance architecture provided by this invention consists of an AI agent and a four-layer closed-loop governance architecture, with audit support provided by a transverse, end-to-end observable layer (this end-to-end observable layer does not participate in any decision-making or control, but only undertakes data recording functions, and is therefore not included in the four-layer closed-loop governance architecture to avoid semantic confusion in the architecture hierarchy). The four-layer closed-loop governance architecture includes, in sequence, a feedforward guidance layer, a feedback sensing layer, a self-correction layer, and a higher-level system fallback gating layer for the closed-loop governance of the AI agent. The functions of each component of the intelligent storage location pre-allocation system based on the four-layer closed-loop governance architecture are described in detail below: Feedforward guidance layer: used to encode the physical world constraints of warehouse location allocation into a structured rule set, and inject the corresponding rules of the structured rule set into the decision context of the AI agent before each location pre-allocation decision. The structured rule set is divided into a rigid constraint rule subset and a flexible constraint rule subset. AI agent: Under the guidance of the decision context, it outputs a pre-allocation plan for storage locations and a unified confidence level for the selected optimal storage location; the pre-allocation plan for storage locations output by the AI agent is sent to the upper system for execution after being checked by the upper system's rollback gating layer. Feedback sensing layer: used to detect the actual execution deviation of the storage location pre-allocation plan output by the AI agent and calculate the uniform deviation range; Self-calibration layer: used to compare the uniform deviation amplitude with the first threshold and the second threshold, automatically select one of the three calibration modes from parameter fine-tuning, strategy rollback and full replanning according to the dual threshold three-mode mechanism, and update the strategy version used by the AI agent in the next execution of the cargo location pre-allocation decision accordingly. The upper-level system rollback gating layer is used to automatically switch the storage location allocation control process from the AI agent decision-making mode to the upper-level system default allocation mode when the unified confidence level is lower than the confidence level threshold, and the storage location allocation is executed by the deterministic allocation algorithm built into the upper-level system. The end-to-end observable layer is used to record the entire process data of the feedforward guidance layer, feedback sensing layer, self-calibration layer, upper system backoff gating layer and AI agent in a cross-sectional manner. The end-to-end observable layer does not participate in decision-making and control.
[0015] The design of the four-layer closed-loop governance architecture follows these core principles: the governance architecture is not the intelligent agent itself, but rather the complete infrastructure governing the operation of AI intelligent agents. Unlike AI governance methods that address the logical constraints of software systems, the four-layer closed-loop governance architecture of this invention needs to handle physical world constraints and automatically adjust constraint strategies according to three racking configurations: the topological occlusion of double-deep racking locations is not a logical access restriction, but a spatial physical barrier caused by the racking geometry; the stacker crane's running time is not a software delay that can be changed through parameter tuning, but a physical process constrained by Newton's laws of motion; the differentiated allocation strategy for single-deep and double-deep racking locations is not a software function configuration option, but an inevitable requirement determined by the racking physical structure.
[0016] Figure 2 This is a flowchart illustrating an intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture, as provided in one embodiment of the present invention. Figure 2 As shown, the intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture includes steps S10 to S60: S10, Feedforward Constraint Injection: The physical world constraints of warehouse location allocation are encoded into a structured rule set through the feedforward guidance layer, and the corresponding rules of the structured rule set are injected into the decision context of the AI agent before each location pre-allocation decision.
[0017] This step is performed by the feedforward bootstrapping layer. Its core is to construct a structured set of rules: .
[0018] Each rule Represented as a quintuple: .
[0019] The fields in the quintuple are defined as follows: id (rule ID): A unique identifier for a rule within a rule set. For example, rules R1, R2, R4a, R4c, etc.
[0020] type (rule type): Used to distinguish the rigidity of rules. It is explicitly divided into two types: rigid constraint rules (type=RIGID) and flexible constraint rules (type=SOFT).
[0021] Predicate: This is a logical expression used to determine whether a rule is triggered (i.e., whether its condition is true). During decision-making, the rule interpreter in the feedforward layer substitutes "product attributes + candidate location attributes + current warehouse status" into this logical expression for evaluation, returning TRUE (condition true / constraint violated) or FALSE (condition false). For example, the predicate of rule R4a below is "For double-deep locations, when the outer location is occupied, inbound allocation to the inner location in the same column is prohibited." If this condition is substituted into a candidate inner location, and there is indeed goods on the outer location in the same column, then the predicate evaluates to TRUE.
[0022] `action` (triggering action): Defines the specific action the system should perform when the rule is triggered (predicate = TRUE). Its semantics vary depending on the `type` field. (1) For rigid constraint rules (type=RIGID), the main actions are: REJECT: Immediately remove the candidate storage location from the "feasible storage location candidate set" and prevent it from proceeding to the subsequent comprehensive evaluation stage. Applicable to hard constraints that are "unavoidable," such as stacker crane malfunctions, hazardous materials isolation, and batch first-in-first-out (FIFO) order.
[0023] TRIGGER_RELOCATION: This action is specifically for outbound and relocation processes in double-deep storage locations (corresponding to rule R4b). When violated, it is not directly rejected, but the system triggers a complete relocation operation process and searches for a temporary storage location for the obscured outer goods. The candidate storage location will be subject to a "relocation time cost" penalty and then participate in the overall scoring.
[0024] (2) For the flexible constraint rule (type=SOFT), the action is: DEDUCT(x): This means that when the preference (predicate=TRUE) is violated, the candidate location is not removed, but x points (base deduction value) are deducted from the candidate location's original comprehensive score. For example, DEDUCT(30) means deducting 30 points.
[0025] Weight: This field is mainly used for flexible constraint rules. It represents the relative importance or influence of the rule's deduction item in the final comprehensive score calculation. The weight value ranges from (0, 1). The actual deduction amount included in the comprehensive score is calculated as: DEDUCT deduction value × weight. For example, if rule R4c is defined as DEDUCT(30) and weight=0.95, then the actual deduction is 30 * 0.95 = 28.5 points. The larger the weight, the greater the negative impact of violating the rule on the total score. For rigid constraint rules (type = RIGID), whose action is REJECT or TRIGGER_RELOCATION, they do not participate in the score weighting. Therefore, the weight field is reserved in the data structure (it can take a default value such as NULL or 0), but it does not participate in the comprehensive score calculation.
[0026] In this embodiment, the rule set of the feedforward guidance layer is strictly divided into two categories: a rigid constraint rule subset (type=RIGID) and a flexible constraint rule subset (type=SOFT), as detailed below: (1) A subset of rigid constraint rules (violations of which would result in physical infeasibility or compliance violation), including the following rules: Rule R1 (ABC Classification and Time Zone Constraints): ABC classification is based on outbound frequency statistics over the past 90 days: the top 20% by outbound frequency is category A, 20%-50% is category B, and over 50% is category C. The predicate is: "Using the stacker crane's inbound / outbound platform as the origin, calculate the estimated arrival time of each storage location based on the stacker crane's kinematic model, and divide all storage locations into three time zones in ascending order of estimated arrival time; the number of storage locations in each of the three time zones matches the expected proportion of the corresponding product category, i.e., the proportion of storage locations in zone T1 (fast zone) equals the proportion of category A products (20%), the proportion of storage locations in zone T2 (medium-speed zone) equals the proportion of category B products (30%), and the proportion of storage locations in zone T3 (slow zone) equals the proportion of category C products (50%); category A products are preferentially allocated to zone T1, category B products to zone T2, and category C products to zone T3." 'e' represents RIGID, and 'action' represents REJECT. When the target time zone's storage capacity is insufficient, overflow allocation is allowed: Category C goods can overflow to T2 and T1 zones sequentially, and Category B goods can overflow to T1 zone. However, no category of goods can overflow to a slower zone than its target time zone to ensure the high-frequency operation efficiency of Category A goods. Overflow events are recorded in the end-to-end observable layer for subsequent rule self-evolution mechanism analysis. This time zone division comprehensively considers the combined impact of both the layer (vertical direction) and column (horizontal direction) of the storage location on the stacker crane's operation time, avoiding the shortcomings of traditional layer-only allocation methods that neglect column distance and allocate Category A goods to lower layers but farthest columns.
[0027] Rule R2 (Stacker crane accessibility constraint): The predicate is "the stacker crane in the aisle where the target storage location is located is currently available", the type is RIGID, and the action is REJECT. When a stacker crane is in a faulty or maintenance state, all storage locations in its aisle are marked as unreachable.
[0028] Rule R3 (Hazardous Materials Segregation Constraint): The predicate is "the distance between hazardous materials storage locations and general goods storage locations shall not be less than 2 storage locations", the type is RIGID, and the action is REJECT.
[0029] Rule R4a (Physical Accessibility Constraint for Double-Deep Racks, applicable only to configurations with double depth): The predicate is "For double-deep racks, when the outer rack is occupied, inbound allocation to the inner rack in the same row is prohibited," the type is RIGID, and the action is REJECT. This constraint reflects the most fundamental physical limitation of double-deep racks—the stacker crane's fork extension arms must pass through the outer position to reach the inner position; therefore, the inner rack can only be inbound when the outer rack is empty. This rule is automatically disabled in pure single-deep configurations.
[0030] Rule R4b (Double-deep storage location outbound transfer process constraint, applicable only to configurations with double depth): predicate is "When goods inside a double-deep storage location need to be outbound, if the outer storage location is occupied, the system must first execute the transfer operation process for the outer goods", type is RIGID, and action is TRIGGER_RELOCATION. This rule encodes the complete warehouse transfer operation process: First, the feedforward guidance layer performs an independent pre-allocation decision for the obscured outer pallet, searching for a temporary or permanent target pallet that meets the constraints within the same aisle. Second, the upper-level system's task scheduling module inserts the transfer task into the operation queue, setting its priority higher than ordinary inbound tasks but lower than emergency outbound tasks. Third, the upper-level system controls the stacker crane to sequentially execute a multi-step continuous operation: "remove the outer pallet → place it in the temporary pallet → return to retrieve the inner target pallet → deliver it to the outbound platform." The entire transfer process transforms the stacker crane's single-command operation into a multi-command composite operation, with the estimated operation time calculated by summing the kinematic times of each step. Fourth, after the transfer is completed, the feedback sensing layer records the transfer event and updates the inventory status. In the allocation decision scoring, candidate schemes that need to trigger a transfer operation will be subject to a transfer time cost penalty, which is equal to the total estimated operation time of the multi-step transfer multiplied by the transfer cost coefficient. This rule is automatically disabled in pure single-depth configurations.
[0031] Rule R9 (First-In, First-Out (FIFO) Constraint): The predicate is "Goods of the same SKU are sorted by the timestamp of the receiving batch. When shipping, the earliest receiving batch is prioritized for the storage location; during inventory allocation, a new receiving batch of the same SKU cannot be assigned to a storage location closer to the shipping outlet than an existing earlier batch of SKU". The type is RIGID, and the action is REJECT. This constraint ensures that the warehouse follows the FIFO principle. For goods with shelf-life requirements (such as food and pharmaceuticals), this rule is extended to the FEFO (First-to-Expires, First-Out) model, where the expiration date of the product batch replaces the receiving timestamp as the sorting basis, prioritizing the shipment of the batch with the closest expiration date.
[0032] (2) A subset of flexible constraint rules (which can be weighed through a scoring and penalty mechanism) includes the following rules: Rule R4c (SKU compatibility preference rule for inner and outer sides of double-deep storage locations, applicable only to configurations with double depth): predicate is "For double-deep storage locations, when the inner position already stores product A, the outer position in the same column will preferentially store products with the same SKU as A", type is SOFT, action is DEDUCT(30), weight is 0.95. This rule is explicitly set as a flexible constraint rather than a rigid constraint because: in scenarios with many SKU types but small inventory of each type (such as e-commerce warehouses), if it is mandatory that the inner and outer positions must be the same SKU, it will result in a large number of vacant outer positions of double-deep storage locations, seriously reducing space utilization and violating the original intention of using double-deep racks to increase storage density. Through a scoring and penalty mechanism, the system achieves a balance between space utilization and outbound efficiency. This rule is automatically disabled in pure single-deep configurations.
[0033] Rule R5 (Adjacent Optimization for Similar Goods): The predicate is "There is no occupied storage location with the same SKU as the goods to be received in the adjacent storage location set of the candidate storage location", the type is SOFT, the action is DEDUCT(15), and the weight is 0.8.
[0034] Rule R6 (Outbound Frequency Matching): The predicate is "When the goods to be received are high-frequency outbound goods (the outbound frequency normalization ranking is above the preset high-frequency threshold), the candidate storage location is a difficult-to-access storage location such as the double-deep inner storage location", the type is SOFT, the action is DEDUCT(20), the weight is 0.9, and high-frequency outbound goods are preferentially allocated to single-deep storage locations or double-deep outer storage locations.
[0035] Rule R7 (Warehouse Area Load Balancing): predicate is "the current capacity utilization rate of the warehouse area where the candidate storage location is located is higher than the average capacity utilization rate of all warehouse areas in the whole warehouse", type is SOFT, action is DEDUCT(10), and weight is 0.7.
[0036] Rule R8 (Stacker crane dual-command cycle optimization): The predicate is "When allocating inbound tasks, prioritize locations in the same aisle as the current outbound task and where the stacker crane can complete both inbound and outbound operations within one dual-command cycle", the type is SOFT, the action is DEDUCT(25), and the weight is 0.85. This rule utilizes the stacker crane dual-command cycle to reduce idle travel. The path planning for the dual-command cycle needs to consider four key nodes simultaneously: the inbound platform pickup location P1, the inbound target location P2, the outbound source location P3, and the outbound platform placement location P4. The system calculates the total estimated operation time for the running path P1→P2→P3→P4 based on the stacker crane kinematic model and selects the inbound target location that minimizes the total time of the dual-command cycle.
[0037] The criteria for determining the action deduction and weight values for each flexible rule are explained below: The base deduction (the value in DEDUCT) reflects the "single penalty intensity" for violating the rule; the more the rule emphasizes physical / business criticality, the larger the deduction. The weight reflects the rule's "relative influence" in the overall comprehensive score; the larger the weight, the more difficult it is for other scoring factors to compensate for the violation. The actual deduction in the comprehensive score = DEDUCT deduction value × weight. The specific criteria for these values are as follows: (A) Rule R4c: Directly related to the future transfer cost of double-deep racking, with the highest deduction score and weight, take action=DEDUCT(30), weight =0.95; (B) Rule R8: Directly saves the stacker crane's idle travel, with a higher deduction score and weight. Take action = DEDUCT (25) and weight = 0.85; (C) Rule R6: Affects the average outbound time of the entire warehouse, with a high weight and a medium deduction. Take action=DEDUCT(20) and weight =0.9; (D) Rule R5: biased towards operational convenience and picking efficiency, with moderate deductions and weights, take action=DEDUCT(15), weight =0.8; (E) Rule R7: Long-term balance indicator, with the least impact of a single violation, and the lowest deduction and weight. Take action=DEDUCT(10) and weight =0.7.
[0038] All weight values are in the range (0, 1); the initial value is configured by the warehouse administrator during the system initialization phase according to business priorities, and is subsequently dynamically fine-tuned by the self-calibration layer based on the deviation report from the feedback sensing layer.
[0039] S10 further includes: S11: The cargo structure perception module of the feedforward guidance layer reads the cargo structure identifier from the warehouse configuration parameters during system initialization; S12: Before each pre-allocation decision of the storage location, the feedforward guidance layer activates the corresponding rule in the structured rule set according to the current storage location structure and injects the decision context of the AI agent.
[0040] Specifically, when the storage structure is a pure single-depth structure, the physical accessibility constraint for double-depth storage locations (rule R4a), the outbound transfer process constraint for double-depth storage locations (rule R4b), and the SKU compatibility preference rule for the inner and outer sides of double-depth storage locations (rule R4c) are automatically disabled. The scoring factors related to the transfer risk of double-depth storage locations are removed from the scoring function and their weights are proportionally redistributed among the remaining scoring factors. All other rules in the structured rule set are activated. When the storage structure is a pure double-depth structure, all rules in the structured rule set are activated, and the scoring factors related to the transfer risk of double-depth storage locations are enabled. When the storage structure is a mixed single-depth and double-depth structure, the double-depth related constraint rules (rules R4a, R4b, and R4c) and the scoring factors related to the transfer risk of double-depth storage locations are determined for each storage location based on the actual shelving type of the aisle where the storage location is located.
[0041] In some embodiments, the structured rule set of the feedforward guidance layer supports knowledge transfer between multiple warehouses: the validated structured rule set in the running warehouse is exported as a standardized rule template, loaded when the new warehouse is cold-started, and the activation status and numerical constraint parameters of the rules are automatically adjusted according to the warehouse's structure and physical parameters.
[0042] The migration process of the structured rule set specifically includes the following steps: Step 1: Export the structured rule set of the running warehouse by decoupling the rule structure and rule parameters. The rule structure includes rule ID, type (RIGID / SOFT), predicate logical expression, action type and initial weight value. The rule parameters include numerical values related to warehouse geometric parameters (such as time zone thresholds such as ABC, warehouse capacity ratio target, SKU compatibility DEDUCT deduction value, and speed / acceleration related to stacker crane kinematics). The exported file format is a standardized rule template file. Step 2: During the cold start of the new warehouse, the rule template file is loaded and the warehouse structure perception module is used to read the warehouse structure identifier (pure single-depth / pure double-depth / mixed), physical dimensions (number of layers, number of columns, number of aisles, double-depth ratio) and stacker crane kinematic parameters. Step 3: Automatically determine the enabling / disabling of the scoring factors related to the risk of double-deep cargo relocation based on the configuration identifier of the new warehouse (rules R4a, R4b and R4c). Step 4: The numerical constraint parameters are scaled proportionally to the measured kinematic parameters and storage capacity of the new warehouse. The time boundaries of time zones ABC are recalculated by the kinematic model of the stacker crane in the new warehouse. The utilization rate of the warehouse area load balancing target is automatically normalized according to the number of aisles in the new warehouse. The cost coefficient of double-deep warehouse relocation retains the empirical value of the original warehouse but is allowed to be automatically adjusted by the self-correction layer according to feedback after the first round of operation. Step 5: The migrated structured rule set is launched in the new warehouse in a trial operation state. During the first N statistical periods (default N=7 days), the confidence threshold of the upper system rollback gating layer is automatically relaxed (the first confidence threshold and the second confidence threshold are temporarily increased by 0.1 each) to allow the AI agent to obtain sufficient feedback samples in the new environment. Subsequently, it is automatically tightened to the normal threshold according to the feedback sensing layer deviation report.
[0043] S20. Intelligent Decision-Making and Confidence Assessment: Guided by the decision-making context, the AI agent outputs a pre-allocation plan for cargo locations and a unified confidence level for the selected optimal cargo location.
[0044] This step is performed by the AI agent. The AI agent receives the attributes of the goods to be allocated, the list of candidate storage locations (including the structure identifier and the occupancy status of adjacent storage locations), and the rules injected by the feedforward guidance layer. Based on this, it makes a storage location pre-allocation decision and outputs the storage location pre-allocation scheme and the unified confidence level of the selected optimal storage location.
[0045] The AI agent uses a standardized interface; the input protocol defines the data format received by the AI agent, including a list of goods attributes to be allocated, a list of candidate storage location attributes (including storage location depth type identifiers and adjacent storage location occupancy status), and a set of constraint rules injected by the feedforward guidance layer; the output protocol defines the data format returned by the AI agent, including a list of mapping relationships between goods and storage locations, a unified confidence score for each mapping relationship, and an estimated stacker crane running time; the standardized interface supports replacing the AI agent with different types of decision models without modifying the other layers of the closed-loop governance architecture.
[0046] In this invention, the AI agent can perform pre-allocation decision-making for cargo locations in two specific ways: one is to use a multi-factor weighted scoring model to perform pre-allocation decision-making for cargo locations; the other is to use a reasoning decision-making model based on a large language model to perform pre-allocation decision-making for cargo locations.
[0047] (1) When the AI agent uses a multi-factor weighted scoring model to make a pre-allocation decision for cargo locations, it specifically includes steps S21 to S24: S21: For each candidate vacant storage location in the automated warehouse, perform feasibility filtering based on the rigid constraint rules injected by the feedforward guidance layer, exclude storage locations that do not meet the rigid constraints, and generate a set of feasible storage location candidates.
[0048] Specifically, each candidate vacant storage location in the automated warehouse is traversed, and the feasibility of each candidate vacant storage location is filtered according to the rigid constraint rules injected by the feedforward guidance layer. In particular, for double-deep storage locations, the occupancy status of the inner and outer sides of the same column is checked to exclude storage locations that do not meet the rigid constraints, thereby generating a set of feasible storage location candidates.
[0049] S22: Calculate the comprehensive score for each candidate location in the feasible location candidate set.
[0050] The comprehensive score for each candidate location in the feasible candidate location set is given by the following function: ; in Let i be the weight of the i-th rating factor. , Let be the calculation function for the i-th rating factor.
[0051] Total number of scoring factors in the above comprehensive scoring function Determined by the cargo structure: Pure single-deep structure (only Activated), including double-deep configurations (pure double-deep or hybrid) ( (All activated). The following provides a detailed explanation of each rating factor: (Stacker crane running time factor): Based on the stacker crane's kinematic model, the estimated running time from the current position to the candidate storage location is calculated. The normalized reciprocal of the time is taken as the scoring factor, meaning the shorter the running time, the higher the score. ; in, For stacker crane to candidate storage location The estimated run time, This represents the maximum estimated travel time from the stacker crane to the candidate storage location. Initial weights. .
[0052] (Outbound Frequency Matching Factor): Evaluates the degree of matching between the frequency of goods outbound and the location depth of the storage location. Frequently outbound goods are preferentially assigned to single-depth or double-depth outer storage locations to avoid frequent relocations caused by assigning them to double-depth inner storage locations. Initial Weight . Specific calculation formula: ; in, For goods The normalized ranking of outbound frequencies is obtained by normalizing the outbound frequencies of the past 90 days in descending order, with the highest frequency being [the highest frequency]. When the frequency is lowest ; Candidate storage location The normalized ranking of retrieval and placement difficulty shows that single-deep and double-deep outer storage locations have a normalized value of 0 (easiest to retrieve / place), while double-deep inner storage locations have a normalized value of 1 (most difficult to retrieve / place). and When consistent (e.g., high-frequency goods are assigned to easily accessible depths, and low-frequency goods are assigned to difficult-to-access depths) When the two are completely opposite (e.g., high-frequency goods are allocated to the inner side of the double depth), In other cases, the decay rate is linear. For mixed-configuration reservoirs, Each reservoir is independently normalized within its own reservoir area to avoid incomparability across different reservoir areas.
[0053] (Reservoir Area Balanced Load Factor): Evaluates the degree of balance in the utilization rate of reservoir capacity in each roadway after allocation. ; in, To select candidate storage locations The standard deviation of the storage capacity utilization rate of each lane is calculated after incorporating the pre-allocation scheme for storage locations; The theoretical maximum standard deviation is given when the storage capacity utilization rate between roadways is extremely unbalanced (one roadway is 100% occupied, and the rest are 0%). In the warehouse of the alleyway,
[0054] When the candidate solutions make the storage capacity utilization rate between roadways tend to be consistent The initial weights should be close to 1, otherwise close to 0. .
[0055] (Proximity Factor for Similar Goods): Assess the quantity of the same SKU in adjacent storage locations. ; Candidate storage location The adjacent storage locations already contain goods awaiting warehousing. The number of storage locations for the same SKU; Candidate storage location The total number of adjacent storage locations (including occupied and unoccupied locations, determined by the warehouse geometry); this embodiment uses candidate storage locations. The 3x3 storage window centered on the center (excluding itself) has a total of 8 adjacent storage spaces as candidate storage spaces. The set of adjacent storage locations, candidate storage locations When located in a corner position, the actual number of adjacent storage positions will be less than 8; A larger value indicates that the candidate storage location has more of the same SKU, which is beneficial for subsequent picking of adjacent SKUs and reduces the idle travel of the stacker crane; when (In the theoretical extreme case) Force the assignment to 0 to avoid division by zero errors. Initial weights. .
[0056] (Double-deep storage location relocation cost factor, activated only in configurations containing double-deep locations): Comprehensive assessment of future relocation risks and costs for double-deep storage locations: ; in, This represents the expected cost of transferring inventory.
[0057] This includes two sub-dimensions: (a) Future relocation risk penalty for outer candidate storage locations: When a candidate storage location is a double-deep outer storage location and the inner storage location in the same column already stores goods with a different SKU than the one to be received, the outer pallet of the current receipt will be moved out before the inner goods are retrieved in the future. Therefore, the candidate storage location incurs an additional expected relocation cost penalty. The penalty is zero when the inner storage location is empty or the inner and outer SKUs are the same (Note: Rule R4a prohibits the retrieval of goods into the inner storage location of an occupied double-deep outer storage location. Therefore, this sub-dimension only assesses the forward-looking risk of outer candidate storage locations and does not overlap with Rule R4a); (b) Additional time cost of relocation operations: The total time of multi-step continuous relocation operations is estimated based on the stacker crane kinematics model and normalized after multiplying by the relocation cost coefficient (default value 1.5). The two sub-dimensions are combined into a normalized expected relocation cost. ,final Initial weights In a pure single-depth configuration, this scoring factor is removed, and its weight is proportionally redistributed among the remaining scoring factors.
[0058] The normalized benefit model is adopted, where higher costs result in lower scores. The value range is [0, 1]. The higher the expected cost of the transfer, the better. The closer it is to 0, the less risk there is of transferring inventory. The value is 1, therefore, and When calculating the overall score, the directions are consistent, and the weighted sum can be directly calculated.
[0059] In addition, the flexible constraint rule R4c (SKU compatibility preference rule for inner and outer sides of double-deep storage locations) uses an independent penalty item. Form is incorporated into the overall score: When the candidate storage location is double-deep inner or outer, and the SKU already stored on the opposite side of the same column is different from the SKU to be stored, the score is deducted according to the rule defined DEDUCT(30) and weight 0.95; this penalty item is automatically disabled in pure single-deep configuration.
[0060] S23: Sort candidate storage locations in descending order of overall score, and select the best storage location with the highest overall score. As a result of the allocation, a pre-allocation scheme for storage locations is generated, and the uniform confidence level of the optimal storage location selected by the AI agent is calculated.
[0061] The uniform confidence level of the selected optimal storage location is calculated using a two-dimensional formula that combines absolute score level and relative score difference: , in, Optimal storage location selected for the AI agent Overall score; The historical average comprehensive score of all optimal storage locations within the most recent preset period (e.g., the most recent 100 times); This is an absolute scoring water level component used to reflect... Compared to The relative intensity; This is a relative score difference component. In a multi-candidate scenario (when there are two or more candidates), select... , The overall score for the second-best cargo location. To preset a minimum positive value, In the case of only a single candidate scenario, a preset neutral value of 0.5 is used; The preset range of values for the fusion weight coefficients is as follows: This is used to balance the relative importance of the absolute score level component and the relative score difference component; in this embodiment, we take... This means that the absolute score level component and the relative score difference component are given the same weight in the unified confidence calculation, so as to avoid any one-sided bias in the backtracking decision by any dimension. Indicates will Truncate to the [0, 1] interval and use a uniform confidence level. The value of is constant in the interval [0, 1], serving as the sole confidence basis for the rollback determination by the upper-level system rollback gating layer.
[0062] The calculation results are all passed The function is truncated to the interval [0, 1] to ensure that the scoring... or When negative values occur, the uniform confidence level remains constant at [0, 1]; when Or when the feasible candidate location set is empty, Forced zeroing automatically triggers the upper-level system to roll back the gating layer.
[0063] S24: Output the pre-allocation scheme for storage locations and the uniform confidence level of the optimal storage location.
[0064] The pre-allocation scheme for cargo locations generated by the AI agent and the unified confidence level of the optimal cargo location selected by the AI agent are compiled into a structured JSON format and then output.
[0065] (2) When the AI agent uses a reasoning and decision-making model based on a large language model to execute the pre-allocation decision of the cargo location, the specific steps include the following: The constraint rules, candidate storage location attributes, and product attributes injected by the feedforward guidance layer are input into the large language model in the form of structured prompts. These prompts include the system role definition (storage location allocation expert for automated warehouses), a structured description of the constraint rule set, a list of candidate storage locations (including depth type, occupancy status, and adjacent storage location information), and product attribute information. The large language model then outputs a storage location pre-allocation scheme based on its reasoning capabilities. The output format of the large language model is conventionally defined as a structured JSON format containing storage location numbers, uniform confidence scores, and decision rationale, for subsequent layers to parse and process.
[0066] Specifically, the large language model is responsible for generating the pre-allocation scheme of cargo locations (i.e., providing a preferred ranking and optimal recommendation of candidate cargo locations based on reasoning). The calculation method of the comprehensive score of cargo locations is consistent with the implementation method based on the multi-factor weighted scoring model: after the large language model outputs the pre-allocation scheme of cargo locations, the AI agent still calculates the comprehensive score of cargo locations according to the comprehensive score function defined in step S22. The comprehensive scores of the optimal and second-best cargo locations recommended by the large language model are calculated separately and substituted into the unified confidence formula in step S23 to calculate the unified confidence score. Thus, the unified confidence scores output by the two implementation methods are from the same source and have the same caliber, and can be directly compared. Both can be used as the basis for the rollback judgment of the rollback gating layer of the upper system. The "decision reason" field output by the large language model is only used for the audit record and interpretability display of the end-to-end observable layer and does not participate in the numerical calculation of the unified confidence score.
[0067] The advantage of using a large language model to infer and execute the pre-allocation decision of the storage location is that it can understand and process unstructured business constraints described by natural language. The disadvantage is that the latency of a single inference is relatively high (about 3-5 seconds). It is suitable for scenarios where the decision-making time requirement is not high but the business rules are complex and changeable.
[0068] To address the throughput demands during peak arrival periods, the large language model inference implementation employs a batch inference mode: multiple product attributes and candidate storage location information from the same batch of arrivals are merged into a single structured prompt input. The large language model then outputs pre-allocation schemes for multiple products at once, thereby reducing the average decision time per product in batch scenarios to less than 1 second. When the single response time of the large language model exceeds a threshold (default 10 seconds), the upper-level system automatically switches the batch to allocation based on a multi-factor weighted scoring model, ensuring that system throughput is not affected.
[0069] The large language model can be a locally deployed open-source model series such as DeepSeek, Qwen, or GLM, or it can be a model service provided by a large model vendor through an API service request, or it can be a self-trained dedicated large language model, etc., and is not limited to a specific large model itself.
[0070] S30, Execution Deviation Feedback: The actual execution deviation of the cargo location pre-allocation plan output by the AI agent is detected through the feedback sensing layer, and the uniform deviation magnitude is calculated.
[0071] This step is executed by the feedback sensing layer. The feedback sensing layer connects to the data interface of the upper-level system and periodically (e.g., after every 100 operations) collects actual operational data from the upper-level system to calculate the actual execution deviation of the storage location pre-allocation plan output by the AI agent. Deviation indicators include the stacker crane running time deviation rate. Double-deep cargo displacement deviation rate Storage capacity utilization deviation Batch First-In-First-Out Deviation Rate The weighted fusion of each deviation index yields a unified deviation range.
[0072] (a) Stacker crane running time deviation rate : The actual running time of the stacker crane for each inbound or outbound operation is recorded by sensor G1 (stack crane running time sensor) in the upper-level system. The estimated running time calculated by the AI agent based on the kinematic model of the stacker crane in the pre-allocation scheme of the storage location is also obtained. The stacker crane running time deviation rate is calculated in this way. .
[0073] When the AI agent calculates the estimated running time of the stacker crane, the stacker crane kinematic model automatically selects a speed curve model based on the travel distance: (1) When the journey is long enough (journey distance) When using a trapezoidal velocity curve model, the unidirectional running time is: .
[0074] Trapezoidal velocity curve model: This means the stacker crane has enough distance to accelerate first (the stacker crane's acceleration is...). (to maximum speed) When a vehicle travels at a constant speed for a period and then decelerates to zero, its velocity curve exhibits a trapezoidal shape with three stages: acceleration, constant speed, and deceleration. This is called the trapezoidal velocity curve model. Its unidirectional travel time is the sum of these three time segments. Acceleration phase: Accelerating from 0 to... Time taken: ; Distance traveled during acceleration phase: ; Deceleration phase: symmetrical to acceleration phase, duration: ; Distance traveled during the deceleration phase: ; Uniform speed segment: Remaining distance traveled, time taken: ; The total time is obtained by adding the three time intervals together: .
[0075] This trapezoidal velocity curve model depicts the actual running time of the stacker crane under long-stroke conditions, constrained by both acceleration / deceleration and maximum speed. It is a computational tool for AI agents. (Stacker crane running time factor) and the physical basis for estimated running time.
[0076] (2) When the journey is short (journey distance) When the stacker crane cannot accelerate to its maximum speed and needs to decelerate, using a triangular velocity curve model, its one-way travel time is: ; Note: The stacker crane cannot accelerate to its maximum speed in time. We must start slowing down in order to When the vehicle comes to a stop, its velocity curve exhibits a triangular shape with two phases: acceleration and deceleration (without a constant velocity phase), which is called the triangular velocity curve model. The derivation of its unidirectional travel time is as follows: Let the peak speed reached be If acceleration and deceleration are symmetrical and each accounts for half of the stroke, then: ; Solve for the peak speed: ; Total running time: .
[0077] The travel distance of the stacker crane Decompose the data to obtain the horizontal travel distance. Vertical travel distance The corresponding trapezoidal velocity curve model or triangular velocity curve model is used to calculate the time taken for horizontal and vertical motion.
[0078] For the horizontal direction: ; ; in, This indicates the maximum horizontal speed of the stacker crane. This indicates the horizontal acceleration of the stacker crane. The running time of the stacker crane in the horizontal direction.
[0079] The same applies to the vertical direction: ; ; in, This indicates the maximum speed of the stacker crane in the vertical direction. This indicates the vertical acceleration of the stacker crane. The vertical travel time of the stacker crane.
[0080] Since the stacker crane performs horizontal and vertical movements in parallel, the larger of the two estimated operation times for a single stacker crane operation is taken: .
[0081] Stacker crane runtime deviation rate: ; in, This refers to the actual running time of a single operation by the stacker crane. Estimate the running time for a single operation of the stacker crane.
[0082] (ii) Deviation rate of double-deep cargo transfer : Using sensor G2 (a double-deep cargo relocation sensor, activated only in configurations with double depth), the system counts the number of relocation operations triggered by outer goods obstructing inner target goods within each statistical period, as well as the actual time consumed for each relocation operation within the statistical period (including the total time for location finding, queuing, and multi-step execution by the stacker crane). It also obtains the number of relocation operations and the total relocation time within the same statistical period estimated by the AI agent based on the pre-allocation scheme, and calculates the double-deep cargo relocation deviation rate accordingly. .
[0083] The deviation rate of double-deep cargo transfer It is characterized by two deviation indices of the same dimensions but different dimensions, which measure the deviation in the dimensions of transfer frequency and transfer time, respectively: A. Transfer frequency deviation rate: ; B. Deviation rate of transfer time: ; in This refers to the actual number of data transfer operations triggered within the statistical period. The number of warehouse transfer operations estimated by the AI agent within the same statistical period based on the warehouse location pre-allocation scheme; This represents the sum of the actual time spent on all warehouse transfer operations within the statistical period. The total time for warehouse transfer operations estimated by the AI agent within the same statistical period based on the pre-allocation plan for warehouse locations. To prevent division by zero by setting a minimum positive value.
[0084] Inside the feedback sensing layer and First, each component is normalized to the [0, 1] interval using max-min normalization, and then weighted and fused to obtain the double-deep cargo displacement deviation rate. : ; in: The internal fusion weight between the transfer frequency dimension and the transfer time dimension is configured by the warehouse administrator based on business priorities (e.g., if prioritizing the number of transfer operations, then...). Prioritize the time spent on warehouse transfer operations. (< 0.5), or the self-calibrating layer can adaptively adjust based on the correlation analysis of historical samples. In this embodiment, the default value is taken as 0.5. (i.e., frequency and time dimensions are equally weighted); norm(·) indicates mapping to the [0, 1] interval by max-min normalization. After the above fusion, Deviation rate with stacker crane running time Storage capacity utilization deviation Batch First-In-First-Out Deviation Rate The dimensions are consistent and can be directly substituted into the following text. The weighted fusion formula.
[0085] The self-calibrating layer adaptively adjusts based on correlation analysis of historical samples. The specific process is as follows: Within the most recent sliding window, the Pearson correlation coefficients between the deviation rate of the transfer frequency dimension, the deviation rate of the transfer time dimension, and the actual total transfer cost of that period are calculated. and ,Pick That is, deviation dimensions that are more correlated with actual inventory transfer costs automatically receive higher fusion weights, thereby enabling... Dynamically adapt to the actual bottlenecks in the current warehouse. To prevent division by zero, a preset minimum positive value is used.
[0086] The above is used to adaptively determine the fusion weights. correlation coefficient and This is not a value derived from general experience, but a statistical quantity calculated online based on the relevant parameters of this patent's database transfer. Its specific meaning and calculation method are as follows.
[0087] Set the most recent sliding window The content contains The statistical period, for the first one Statistical periods ( ), and its transfer frequency dimension deviation rate is denoted as (abbreviated as) The deviation rate of the transfer time dimension is ), (abbreviated as) The actual total cost of the transfer was .but Defined as a sequence and The Pearson correlation coefficient between them Defined as a sequence and The Pearson correlation coefficients between the two variables measure the strength of the linear correlation between the frequency dimension bias, the time dimension bias, and the actual total cost of inventory transfer, respectively. The calculation formula is as follows: ; .
[0088] In the formula, , , Sliding windows Inside , , Sample means of the three sequences: ; and , The range of values is The closer the absolute value is to 1, the stronger the explanatory power of the deviation in that dimension on the fluctuation of the actual total cost of inventory transfer; , Substitute into the previous equation This allows the deviation dimension, which is more correlated with the actual transfer cost, to be included in the fusion weight. It automatically obtains a higher percentage.
[0089] Among them, the actual total cost of the transfer The transfer-related parameters defined in this patent characterize the first... The comprehensive cost incurred by all warehouse transfer operations triggered within a statistical period due to goods on the outer side of a double-deep storage location obstructing the target goods on the inner side is calculated by weighting the actual number of warehouse transfer operations and the actual man-hours occupied within that period, as shown in the following formula: .
[0090] In the formula, For the first The number of actual warehouse transfer operations collected by sensor G2 within a statistical period The total actual time spent on warehouse transfer operations collected by sensor G2 during the statistical period (including the total time for location finding, queuing, and multi-step execution by the stacker crane) is consistent with the original observations used in the warehouse transfer frequency dimension deviation rate and warehouse transfer time dimension deviation rate mentioned above. This is the fixed cost coefficient for a single warehouse transfer operation, which describes the fixed overheads introduced for each warehouse transfer, such as equipment start-up, shutdown, scheduling, and location finding. The cost coefficient per unit of warehouse transfer time describes the time cost of warehouse transfer operations using stacker cranes and aisle resources. , The energy consumption and labor cost structure of the warehouse are calibrated by the warehouse manager, and can be normalized under default conditions. .
[0091] Since the actual total cost of warehouse relocation is driven by both the number of relocations and the relocation man-hours, and the degree to which these two factors dominate the total cost varies in different warehouses and different operational stages, this patent addresses this issue by... , The statistical correlation between the frequency dimension deviation and the time dimension deviation of inventory transfer and the actual total cost of inventory transfer are measured respectively, and the fusion weights are adaptively assigned accordingly. This reduces the inventory transfer deviation rate. The internal integration always aligns with the actual cost structure of warehouse relocation, and can more objectively reflect the actual relocation bottlenecks compared to fixed weights.
[0092] (iii) Deviation in warehouse capacity utilization : The actual storage capacity utilization rate of single-deep and double-deep storage locations is statistically analyzed using sensor G3 (storage capacity utilization sensor) in the upper-level system. The target storage capacity utilization rate output by the AI agent when executing storage location pre-allocation decisions is also obtained, and the storage capacity utilization rate deviation is calculated accordingly. .
[0093] Actual utilization rate of warehouse capacity Defined as the ratio of the number of occupied storage spaces in storage area z to the total number of storage spaces, where This represents the number of currently occupied storage spaces within storage area z. The total number of storage locations within storage area z: .
[0094] Storage capacity target utilization rate The AI agent, when making pre-allocation decisions for storage locations, considers the warehouse area load balancing factor. Synchronous output; calculate the absolute deviation between the actual storage capacity utilization rate and the target storage capacity utilization rate for each storage area, and obtain the storage capacity utilization rate deviation for the corresponding storage area. : .
[0095] Storage capacity utilization deviation reported to the self-correction layer The deviation is weighted and averaged based on the number of warehouse units in each area, ensuring that the deviation weight of the larger warehouse area is higher than that of the smaller warehouse area: .
[0096] Under the mixed configuration, the deviation of single-depth storage capacity utilization rate was calculated independently for single-depth and double-depth areas. Deviation from double-deep storage capacity utilization Two deviation indicators: ; ; in, This represents the actual utilization rate of single-depth storage space. The target utilization rate of single-depth storage space. To determine the actual utilization rate of double-deep storage capacity, The target utilization rate for double-deep storage capacity.
[0097] (iv) Batch First-In-First-Out Deviation Rate : The batch first-in-first-out (FIFO) deviation rate is calculated by using sensor G4 (batch FIFO sensor) in the host system to determine the ratio of abnormal outbound events that did not follow the FIFO / FEFO order to the total number of outbound operations within a preset period. .
[0098] The feedback sensing layer will calculate the four types of raw deviation indicators (stacker crane runtime deviation rate). , library transfer deviation rate Storage capacity utilization deviation Batch First-In-First-Out Deviation Rate After being normalized to the [0, 1] interval by max-min normalization, the uniform bias amplitude obtained by weighted fusion is... The calculation formula is: ; in In this embodiment, the default value is... ; This indicates that the map is normalized to the [0, 1] interval by max-min. This ensures... With the first threshold preset by the self-calibration layer Second threshold (Also within the [0, 1] interval) Dimensions are consistent and directly comparable. Weights can be customized by warehouse administrators based on business concerns, or adaptively adjusted by the self-correction layer based on correlation analysis of historical samples. For example, for pharmaceutical or food warehouses with shelf-life management requirements, the administrator can increase the weight of the batch FIFO deviation rate (e.g., from 0.25 to 0.40) and proportionally decrease the weights of the other three items to give greater weight to FEFO / FIFO compliance deviation within a unified deviation range; for warehouses sensitive to stacker crane energy consumption, the weight of the stacker crane running time deviation rate can be increased. When the self-calibrating layer adaptively adjusts the weights based on the correlation analysis of historical samples, it calculates the correlation coefficients between the four types of deviation indicators and the overall system objective (such as comprehensive operating cost or KPI achievement) within the most recent sliding window. Deviation dimensions with higher correlations automatically receive greater weights and are normalized so that the sum of the four weights is always 1. For example, when warehouse transfer problems occur frequently in a certain stage and the correlation between the warehouse transfer deviation rate and comprehensive operating cost of double deep cargo transfer increases significantly, the self-calibrating layer automatically increases the weight corresponding to the warehouse transfer deviation.
[0099] If a certain deviation is severely unbalanced (e.g., a single...) It is already far superior to other dimensions, but after integration Still below If the deviation is severely imbalanced, the fractal deviation alarm submodule of the end-to-end observable layer will issue an independent warning and hand it over to the warehouse administrator for traceability as needed, without participating in the automatic switching of the self-correction layer mode. The criterion for determining a severe imbalance is: among the four types of normalized deviation indicators, there is a fractal deviation that exceeds k times the average of the other three fractal deviations (k=3 in this embodiment), or exceeds the fractal alarm threshold preset for that dimension. At this time, even if the uniform deviation amplitude is still lower than the first threshold after weighted fusion, and should enter the parameter fine-tuning mode according to the dual-threshold three-mode mechanism, the self-correction layer will also suppress this automatic mode switching, and instead issue an independent warning by the fractal deviation alarm submodule and hand it over to the warehouse administrator for traceability as needed, so as to avoid the parameter fine-tuning for the overall deviation from masking or aggravating the structural imbalance of a single dimension.
[0100] S40, Hierarchical self-correction: The self-correction layer compares the uniform deviation amplitude with the first threshold and the second threshold, and automatically selects one of the three correction modes—parameter fine-tuning, strategy rollback, and complete replanning—according to the dual-threshold three-mode mechanism, and updates the strategy version used by the AI agent in the next execution of the pre-allocation decision of the storage location.
[0101] This step is performed by the self-calibration layer. The self-calibration layer receives the uniform deviation amplitude output from the feedback sensing layer. Then, the deviation range will be standardized. With two thresholds (first threshold) Second threshold , The system compares the three correction modes and automatically selects one to execute, thereby updating the strategy version used by the AI agent in the next execution of the pre-allocation decision, forming a dual-threshold, three-mode hierarchical correction mechanism.
[0102] Mode 1: Parameter fine-tuning mode ( The deviation is small, and only the weight parameters of the scoring factors in the cargo location comprehensive scoring function are adjusted.
[0103] Mode 2: Strategy Rollback Mode If the deviation is too large, the allocation strategy will be rolled back to the historical strategy version with the best deviation index in the strategy version stack.
[0104] Mode 3: Complete Replanning Mode The deviation is too large, and the historical strategy version in the current strategy version stack is no longer applicable. This triggers the AI agent to re-plan the location allocation based on the current warehouse status. The newly generated strategy is added to the strategy version stack as a new strategy version.
[0105] The parameter fine-tuning mode and the strategy rollback mode are strictly distinguished in terms of correction granularity: the parameter fine-tuning mode only makes small in-situ adjustments to the weight parameters of one or a few scoring factors in the cargo location comprehensive scoring function (for example, adjusting a certain weight by a preset step size based on the current strategy version). The adjusted parameter vector still belongs to the continuation of the current strategy version and does not trigger the version switching of the strategy version stack. The strategy rollback mode, on the other hand, replaces the entire strategy version. The complete strategy version includes all weights of the comprehensive scoring function, the enabled / disabled status of the corresponding rules, the threshold and preference parameters of each rule, and the prompt word template or model parameter snapshot of the AI agent. The self-correction layer selects the historical version with the best deviation index in the strategy version stack based on the deviation index generated by the feedback sensing layer and replaces it as a whole (rather than just a single weight) with the current working strategy version.
[0106] The full replanning mode is triggered when there are no available qualified strategy versions in the strategy version stack or when the deviation is extremely large. The AI agent then completely regenerates the location allocation plan based on the current warehouse state. The strategy version stack maintains snapshots of the most recent K (e.g., 10) historical strategy versions that have passed the deviation detection of the feedback sensor layer. Each time a strategy version is completely switched or the AI agent generates a new strategy through full replanning and it passes verification by the feedback sensor layer, the new version is added to the stack. When the stack is full, the oldest version is discarded according to the FIFO rule. When the strategy version stack is empty or the deviation indicators of all historical versions do not meet the qualified threshold, the strategy rollback mode is automatically upgraded to the full replanning mode. The switching of the above three correction modes, parameter adjustment details, and strategy version change records are all recorded in a cross-cutting manner by the end-to-end observable layer, forming an auditable correction link.
[0107] In one embodiment, the first threshold Second threshold Determined in the following manner: (1) When the number of samples for system initialization or uniform deviation amplitude is less than N, the first threshold and the second threshold are set using the initial values. , ; (2) When the number of samples with uniform deviation amplitude reaches N, the mean of the samples with uniform deviation amplitude within the statistical window is used. and standard deviation Based on this, the first threshold is adaptively and dynamically adjusted. Second threshold : Specifically, based on a sample sequence of uniform deviation amplitudes over the most recent N statistical periods. Calculate the sample mean of uniform deviation amplitude according to the standard formula. and standard deviation : ; ; Where N is the statistical window length (in this example, N=50 by default). This represents the uniform deviation amplitude calculated and archived by the feedback sensing layer during the i-th statistical period. At the start of each new statistical period, the earliest... Move out of window, latest Add a window and recalculate using the formula above. and Then update dynamically based on this. If dynamic calculation leads to (For example (At that time), it will automatically Set as ( To preset the minimum threshold interval, this embodiment uses 0.01 to ensure that it remains constant at any given time. ; (3) When the uniform deviation amplitude is within M consecutive statistical periods (M = 5 in this embodiment), All are below the first threshold At that time, and Each threshold is tightened by multiplying by a tightening factor less than 1. The tightening factor ranges from (0, 1), and in this embodiment, the default tightening factor is 0.9, i.e., a 10% reduction. When the number of times the full replanning mode is triggered in a single cycle exceeds the frequency threshold (default is 3 times per cycle), [the following will occur]. and Each threshold is relaxed by multiplying by a relaxation factor greater than 1. The relaxation factor ranges from (1, +∞). In this embodiment, the default relaxation factor is 1.1, which is an increase of 10%. All correction actions are output to the end-to-end observable layer for auditing and logging.
[0108] In one embodiment, when the current cumulative number of uniform deviation amplitude samples within the statistical window... At that time, the two threshold values are: ; ; in, and They are respectively , The lower limit of the catch-all value (equal to the initial value), in this embodiment for and Set separate lower limit values; when the dynamic calculation result is lower than the corresponding lower limit, the lower limit value will be used to avoid an overly narrow sample distribution. When the threshold is too low (in the case of extremely small values), even minor deviations will frequently trigger corrections.
[0109] S50. Automatic System Rollback: When the unified confidence level is lower than the confidence level threshold, the upper-level system rollback gating layer automatically switches the cargo location allocation control process from the AI agent decision-making mode to the upper-level system's default allocation mode, and the cargo location allocation is executed by the deterministic allocation algorithm built into the upper-level system.
[0110] This step is executed by the upper-level system's rollback gating layer, which is one of the core innovations of this invention's four-layer closed-loop governance architecture and is crucial for ensuring the safety of fully automated operation. In a fully automated warehouse scenario, manual intervention would significantly reduce system throughput, and manual intervention is impractical in a 24 / 7 unattended operation mode. Therefore, this invention innovatively replaces the human-in-the-loop in traditional AI governance with automatic system rollback (System-in-the-loop). When the AI agent makes an abnormal decision, it automatically rolls back to the deterministic allocation algorithm built into the upper-level system.
[0111] The deterministic allocation algorithm built into the host system is a rule-based fixed process algorithm pre-set at the factory by host systems such as WMS / WES / WCS. It exists in parallel with the AI agent to ensure reliable rollback.The typical process steps are as follows: Step 1 (Candidate Filtering): Perform hard constraint filtering (stacker crane accessibility, hazardous materials isolation, double-deep inner accessibility, batch FIFO / FEFO) on all available storage locations to eliminate infeasible locations; Step 2 (ABC Classification Matching): Classify the goods to be received into categories A / B / C based on their outbound frequency over the past 90 days, and filter candidate locations within the target time zone according to the preset mapping of "A→T1 (fast zone near platform), B→T2 (medium speed zone), C→T3 (slow speed zone)"; Step 3 (Proximity Allocation): Sort the candidate locations within the target time zone in ascending order of "estimated travel time from the stacker crane to the location", and select the location with the shortest estimated travel time as the preferred location; Step 4 (Warehouse Area Balancing): When the actual utilization rate of the warehouse area where the preferred location is located has reached the preset upper limit... When the threshold (default 95%) is reached, the remaining candidate storage locations within the target time zone are selected sequentially towards the storage area with the lowest actual utilization rate. (Specifically: the remaining candidate storage locations within the target time zone that have not yet been excluded are sorted in ascending order according to the current actual utilization rate of their respective storage areas, and candidate storage locations in the storage area with the lowest actual utilization rate are selected as allocation targets. If there are still multiple candidate storage locations in that storage area, the one with the shortest estimated stacker crane running time is selected according to the proximity principle in step 3. This guides goods from the hot storage area that is close to full capacity to the least idle storage area to achieve load balancing between storage areas and avoid local congestion. For example, in the target T1 area, the utilization rate of storage area A has reached 96%, storage area B is 60%, and storage area C is 75%, so candidate storage locations are selected first in storage area B.) Step 5 (backup strategy): If there are no feasible candidates in the target time zone in step 2, then the selection is done in the order of C→B→A. The system sequentially overflows to slower time zones (ensuring that high-frequency goods of category A are always selected from the nearest available location). If no candidates are found, it reverts to the "first feasible storage location FIFO selection" strategy (here, the "first feasible storage location FIFO selection" strategy means that when there are still no feasible candidates in the target area of the entire warehouse after overflowing from time zones A, B, and C, the system abandons the A, B, C partitioning and proximity optimization constraints, traverses all available storage locations in the entire warehouse, and only retains the feasible storage locations filtered by the hard constraints in step 1. These storage locations are then sorted in a first-in-first-out manner according to the timestamp when they become available (the earliest available storage location is selected first), and the first feasible storage location in the sorted list is selected). The storage location serves as the final allocation result. It should be noted that the FIFO here refers to a first-in-first-out (FIFO) fallback order based on storage location idle time, which differs in meaning and application from the batch FIFO / FEFO for the same SKU of goods entering the warehouse as described in Rule R9. This fallback strategy aims to ensure that even under extreme conditions where the entire warehouse is nearly full, the deterministic allocation algorithm can still stably and definitively provide a unique executable result, preventing deadlocks or order loss. Step 6 (Execution Distribution): The allocation result is distributed to the stacker crane as an instruction and written into the upper-level system's workflow. This deterministic allocation algorithm does not rely on any model inference, and its response time is consistently in the millisecond range, serving as a reliable alternative when the AI agent is unavailable.The deterministic allocation algorithm built into the upper-level system allocates storage locations to goods according to ABC classification rules, proximity allocation principles, and warehouse area balancing strategies. Although it does not have the adaptive optimization capabilities of AI, it has the advantages of strong determinism, fast response speed, and no need to rely on external models, making it suitable as a reliable rollback solution when AI malfunctions.
[0112] The upper-level system rollback gating layer is based on the unified confidence level output by the AI agent. The tiered rollback strategy is implemented, and the rollback gating layer of the upper-level system presets the following rollback trigger conditions: Conditional F1 (Low Confidence): The uniform confidence level of the AI agent's output. Below the first confidence threshold (like But not lower than the second confidence threshold (like Partial rollback triggered at certain times; uniform confidence level. Below the second confidence threshold A complete rollback is triggered at that time.
[0113] Condition F2 (Continuous Replanning): When the number of consecutive triggers of the full replanning mode in the same batch of assigned tasks reaches a preset threshold (e.g., 3 times), a full rollback is triggered.
[0114] Condition F3 (Main decision model has an availability anomaly): A complete rollback is triggered when the AI agent's main decision model times out (e.g., more than 10 seconds) or when 3 consecutive API calls fail.
[0115] The specific procedures for partial and full rollback are as follows: (1) Partial rollback: If The upper-level system rolls back to the gating layer to independently calculate the uniform confidence level of each item in the set of goods to be allocated. ; regarding one of them < The allocation of "low-confidence individual goods" is handled solely by a deterministic allocation algorithm built into the higher-level system; the rest... ≥ The allocation of "healthy individual goods" is still performed by the pre-allocation scheme of the storage location output by the AI agent. When performing partial rollback, the switch is based on the "single goods" as the smallest decision unit to avoid the impact of a single rollback on all goods and to retain the optimization benefits of the AI agent to the maximum extent; all switching events are recorded one by one by the end-to-end observable layer.
[0116] Full rollback: If If either condition F2 or F3 is triggered, the upper-level system's rollback gating layer will transfer the allocation rights of the entire batch of goods to be allocated to the upper-level system's built-in deterministic allocation algorithm in one go. The AI agent will not output any execution instructions in this batch, but will only use shadow mode for parallel inference to recover the decision. During the full rollback, the shadow decision results will not enter the actual execution chain, but will only be written to the full-link observable layer for subsequent auditing and model recovery evaluation.
[0117] It should be noted that the AI agent's main decision-making model refers to the core model that the AI agent currently undertakes the task of reasoning for pre-allocation of cargo spaces, and its output is the actual pre-allocation plan that is issued and executed. When the AI agent uses a multi-factor weighted scoring model to execute the pre-allocation decision, its main decision-making model is a comprehensive scoring function for cargo spaces based on multi-factor weighting. Its parameter set; when the AI agent uses a reasoning and decision-making model based on a large language model to perform the pre-allocation decision of the cargo location, its main decision-making model is the large language model (including its structured prompt word template and output parser) that undertakes the reasoning of the cargo location.
[0118] A timeout (exceeding 10 seconds) in the main decision model response or three consecutive API call failures indicate an availability anomaly in the model service (local inference or remote API) under the large language model implementation; under the multi-factor weighted scoring implementation, it indicates that the comprehensive score calculation is timed out due to blockage of the real-time data source it depends on (such as the outbound frequency table or the stacker crane status table). Either scenario will trigger a complete rollback of the upper-level system's rollback gating layer.
[0119] The upper-level system rollback gating layer also includes an automatic recovery mechanism: during the system switch to the upper-level system's default allocation mode, the AI agent still uses shadow decision-making mode to perform parallel decision reasoning for each actual allocation task and outputs a unified confidence level. The shadow decision results do not enter the actual execution link and are only used as a basis for recovery judgment. When the unified confidence level of the AI agent's shadow decisions is higher than the recovery threshold for 10 consecutive times... (like Furthermore, when the response latency and availability health indicators of the master decision-making model return to normal, it automatically switches back from the default allocation mode of the upper-level system to the AI agent decision-making mode. The mode switching event is output to the end-to-end observable layer.
[0120] S60. End-to-end observation: The end-to-end observable layer records the entire process data from steps S10 to S50 in a cross-cutting manner. The end-to-end observable layer does not participate in decision-making or control.
[0121] This step is executed by the end-to-end observability layer. The end-to-end observability layer is a cross-cutting layer in the architecture of this invention, its responsibilities strictly limited to data recording and audit support, without participating in any decision-making or control logic. This clear delineation of responsibilities is a crucial design element for the clear hierarchical structure of this invention: the four-layer closed-loop governance architecture (feedforward guidance layer, feedback sensing layer, self-calibration layer, and upper-level system backoff gating layer) is responsible for the decision-making and control of the AI agent, while the end-to-end observability layer only cross-cuts and collects process data from the above four layers and the AI agent in a "side-tracking" manner, ensuring the decoupling of governance logic from observability capabilities.
[0122] The process data recorded by this end-to-end observable layer includes: decision trigger timestamps, snapshots of input product attributes and warehouse status, a list of rule identifiers injected by the feedforward guidance layer, the pre-allocation scheme and unified confidence level output by the AI agent, deviation detection results from the feedback sensing layer, correction action records from the self-correction layer, and mode switching records from the upper-level system rollback gating layer. All data is persistently stored in a structured log format, supporting querying and audit analysis by time range, product type, deviation type, and correction mode.
[0123] In one embodiment, the intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture is further configured with a storage location pre-allocation triggering strategy: before the actual arrival of an inbound order, the system triggers the storage location pre-allocation process in advance based on the forecast arrival information received by the upper-level system (the arrival time window accuracy is ±30 minutes). The pre-allocation result marks the target storage location in a soft-locked state (the default lock timeout is 4 hours), and it is automatically released after the timeout. When the actual arrival is inconsistent with the forecast arrival information, the storage location allocation is re-executed through the full replanning mode of the self-correction layer.
[0124] In one embodiment, the intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture further includes a rule self-evolution step, which is executed by an AI agent and includes: receiving changes in business needs described by warehouse management personnel in natural language; the AI agent parses the natural language description into structured rule update instructions; and the parsed rule update instructions take effect after being confirmed by a rule verification process.
[0125] The natural language description supports at least the following three typical business scenarios: Scenario 1, peak outbound adjustment for major promotions / seasonal sales, example input "Next week's major promotion, the outbound volume of Class A goods is expected to double, and priority will be given to single-depth storage locations on the 1st and 2nd floors near the outbound exit", the parsing result is to modify the weight of the outbound frequency matching rule R6 from 0.9 to 1.1 and add a priority flag for single-depth storage locations on the 1st and 2nd floors for Class A goods; Scenario 2, adding special goods classification constraints, example input "Add hazardous chemical type X, this type of goods must be far away from electrical equipment rooms and separated from each other by no less than 3 storage locations", the parsing result is to add rule R3-X to the rigid constraint subset and set the corresponding predicate and action; Scenario 3, pausing or restricting existing rules, example input "Due to equipment maintenance, pausing all inbound allocations in the upper half of the 3rd aisle", the parsing result is to modify the predicate of the stacker crane accessibility constraint in rule R2 to add an exclusion condition for the upper half of aisle 3 and set an effective time window.
[0126] The parsing process is executed by the AI agent based on the understanding ability of the large language model, and the output is a structured rule update instruction (JSON format), which includes five fields: operation type, target rule ID, new predicate expression, new weight / threshold, and effective time window. The operation type includes three types: adding a rule, modifying the rule weight, and pausing the rule.
[0127] The rule verification process is as follows: First, the parsed rule update command is subjected to syntax and semantic verification, and the validity of the rule ID (whether the rule ID exists), the compilability of the predicate (whether the predicate can be compiled by the rule interpreter of the feedforward guidance layer), and the legality of the weight range (whether the weight is in the legal range of [0, 1]). After the verification is passed, the impact of the update on the deviation index is evaluated in the simulation sandbox by playing back the historical input of the most recent week. After the simulation evaluation is passed and the warehouse administrator confirms it a second time, the rule update command is officially issued to the feedforward guidance layer to take effect.
[0128] The criteria for passing the simulation evaluation are as follows: by replaying the historical inputs of the most recent week, the key indicators before and after the rule update are compared. The simulation is considered passed when all three of the following conditions are met: ① The uniform deviation after the update is not higher than that before the update; ② The number of violations of rigid constraint rules during the simulation replay period is zero; ③ The degradation of key business indicators, including batch first-in-first-out / first-out-of-expiration compliance rate, average outbound operation time, and the balance of warehouse capacity utilization in each warehouse area relative to the pre-update level, does not exceed the preset tolerance threshold. If any condition is not met, the simulation is considered failed, the rule update instruction is rejected, not issued, and an alarm is triggered. The reason for rejection and the comparison data are recorded by the end-to-end observable layer.
[0129] All natural language inputs, parsing results, verification and evaluation data, and final effective status are recorded by the end-to-end observable layer.
[0130] In one embodiment, the intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture further includes a rule self-evolution step: during the execution of storage location pre-allocation decisions, the AI agent performs correlation analysis on the contribution of the scoring factors of each decision and the actual execution deviation, and conducts statistical significance evaluation through a combination of chi-square independence test and mutual information measurement. When there is a significant statistical correlation between a combination of goods attributes not covered by existing rules and the actual execution deviation, the AI agent automatically generates candidate constraint rule proposals. After the candidate constraint rule proposals are verified through a preset rule verification process, they are incorporated into the structured rule set of the feedforward guidance layer.
[0131] This step is executed by the AI agent. During the pre-allocation of storage locations, the AI agent performs correlation analysis on the contribution of scoring factors to each decision and the actual execution deviation. It identifies the statistical correlation between combinations of goods attributes not covered by existing rules and the execution deviation, and automatically generates candidate constraint rule proposals. The candidate constraint rules are run in a trial state for 2 weeks and their effectiveness is automatically evaluated. The upper-level system's rollback gating layer monitors the effectiveness of the rules during the trial run.
[0132] The specific implementation of the correlation analysis is based on a two-dimensional contingency table: Record the combination of product attributes to be examined, A (e.g., "a certain product category × high outbound frequency" Boolean value), and the execution deviation, B (e.g., "whether the operation time deviation rate exceeds..."). The "Boolean values" represent two binary categorical variables; in the most recent sliding window (default On the end-to-end observable layer log, a 2×2 contingency table is formed by statistically analyzing the joint frequency of A and B for each allocation decision. (contingency table): .
[0133] Contingency Table It is A two-dimensional contingency table is a statistical table used to describe the joint frequency distribution of two categorical variables. Indicates whether the product attribute combination A is true (0 means false, 1 means true). This indicates whether the execution deviation B is true. This represents the number of allocation decision events (i.e., observed frequency) where A = i and B = j in the most recent sliding window W (default 30 days) of the end-to-end observable log. Specifically: Let A be the number of events where A is false and B is false. Let A be the number of events where A is false but B is true. Let A be the number of events where A is true but B is false. Let A be the number of events where both A and B are true. The entire... The contingency table fully describes the joint frequency distribution of the two binary variables A and B, and serves as a common input for subsequent chi-square independence tests and mutual information measurements.
[0134] Among them, the total Equivalent to the most recently sliding window The total number of allocation decision events within; rows and A represents the total number of samples in class i, i.e. The number of events in which the attribute combination A of the goods in the window takes the value i; columns and B represents the total number of samples in class j, i.e. The number of events in the window where the execution deviation B has a value of j. Substituting this into the expected frequency formula, the cell reflects the number of events under the assumption that A and B are independent. Theoretical expected frequency: ; Calculate the chi-square statistic: .
[0135] Based on this, calculate the P-value with one degree of freedom; when P < 0.01, determine that there is a statistically significant association between A and B.
[0136] Regarding the P-value ( The meaning and calculation method of P-value: The P-value is a quantitative indicator in statistical hypothesis testing. It represents the probability of observing the current chi-square statistic and more extreme results, assuming the null hypothesis (A and B are independent) holds true. A smaller P-value indicates that the observed biased result is less likely to be caused by chance under the independence hypothesis, thus supporting the alternative hypothesis that A and B are significantly correlated. The specific calculation method is as follows: (1) Calculate the contingency table Degrees of freedom: The method used in this invention Contingency table, ; (2) Chi-square statistic Substitute degrees of freedom Chi-square cumulative distribution function: ; (3) Take the area of the right tail as the p-value: That is, the chi-square distribution is in The area of the right tail at that location, where For degrees of freedom The chi-square probability density function. This invention uses a significance threshold. When p < Then determine whether A and B are statistically significantly correlated.
[0137] As a complementarity indicator, the mutual information measure between A and B is calculated using the following formula: ; in, For the value of the product attribute combination A, in this invention A is a binary variable, therefore ; The value of the deviation B is determined by the fact that it is also a binary variable. Joint probability This represents the joint probability of A = a and B = b; Marginal probability of A: ; Marginal probability of B: ; Let A = a and B = b be the number of allocation decision events. Represent the logarithm to the base 2 such that The dimension of is bit; hour, The larger the value, the stronger the correlation between A and B; This indicates that A and B are completely independent; summation iterates through all possible combinations of values for A and B (for each pair of pairs). (Total of 4 terms, summation); when As is customary: , To avoid log evaluation errors.
[0138] This invention takes threshold ,when The time was used to determine that A and B were significantly correlated. The two statistical methods were used according to... The final decision is made using either OR logic to accommodate the sensitivity differences between the two statistics in sparse sample scenarios.
[0139] Once product attribute combination A is determined to be significantly correlated with execution deviation B, the system automatically generates a form like... Candidate constraint rule proposals, each accompanied by metadata, are generated. These proposals enter the shadow rule set of the feedforward guidance layer in a trial run state, where they are evaluated in parallel with existing rules without affecting the execution of the main decision. The upper-level system's rollback gating layer continuously compares the degree of deviation improvement under the assumptions of enabling and disabling the rule within a trial run window of T (default 2 weeks). Only when the degree of improvement reaches a preset threshold is the candidate rule formally incorporated into the structured rule set of the feedforward guidance layer. All correlation analysis processes, candidate rule generation, trial run evaluation, and rule upgrade events are recorded by the end-to-end observable layer.
[0140] Further explanation regarding candidate constraint rule proposals: The format of candidate constraint rule proposals follows the standard quintuple of the feedforward guidance layer structured rule set and is expressed in two typical paradigms: Paradigm 1 (Flexible Preference): "IFattribute_combo(A) THEN DEDUCT(score_penalty)" indicates that when a candidate assignment simultaneously possesses product attribute combination A, its comprehensive score is reduced by score_penalty points; Paradigm 2 (Rigid Constraint): "IFattribute_combo(A) THEN REJECT" indicates that when a candidate assignment simultaneously possesses product attribute combination A, it is directly removed from the feasible candidate set. The paradigm selection is automatically determined by the "bias severity" obtained from the association analysis: when the execution bias B corresponding to product attribute combination A is absolutely physically infeasible (such as double-depth reachability failure, stacker crane overload), Paradigm 2 is used; otherwise, Paradigm 1 is used. The candidate constraint rule proposal comes with 6 metadata items: (a) the precise formal definition of attribute_combo(A) (Boolean expression or SQL-like predicate); (b) the association strength (in words); and (c) the relationship strength. Value and (a) Values given in parallel; (c) Number of historical samples; (d) 95% confidence interval; (e) Suggested initial weights (based on...) (f) Suggested initial score_penalty value. All candidate constraint rule proposals do not take effect directly, but instead enter the shadow rule set of the feedforward guidance layer in a trial state. Within the trial window of T (default 2 weeks), the upper system back-off gating layer evaluates the degree of deviation improvement under the two assumptions of "enabling / disabling the rule" in parallel. Only when the degree of improvement reaches the preset threshold is it formally included in the rule set. Proposals that fail the evaluation are automatically archived to the end-to-end observable layer for subsequent traceability.
[0141] In one embodiment, the intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture further includes a multi-agent parallel allocation step: when the number of goods to be allocated exceeds a preset batch threshold, the main AI agent first establishes a global SKU batch order index according to the SKU identifier; then, the main AI agent splits the task to be allocated into multiple sub-tasks according to the lane dimension and distributes them to each sub-agent; each sub-agent executes the storage location pre-allocation decision in parallel within its respective lane, and references the global SKU batch order index when filtering rigid constraints to ensure that the first-in-first-out order of cross-lane batches is not violated; when the local optimal solution of a sub-agent conflicts with the global SKU batch order or requires cross-lane overflow allocation, the main AI agent coordinates the cross-lane resource competition and the global SKU batch order constraint through a distributed lock mechanism to avoid multiple sub-agents competing for the same storage location at the same time or violating the global SKU batch order.
[0142] In one embodiment, the AI agent includes an AI agent and multiple sub-agents, wherein the sub-agents inherit the rule set injected by the feedforward guidance layer in the main AI agent and share the same end-to-end observable layer and upper-level system fallback gating layer.
[0143] When the quantity of goods received in a batch exceeds the preset batch threshold, the system activates a multi-agent parallel allocation mechanism. To address the issue that independent decision-making by sub-agents according to lanes may violate the global batch first-in-first-out (FIFO / FEFO) order, a global SKU-batch order index is introduced as a coordination middleware. Its workflow is as follows: Step A: Before splitting the sub-tasks, the main AI agent first scans all the goods to be allocated and the batches already stored in the warehouse through the global SKU coordinator, clusters them by SKU identifier, and sorts each SKU in ascending order by the entry timestamp (or expiration date, for the FEFO scenario) to generate the global SKU-batch sequence index table GlobalIdx; Step B: The main AI agent splits the task according to the lane dimension and sends the sub-tasks along with a read-only copy of the global SKU-batch sequence index table GlobalIdx to each sub-agent; Step C: Each sub-agent executes the pre-allocation decision of storage locations in parallel within its respective aisle. When performing feasibility filtering under rule R9 (batch FIFO / FEFO constraint), the sub-agent queries the global SKU-batch order index table GlobalIdx for each candidate storage location to confirm that the relative position of the batch to be allocated in the SKU sequence will not disrupt the global order. If there is a potential allocation that disrupts the global order (e.g., a newer batch is allocated to a location closer to the outbound outlet than an older batch), the candidate storage location is marked as infeasible. Step D: When a sub-agent's local optimal solution conflicts with the global SKU batch order or when goods need to be overflowed and allocated to other lanes, the overflow request is uniformly handled by the master AI agent: the master agent acquires a distributed lock on the candidate storage location in the target lane, executes the allocation decision, and releases the lock after completion, ensuring that multiple sub-agents do not compete for the same storage location simultaneously and that the global FIFO / FEFO order is not disrupted. The lock waiting timeout is set to 2 seconds. After the timeout, the master agent automatically selects other available lanes or triggers the upper-level system to roll back the gating layer.
[0144] This multi-agent parallel allocation mechanism achieves both the speed advantage of multi-agent parallelism and the correctness of the global SKU batch order, solving the cross-lane FIFO / FEFO violation problem that may be caused by traditional independent parallel allocation by lane.
[0145] The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture provided in this embodiment of the invention has the following beneficial technical effects compared with existing storage location allocation methods: (1) By actively injecting the structured rule set that separates rigid constraints and flexible constraints into the decision context of the AI agent through the feedforward guidance layer, the physical feasibility problem of the AI agent in the physical execution scenario is fundamentally solved, and the reliability of the AI agent's decision is guaranteed. (2) By forming a closed loop through the feedback sensing layer and the self-correction layer, the system can automatically correct the allocation strategy according to the actual execution deviation, and realize the continuous adaptive optimization of the allocation strategy; the hierarchical correction mechanism of dual threshold and three modes (parameter fine-tuning, strategy rollback, and full replanning) enables the system to automatically select the most economical repair path according to the severity of the deviation, which significantly reduces system jitter compared with the single full replanning method. (3) The system automatically rolls back based on a unified confidence level by using the rollback gating layer of the upper system. When the AI agent makes an abnormal decision, it can automatically switch to the default algorithm of the upper system without manual intervention, and can automatically switch back after the AI agent recovers, ensuring the continuity and security of fully automated operation. (4) The end-to-end observable layer records the entire chain of data from decision triggering to final execution, making the basis, process, result and subsequent adjustment of each allocation decision traceable, realizing the auditability of the complete decision chain, and greatly improving the operability and credibility of the system; the end-to-end observable layer and the four-layer closed-loop governance architecture are strictly separated in terms of responsibilities (do not participate in decision-making and control), ensuring that the governance logic and observability are decoupled, which makes it easy for the operation and maintenance side to independently expand the audit strategy without affecting the main decision path; (5) Through the cargo architecture perception mechanism, a single system can uniformly adapt to three cargo architecture types: pure single-deep, pure double-deep, and hybrid, without the need to develop independent systems for different architectures; (6) The AI agent, as a replaceable component, supports two implementation methods: a decision-making model based on multi-factor weighted scoring and a reasoning decision-making model based on a large language model, which facilitates flexible selection in different warehouse scales and business complexity scenarios.
[0146] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0147] Specific limitations regarding the intelligent storage location pre-allocation system based on a four-layer closed-loop governance architecture can be found in the limitations of the intelligent storage location pre-allocation method based on this architecture mentioned above, and will not be repeated here. Each module in the aforementioned intelligent storage location pre-allocation system based on a four-layer closed-loop governance architecture can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0148] In one embodiment, a computer device is provided, the internal structure of which can be shown as follows: Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with other external electronic devices via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the aforementioned intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture.
[0149] In one embodiment, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the functions or steps of the above-described intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture.
[0150] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0151] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A smart storage location pre-allocation method based on a four-layer closed-loop governance architecture, characterized in that, Includes the following steps: Feedforward constraint injection: The physical world constraints of the allocation of storage locations in the automated warehouse are encoded into a set of structured rules through the feedforward guidance layer, and the corresponding rules of the set of structured rules are injected into the decision context of the AI agent before each storage location pre-allocation decision. The set of structured rules is divided into a subset of rigid constraint rules and a subset of flexible constraint rules. Intelligent decision-making and confidence assessment: Guided by the decision-making context, the AI agent outputs a pre-allocation plan for cargo locations and a unified confidence level for the selected optimal cargo location; Execution deviation feedback: The actual execution deviation of the storage location pre-allocation plan output by the AI agent is detected by the feedback sensing layer and the uniform deviation magnitude is calculated; Hierarchical self-correction: The self-correction layer compares the uniform deviation amplitude with the first threshold and the second threshold, and automatically selects one of the three correction modes—parameter fine-tuning, strategy rollback, and complete replanning—according to the dual-threshold three-mode mechanism, and updates the strategy version used by the AI agent in the next execution of the pre-allocation decision of the storage location. Automatic system rollback: When the unified confidence level is lower than the confidence level threshold, the upper-level system rollback gating layer automatically switches the storage location allocation control process from the AI agent decision-making mode to the upper-level system's default allocation mode, and the storage location allocation is executed by the upper-level system's built-in deterministic allocation algorithm; End-to-end observation: The end-to-end observable layer records data from the entire process of feedforward constraint injection, intelligent decision-making and confidence assessment, execution deviation feedback, hierarchical self-correction and automatic system rollback in a cross-cutting manner. The end-to-end observable layer does not participate in decision-making and control.
2. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, The intelligent decision-making and confidence assessment steps further include: For each candidate vacant storage location in the automated warehouse, feasibility filtering is performed based on the rigid constraint rules injected by the feedforward guidance layer to exclude storage locations that do not meet the rigid constraints and generate a set of feasible storage location candidates. For each candidate location in the feasible location candidate set, calculate the comprehensive score. ,in Let i be the weight of the i-th rating factor. , Let be the calculation function for the i-th rating factor. The total number of rating factors; Candidate storage locations are sorted in descending order of their overall scores, and the optimal storage location with the highest overall score is selected. As the allocation result, a pre-allocation scheme for storage locations is generated, and the unified confidence score of the optimal storage location selected by the AI agent is calculated: , in Optimal storage location selected for the AI agent Overall score; The historical average comprehensive score of all optimal storage locations within the most recent preset period; This is a relative score difference component. In a multi-candidate scenario, select , The overall score for the second-best cargo location. To preset a minimum positive value, In the case of only a single candidate scenario, a preset neutral value of 0.5 is used; Preset fusion weight coefficients; Indicates will Truncate to the [0, 1] interval and use a uniform confidence level. The value of is constant in the interval [0, 1].
3. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, The execution deviation feedback step further includes: Based on the estimated running time and actual running time of the stacker crane calculated using the stacker crane's kinematic model, the stacker crane's running time deviation rate is calculated. ,in, This refers to the actual running time of a single operation by the stacker crane. Estimate the running time for a single operation of the stacker crane; Calculate the deviation rate of double-deep cargo displacement. ,in, The deviation rate of the transfer frequency dimension. , This refers to the actual number of data transfer operations triggered within the statistical period. The estimated number of warehouse transfer operations within the same statistical period for the warehouse location pre-allocation scheme output by the AI agent; The deviation rate in the time dimension of the transfer of goods. , This represents the actual total time spent on all warehouse transfer operations within the statistical period. The estimated total time for all warehouse transfer operations within the same statistical period is calculated based on the pre-allocation plan for storage locations output to the AI agent. To prevent division by zero by setting a minimum positive value, The internal fusion weights between the transfer frequency dimension and the transfer time dimension are defined as follows: norm(·) represents a function that is normalized to the interval [0, 1] by max-min. Calculate the deviation of warehouse capacity utilization ,in, Represents the reservoir area Deviation in storage capacity utilization , Represents the reservoir area The target utilization rate of storage capacity Represents the reservoir area Actual utilization rate of storage capacity , This represents the number of currently occupied storage spaces within storage area z. This represents the total number of storage locations within storage area z. The ratio of the number of abnormal outbound events that did not follow the FIFO / FEFO order to the total number of outbound operations within a preset period is used as the batch first-in-first-out deviation rate. ; The stacker crane running time deviation rate Double-deep cargo displacement deviation rate Storage capacity utilization deviation Batch First-In-First-Out Deviation Rate After normalization, the uniform bias amplitude obtained by weighted fusion is... .
4. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 3, characterized in that, The hierarchical self-calibration step further includes: Unify deviation range With the first threshold Second threshold If a comparison is made, In the parameter fine-tuning mode, only the weight parameters in the comprehensive scoring function of the cargo location are adjusted; if Execute the strategy rollback mode, reverting the allocation strategy to the historical strategy version with the best deviation metric in the strategy version stack; if The system executes a full replanning mode, triggering the AI agent to re-plan the allocation of storage locations based on the current warehouse status. The newly generated strategy is added to the strategy version stack as a new strategy version.
5. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 4, characterized in that, If the number of system initialization or unified deviation amplitude samples is less than N, the first threshold is set using the initial value. Second threshold ; If the number of samples with uniform deviation amplitude reaches N, the mean of the samples with uniform deviation amplitude within the statistical window will be used. and standard deviation Based on this, the first threshold is adaptively and dynamically adjusted. Second threshold : , , in, N represents the current cumulative uniform deviation amplitude sample size, and N is the statistical window length. This is the lower limit of the first threshold, which serves as a safety net. This is the lower limit of the second threshold; If the uniform deviation amplitude is lower than the first threshold for M consecutive statistical periods When, the first threshold Second threshold Each threshold is tightened by multiplying by a tightening factor less than 1, where the tightening factor ranges from (0, 1); if the number of times the full replanning mode is triggered within a single statistical period exceeds a frequency threshold, the first threshold is adjusted. Second threshold Each threshold is relaxed by multiplying by a relaxation factor greater than 1, where M is a preset value.
6. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, The automatic rollback step of the system further includes: when When the execution rollback occurs, the uniform confidence level of each item in the set of items to be allocated is calculated independently. Items with a uniform confidence level lower than [a certain value] are considered unqualified. Individual goods are allocated by a deterministic allocation algorithm built into the higher-level system, with a uniform confidence level greater than or equal to... Individual goods are allocated using a pre-allocation scheme for storage locations output by an AI agent; when When the number of times the full replanning mode is triggered consecutively in the same batch of allocation tasks reaches the preset threshold or when the availability of the AI agent's main decision-making model is abnormal, a full rollback is executed. The upper-level system rollback gating layer transfers the allocation rights of the entire batch of goods to be allocated to the upper-level system's built-in deterministic allocation algorithm at once. in, The unified confidence level for the optimal storage location selected by the AI agent. The first confidence threshold is... The second confidence threshold, .
7. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, It also includes the rule self-evolution step: During the process of making pre-allocation decisions for cargo locations, the AI agent performs correlation analysis on the contribution of scoring factors to each decision and the actual execution deviation. It conducts statistical significance evaluation through a combination of chi-square independence test and mutual information measurement criteria. When there is a significant statistical correlation between a combination of cargo attributes not covered by existing rules and the actual execution deviation, the AI agent automatically generates candidate constraint rule proposals. After the candidate constraint rule proposals are verified through a preset rule verification process, they are incorporated into the structured rule set of the feedforward guidance layer.
8. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, It also includes a multi-agent parallel allocation step: When the number of goods to be allocated exceeds the preset batch threshold, the main AI agent first establishes a global SKU batch order index based on the SKU identifier. Then, the main AI agent breaks down the task to be allocated into multiple sub-tasks according to the lane dimension and distributes them to each sub-agent. Each sub-agent executes the pre-allocation decision of the storage location in parallel within its own lane, and references the global SKU batch order index when filtering rigid constraints to ensure that the first-in-first-out order of batches across lanes is not violated. When the local optimal solution of a sub-agent conflicts with the global SKU batch order or requires cross-lane overflow allocation, the main AI agent coordinates the cross-lane resource competition and the global SKU batch order constraint through a distributed lock mechanism to avoid multiple sub-agents competing for the same storage location at the same time or violating the global SKU batch order.
9. The intelligent storage location pre-allocation method based on a four-layer closed-loop governance architecture as described in claim 1, characterized in that, It also includes the natural language policy update step: Upon receiving changes in business requirements described by warehouse managers in natural language, the AI agent parses the natural language description into structured rule update instructions. The rule update instructions include five fields: operation type, target rule ID, new predicate expression, new weight / threshold, and effective time window. Perform syntax and semantic checks on the parsed rule update instructions, and verify the validity of the rule ID, the compilability of the predicate, and the legality of the weight range; After the verification is passed, the impact of the update on the deviation index is evaluated in the simulation sandbox by playing back the historical inputs of the most recent week. After the simulation evaluation is passed and the warehouse administrator confirms it a second time, the rule update instruction is officially issued to the feedforward guidance layer to take effect.
10. An intelligent storage location pre-allocation system based on a four-layer closed-loop governance architecture, characterized in that, include: The feedforward guidance layer is used to encode the physical world constraints of the allocation of storage locations in the automated warehouse into a structured rule set, and inject the corresponding rules of the structured rule set into the decision context of the AI agent before each storage location pre-allocation decision. The structured rule set is divided into a subset of rigid constraint rules and a subset of flexible constraint rules. The AI agent is used to output a pre-allocation scheme for cargo space and a uniform confidence level for the selected optimal cargo space, guided by the decision context. The feedback sensing layer is used to detect the actual execution deviation of the pre-allocation plan for the storage location output by the AI agent and calculate the uniform deviation range. The self-calibration layer is used to compare the uniform deviation amplitude with the first threshold and the second threshold, and automatically select one of the three calibration modes—parameter fine-tuning, strategy rollback, and complete replanning—according to the dual threshold three-mode mechanism, and update the strategy version used by the AI agent in the next execution of the cargo location pre-allocation decision. The upper-level system rollback gating layer is used to automatically switch the storage location allocation control process from the AI agent decision-making mode to the upper-level system default allocation mode when the unified confidence level is lower than the confidence level threshold, and the storage location allocation is executed by the deterministic allocation algorithm built into the upper-level system. The end-to-end observable layer is used to record the entire process data of the feedforward guidance layer, feedback sensing layer, self-calibration layer, upper system backoff gating layer and AI agent in a cross-sectional manner. The end-to-end observable layer does not participate in decision-making and control.