Intelligent mobile terminal-oriented warehouse space gridding AI path planning method

CN122820097APending Publication Date: 2026-09-25WUHAN GAODA SOFTWARE SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611253229.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,在高并发的动态调度场景下,多源终端数据格式的异构无序极大地加重了云端解析负担;同时,连续空间下的寻优计算引发了严重的高维状态空间膨胀现象,特别是节点增多时,浮点运算易导致服务器算力枯竭与阻塞

Benefits of technology

[0055]本发明通过构建关系映射表与离散化映射生成空间事件流和动态网格状态张量,将多源异构状态完全收敛,规避了连续空间下的高维状态空间膨胀,解决传统寻优中解析负担重与算力枯竭痛点;基于强化学习的动作价值张量将复杂调度转化为叠加通行阻抗的量化预期收益评估,生成确定性的网格路径序列,通过网格化离散化有效降低了多终端并发时的轨迹冲突概率;无损压缩的轻量化指令包结合异常状态增量局部更新机制,彻底摒弃全量状态回传,在极小通信开销下实现实时精准同步,大幅缩减计算与通信双重时延。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820097A_ABST
    Figure CN122820097A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of warehouse logistics scheduling and artificial intelligence path planning, and discloses a warehouse space gridding AI path planning method for an intelligent mobile terminal. The method comprises the following steps: constructing a relation mapping table to analyze and normalize original operation data to generate a space event stream, performing space dimension reduction and discretization mapping to generate a dynamic grid state tensor; calling the dynamic grid state tensor to calculate a penalty coefficient to obtain an action value tensor, performing path backtracking to generate a grid path sequence; performing lossless compression on the grid path sequence to generate a lightweight instruction package, and issuing the lightweight instruction package; and based on an abnormal state increment returned by the terminal, locally updating the dynamic grid state tensor. The application normalizes multi-source heterogeneous data, greatly reduces the dimension of space data processing, avoids the calculation power exhaustion caused by floating point operation, and solves the trajectory conflict and system deadlock through the compression and increment mechanism, which significantly reduces the communication load and calculation time delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of warehouse logistics scheduling and artificial intelligence path planning technology, specifically to a gridded AI path planning method for warehouse space for intelligent mobile terminals. Background Technology

[0002] With the development of smart IoT technology, warehousing systems typically employ a collaborative approach between centralized cloud servers and multi-source intelligent mobile terminals to achieve logistics scheduling and route planning functions. In the traditional route optimization process, various terminals continuously report incompatible absolute spatial coordinate data to the cloud. The central server calculates floating-point motion trajectories based on a continuous spatial coordinate system using global or heuristic algorithms and then issues the commands for execution.

[0003] However, in high-concurrency dynamic scheduling scenarios, the heterogeneous and disordered data formats from multiple terminals significantly increase the cloud's parsing burden. Simultaneously, optimization computation in continuous space leads to severe high-dimensional state space expansion, especially as the number of nodes increases, where floating-point operations easily cause server computing power exhaustion and congestion. Furthermore, the traditional plaintext path delivery and full state feedback mechanism after obstacle encounters further exacerbate the communication load. This dual latency caused by excessively high spatial computation dimensionality and redundant communication data results in path instructions received by the terminal lagging behind the actual physical environment, leading to trajectory conflicts and system deadlocks. Therefore, how to normalize multi-source heterogeneous data, reduce the spatial data processing dimensionality, and decrease computation and communication latency in the scheduling process of high-concurrency intelligent mobile terminals is a problem that urgently needs to be solved in this field. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to provide a gridded AI path planning method for warehouse space on intelligent mobile terminals. This method uses a relational mapping table to perform parsing and normalization on heterogeneous and disordered raw operational data to generate a spatial event stream, and then reduces its dimensionality by mapping it to a dynamic grid state tensor. A deep Q-value network is introduced to calculate penalty coefficients and output action value tensors, enabling explicit path backtracking to generate a grid path sequence. Run-length encoding is used for lossless compression to generate lightweight instruction packets for execution, combined with incremental triggering of local updates based on abnormal states. This achieves the effect of resolving heterogeneous and disordered states and effectively cutting off computational blockage caused by the expansion of high-dimensional state space in continuous space. Simultaneously, relying on data dimensionality compression and extremely low communication load mechanisms, it significantly reduces both latency and ensures trajectory safety and closed-loop scheduling.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] In a first aspect, the present invention provides a gridded AI path planning method for warehouse space for smart mobile terminals, comprising:

[0007] Obtain the raw job data, construct a relational mapping table, perform parsing normalization on the raw job data, generate a spatial event stream, perform spatial dimension reduction discretization mapping based on the spatial event stream, and generate a dynamic grid state tensor.

[0008] Obtain the starting grid node and target grid node when the intelligent mobile terminal initiates a job request, call the dynamic grid state tensor to calculate the penalty coefficient of the starting grid node and target grid node to obtain the action value tensor, and perform path backtracking based on the action value tensor to generate a grid path sequence;

[0009] Lossless compression encoding is performed on the grid path sequence to generate a lightweight instruction package, which is then sent to the target smart mobile terminal to execute an adaptive job loop. Each loop of the adaptive job loop is used for detection, and the obstacle detection results are output. The obstacle detection results are then processed by branching and setting updates to obtain the local update results of the dynamic grid state tensor.

[0010] Furthermore, the method for acquiring the spatial event stream includes:

[0011] Parse the network packets encapsulated from the original job data, obtain the device type identifier in the packet header, and extract the local timestamp carried by the network packet;

[0012] Using the device type identifier as an index, the private fields in the corresponding message payload are extracted from the relational mapping table. The private fields are then filled into the standard fields of the preset unified data pattern, and the local timestamp is filled into the standard timestamp field of the unified data pattern to generate the corresponding structured record.

[0013] Multiple structured records generated for each network message are aggregated and arranged to generate a spatial event stream.

[0014] Furthermore, the method for obtaining the structured record includes:

[0015] The unified data schema includes a standard device identifier field, a standard trigger type field, a standard coordinate field, and a standard queuing field;

[0016] The private fields are mapped and filled into the standard device identifier field, standard trigger type field, standard coordinate field, and standard queue field, respectively.

[0017] Convert the local timestamp to a time zone and populate it into the standard timestamp field to output a structured record.

[0018] Furthermore, the method for obtaining the dynamic mesh state tensor includes:

[0019] Within the storage space, the grid is divided along three preset orthogonal axes according to preset grid side lengths to construct a three-dimensional grid index composed of multiple grid nodes.

[0020] Extract the standard coordinate field from each structured record in the spatial event stream as the current physical coordinate. Divide the value of the current physical coordinate in the three orthogonal axis directions by the grid side length and round down to obtain three integer index values, and then concatenate them to form the discrete grid coordinate code corresponding to the current physical coordinate.

[0021] Each grid node is assigned a dynamic feature vector, which includes space occupancy bits, passage impedance value, and risk warning threshold.

[0022] By aggregating the dynamic feature vectors based on the coordinate codes of each discrete grid, a dynamic grid state tensor is generated.

[0023] Furthermore, after assigning dynamic feature vectors to each grid node, the process further includes:

[0024] Parse the standard trigger type field in the structured record and extract the event category represented by the standard trigger type field. The event category includes cargo scanning event, warehouse age warning event and access control gate queuing event.

[0025] If the event category is a cargo scanning event, the space occupancy bit of the corresponding discrete grid coordinate encoding associated grid node is set to 1. If the event category is a warehouse age warning event or an access control gate queuing event, the space occupancy bit of the corresponding grid node is set to 1, and incremental superposition calculation is performed on the passage impedance value of the grid node.

[0026] Furthermore, the method for obtaining the action value tensor includes:

[0027] Traverse each grid node in the 3D grid index that has zero spatial occupancy, extract the local state vector of each grid node and input it into the pre-trained reinforcement learning model, and output the original action value vector corresponding to each grid node.

[0028] The penalty coefficient of each grid node is calculated using the starting grid node, the target grid node, and the passage impedance value. The penalty coefficient is then superimposed on the corresponding original action value vector to generate the corrected action value vector for each grid node.

[0029] Combine the modified action value vectors corresponding to all grid nodes to obtain the action value tensor.

[0030] Furthermore, the method for calculating the corrected action value vector includes:

[0031] Calculate the difference between the passage impedance value of the corresponding grid node and the risk warning threshold, extract the maximum value between the difference and zero using the maximum value function, and multiply it by the preset over-limit penalty coefficient to obtain the over-limit penalty item;

[0032] The gradient penalty term is obtained by multiplying the impedance value by a preset impedance gradient penalty coefficient.

[0033] The penalty coefficient is obtained by algebraically summing the over-limit penalty term and the gradient penalty term and taking the negative value. The penalty coefficient is then added to the directional component of the corresponding original action value vector as an algebraic bias term to obtain the corrected action value vector.

[0034] Furthermore, the method for obtaining the grid path sequence includes:

[0035] Create a path node list and write the discrete grid coordinate code corresponding to the starting grid node as the first element;

[0036] Set the current grid node pointer to initially point to the starting grid node, triggering a path backtracking loop;

[0037] In the path backtracking loop, extract the corrected action value vector corresponding to the grid node pointed to by the current grid node pointer in the action value tensor, and obtain the component values ​​of multiple discrete movement directions in the corrected action value vector.

[0038] Select the discrete movement direction corresponding to the largest component value for coordinate offset, generate the discrete grid coordinate code of the next grid node and append it to the path node list, and update the current grid node pointer to point to the next grid node;

[0039] The loop terminates when the current grid node pointer points to the target grid node, and the grid path sequence is output.

[0040] Furthermore, the method for obtaining the lightweight instruction package includes:

[0041] For each adjacent discrete grid coordinate code in the grid path sequence, perform a component-wise integer subtraction operation with the previous discrete grid coordinate code to obtain a differential coordinate sequence composed of multiple coordinate increments;

[0042] Replace consecutively repeated identical coordinate increments in the differential coordinate sequence with a tuple containing the coordinate increment and the number of consecutive repetitions to generate a compressed bitstream;

[0043] Obtain the device identifier of the target smart mobile terminal;

[0044] The compressed bitstream, the discrete grid coordinate encoding corresponding to the starting grid node, and the device identifier of the target smart mobile terminal are encapsulated to generate a lightweight instruction package.

[0045] Furthermore, the obstacle detection results include:

[0046] Obtain the obstacle detection results output by the target intelligent mobile terminal in the adaptive job cycle, the obstacle detection results including the detection of temporary obstacles and the absence of temporary obstacles;

[0047] If the obstacle detection result is that a temporary obstacle is detected, the abnormal state increment is received from the target smart mobile terminal. The abnormal state increment is composed of the discrete grid coordinate code of the abnormal grid node and the abnormal impedance setting value.

[0048] The abnormal grid node is the current grid node that has detected a temporary obstacle, and the abnormal impedance setting value is a constant with a value greater than the risk warning threshold.

[0049] Furthermore, the method for obtaining the local update result includes:

[0050] Based on the discrete grid coordinate encoding of the abnormal grid nodes, the corresponding abnormal grid nodes are located in the dynamic grid state tensor.

[0051] Replace the pass-through impedance value of the abnormal grid node directly with the abnormal impedance setting value;

[0052] Maintaining the dynamic eigenvectors of the remaining grid nodes in the dynamic grid state tensor unchanged, the local update result of the dynamic grid state tensor is generated;

[0053] The method for obtaining the action value tensor and the method for obtaining the grid path sequence are retried using the results of local updates, so as to regenerate the grid path sequence.

[0054] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0055] This invention generates spatial event streams and dynamic grid state tensors by constructing a relational mapping table and discretizing the mapping, completely converging multi-source heterogeneous states and avoiding high-dimensional state space expansion in continuous space. This solves the pain points of heavy analytical burden and computational exhaustion in traditional optimization. Based on the action value tensor of reinforcement learning, complex scheduling is transformed into a quantitative expected benefit assessment with superimposed passage impedance, generating a deterministic grid path sequence. The grid discretization effectively reduces the probability of trajectory conflicts when multiple terminals are concurrent. The lossless compressed lightweight instruction package combined with the incremental local update mechanism of abnormal states completely eliminates the need for full state backhaul, achieving real-time and accurate synchronization with minimal communication overhead, and significantly reducing both computation and communication latency. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a flowchart of a warehouse space grid-based AI path planning method for smart mobile terminals provided in an embodiment of the present invention.

[0058] Figure 2 This is a schematic diagram of the physical architecture for multi-source heterogeneous terminal status data acquisition and event stream normalization processing provided in an embodiment of the present invention;

[0059] Figure 3 This is a schematic diagram illustrating the principle of dimensionality reduction discretization of warehouse space and construction of dynamic grid state tensors provided in an embodiment of the present invention.

[0060] Figure 4 This embodiment provides a schematic diagram illustrating the deductive principle of greedy path backtracking based on the action value distribution of grid nodes.

[0061] Figure 5 This invention provides a schematic diagram of the timing and data structure for lossless compression and delivery of a grid path sequence and its decoding execution at the terminal. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Example 1

[0064] Please see Figure 1 As shown, this embodiment provides a gridded AI path planning method for warehouse space for smart mobile terminals, including:

[0065] Step S10: Obtain the original job data, construct a relational mapping table, perform parsing normalization on the original job data, generate a spatial event stream, perform spatial dimensionality reduction discretization mapping based on the spatial event stream, and generate a dynamic grid state tensor.

[0066] Further, step S10 includes:

[0067] Step S11: Obtain the original job data, construct a relational mapping table, perform parsing and normalization operations on the original job data, and generate a spatial event stream.

[0068] During the continuous operation of large-scale intelligent IoT warehousing systems, especially in high-concurrency dynamic scheduling scenarios where centralized cloud servers and multi-source intelligent mobile terminals collaborate to achieve logistics scheduling and route planning, various heterogeneous physical devices distributed within the warehousing area need to continuously report their perceived operational status to the cloud for unified optimization. However, in existing technologies, cloud servers typically directly receive and rely on the absolute spatial coordinates continuously reported by mobile terminals. Since PDA devices, access control gate devices, and weighbridge nodes belong to different manufacturers, use incompatible communication protocols and data field formats, and their respective reporting frequencies are inconsistent with local clock references, the raw operational data aggregated on the cloud server side exhibits a severely heterogeneous and disordered state in terms of data structure, with random misalignments on the time axis. When the data volume and the number of concurrent request nodes increase, if this raw operational data with divergent formats and misaligned time sequences is directly entered into the downstream optimization algorithm without unified abstraction, it will force the optimization algorithm to repeatedly perform field adaptation and coordinate parsing in a continuous spatial coordinate system, exacerbating the high-dimensional state space expansion phenomenon, and ultimately leading to server-side computing power exhaustion and processing blockage. Therefore, a normalization mechanism for multi-source heterogeneous terminal status data based on standard communication interfaces and preset data cleaning rules is established. The purpose is to transform isolated and disordered raw operation data belonging to different physical devices into a spatial event stream carrying a unified timestamp that can be directly processed by a computer.

[0069] Specifically, raw operational data uploaded by smart mobile terminals distributed throughout the warehouse area is acquired in real time through a standard communication interface deployed at the edge access side of the cloud server. The standard communication interface refers to a network communication software module that supports dual-stack access of Hypertext Transfer Protocol and Message Queue Telemetry Transmission Protocol, and exposes a unified message subscription endpoint. This module enables low-coupling data access between the heterogeneous physical device layer and the cloud server computing layer. The smart mobile terminal refers to a collection of physical devices distributed throughout the warehouse area with local data acquisition and wireless backhaul capabilities, including PDA devices, access control gate devices, and weighbridge nodes. A PDA device is a handheld or vehicle-mounted mobile data terminal carried by warehouse personnel or automated guided vehicles, used to scan cargo barcodes and report its own location coordinates. Its internal payload includes cargo identification and the terminal's real-time location coordinates. The cargo identification is a globally unique electronic material code, such as a barcode, QR code, or RFID tag text string, used to uniquely identify the scanned cargo. The terminal's real-time location coordinates are the coordinates obtained by the PDA device at the absolute instant of the scanning operation, using its built-in positioning chip and based on the unified reference origin of the warehouse area. Three-dimensional continuous space floating-point coordinates; Access control gate equipment refers to a fixed gate control node deployed at the entrance and exit of a storage area to detect the passage and queuing status of vehicles. Its internal load consists of a passage trigger signal and the current queue vehicle count. The passage trigger signal is a Boolean value representing a vehicle requesting passage, generated when the inductive loop or infrared sensor built into the access control gate equipment is blocked by a passing vehicle. The current queue vehicle count is an integer value maintained by the access control gate equipment control chip, representing the total number of vehicles currently queuing in the waiting area before the gate. Weighbridge node refers to a fixed weighing node deployed in the receiving and shipping area to detect the weight of goods and trigger inventory aging timekeeping. Its internal load consists of a weighing amplitude and a goods entry timestamp. The weighing amplitude is a floating-point value representing the physical weight output by the gravity sensor of the weighbridge node after weighing the material on the weighing platform. The goods entry timestamp is a digital time tag generated by the weighbridge node when it detects stable weighing and successful code reading, representing the material officially entering the inventory management process. The raw operational data refers to the data set reported by the aforementioned three types of intelligent mobile terminals through standard communication interfaces and encapsulated in a network message format. The network message is structurally composed of a message header and a message payload sequentially concatenated.The message header includes a fixed-length metadata field for uniquely indicating the type of reporting device, defined as a device type identifier, and a fixed-length field for uniquely tracing the terminal identity, defined as a device unique identifier; the specific values ​​of these fields represent PDA devices, access control gate devices, and weighbridge nodes, respectively. The message payload contains data encapsulated using the private field format of each physical device. When the device type identifier represents a PDA device, the corresponding message payload includes a goods identifier encapsulated in a first private field and the terminal's real-time location coordinates encapsulated in a second private field. When the device type identifier represents an access control gate device, the corresponding message payload includes a passage trigger signal encapsulated in a third private field and the current queue vehicle count value encapsulated in a fourth private field. When the device type identifier represents a weighbridge node, the corresponding message payload includes a weighing amplitude encapsulated in a fifth private field and a goods entry timestamp encapsulated in a sixth private field.

[0070] To transform heterogeneous raw operational data into unified and processable structured data, the cloud server constructs a relational mapping table during system initialization. Specifically, a unified data schema is defined in the cloud server's memory. This unified data schema consists of four standard fields: a standard device identifier field, a standard trigger type field, a standard coordinate field, and a standard queuing field. The standard device identifier field is a string or long integer variable used to uniquely trace the core entity of a record within the global warehousing system. Its internal payload can be the hardware number of a fixed physical device or the global electronic material code of the processed goods. The standard trigger type field is an integer enumeration variable used to carry the converged standardized operational event category, with its value space strictly limited to integers representing different business dimensions. The standard coordinate field is a three-dimensional floating-point array variable used to uniformly represent the location of the event, and its data structure is required to include three-dimensional coordinate values ​​relative to a unified reference origin of the warehousing area. The standard queuing field is an unsigned integer variable used to quantify and record the current spatial congestion level at the event location. System administrators, in the configuration library of the cloud server, establish a one-to-one correspondence mapping table between each private field in the message payload and each standard field in the unified data model, based on the identifiers representing three device types: PDA devices, access control gate devices, and weighbridge nodes. Specifically, the mapping table is as follows: For PDA devices, since the main operating subject is the scanned goods, a mapping is established from the global electronic material code of the goods to the standard device identifier field. The hardware number of the PDA device itself is recorded separately through the device unique identification field. The standard trigger type field is fixedly and forcibly assigned the value 1 and defined as a goods scanning event. A mapping is established from the second private field to the standard coordinate field, and the standard queuing field is fixedly assigned the default value of 0. For access control gate devices, the standard device identifier field is assigned the hardware number of the access control gate device. The standard trigger type field is fixedly and forcibly assigned the value 3, and defined as an access control gate queuing event. The standard coordinate field is assigned the fixed physical coordinates of the access control gate device. The mapping of the fourth private field to the standard queuing field is established. The fixed physical coordinates refer to the static three-dimensional coordinates of the actual deployment position of the fixed intelligent mobile terminal in the storage area. During the system initialization phase, the actual spatial displacement of the physical center point of the equipment relative to the unified reference origin of the storage area is measured by the engineers using high-precision surveying tools, and this is hard-coded as a static system parameter into the configuration library of the cloud server. For the weighbridge node, the standard equipment identifier field is assigned the hardware number of the weighbridge node, the standard trigger type field is fixedly and forcibly assigned the value 2, and defined as a storage age warning event. The standard coordinate field is assigned the preset fixed physical coordinates of the weighbridge node, and the standard queuing field is fixedly and defaulted to 0.The received raw job data is parsed and normalized using the aforementioned relational mapping table. Specifically, after receiving raw job data in any network packet format, the cloud server extracts and parses the packet header to obtain the device type identifier. Using this device type identifier as an index, the server retrieves the packet payload from the relational mapping table and extracts the data from each private field in the packet payload: if the device type identifier represents a PDA device, the first and second private fields are extracted; if the device type identifier represents an access control gate device, the third and fourth private fields are extracted; if the device type identifier represents a weighbridge node, the fifth and sixth private fields are extracted. After extraction, the extracted private fields are aligned to the appropriate data types and then filled into the corresponding standard fields in the unified data pattern. During this field filling process, the cloud server extracts the local timestamp carried in the network packet as the local time reference, performs time zone conversion according to the Coordinated Universal Time standard, and fills it into the newly added standard timestamp field of the unified data model. Finally, it outputs a field-aligned structured record. The local timestamp refers to the time tag generated by each smart mobile terminal at the instant it senses and triggers the corresponding job event, based on the local clock maintained by the hardware clock chip or crystal oscillator module built into each smart mobile terminal, and used to digitally carry the physical occurrence time of the job event. When each smart mobile terminal reports the original job data to the cloud server through the standard communication interface, it encapsulates the timestamp as the original time feature payload in the network packet. The structured record refers to a structured data object allocated in the cloud server's memory that strictly conforms to the unified data model and has fixed data alignment boundaries. All structured records output after parsing and normalization are aggregated and arranged in the order of the unified timestamp, forming a spatial event stream in the input buffer of the cloud server.The spatial event stream refers to a structured record sequence containing multiple types of operational events from multi-source intelligent mobile terminals, which is dynamically updated over physical time. Each structured record consists of four fields concatenated sequentially: the first field is the device identifier, mapped from the standard device identifier field in the structured record, used to uniquely identify the core entity of the record. For access control gate devices and weighbridge nodes, it is the hardware node number; for PDA devices, it is mapped to the global electronic material code of the goods. The second field is the trigger type, mapped from the standard trigger type field in the structured record, used to characterize the enumeration identifier of the operational event category corresponding to the record, specifically including... The event stream includes: cargo scanning events (enumerated value 1), warehouse age warning events (enumerated value 2), and access control gate queuing events (enumerated value 3); the third field is the current physical coordinates, mapped from the standard coordinate field in the structured record, referring to the three-dimensional continuous spatial floating-point coordinates of the smart mobile terminal at the moment the event is triggered, based on the unified reference origin of the warehouse area; the fourth field is the queuing status identifier, mapped from the standard queuing field in the structured record, referring to an integer quantified value representing the current queuing congestion level at the location of the smart mobile terminal. For non-access control gate events, this queuing status identifier is fixed at 0 to indicate no queuing congestion. The spatial event stream achieves a complete record of the operational status of multi-source heterogeneous physical equipment under a unified time scale and unified data structure. See also. Figure 2 This is a schematic diagram of the physical architecture for multi-source heterogeneous terminal status data acquisition and event stream normalization processing provided in an embodiment of the present invention. Figure 2 The system is divided into three horizontal tracks from bottom to top: the physical device layer, the standard access layer, and the cloud computing layer. In the physical device layer, three types of rectangular blocks are scattered to represent PDA devices, access control gate devices, and weighbridge nodes, respectively. Arrows drawn upwards from each block, using different line types, represent the heterogeneous raw operational data reported by the three types of smart mobile terminals. The different line types visually represent the heterogeneous and disordered state of the raw operational data before access. In the standard access layer, a rectangular module running horizontally through the middle represents the standard communication interface. The nested grid-like blocks inside represent a preset relational mapping table. A single solid arrow pointing upwards from this module represents the structured record output after parsing and normalization. In the cloud computing layer, the horizontal strip at the top represents the spatial event stream arranged according to a unified timestamp in the cloud server's input buffer.

[0071] Step S12: Perform spatial dimension reduction discretization mapping based on the spatial event flow to generate a dynamic mesh state tensor.

[0072] After obtaining the spatial event stream containing structured records, to avoid high-dimensional state space expansion and floating-point matrix operation blockage caused by directly inputting the spatial event stream carrying three-dimensional coordinate values ​​into downstream calculations, the cloud server performs spatial dimensionality reduction discretization mapping based on the spatial event stream. Since each structured record in the spatial event stream has a fixed trigger type, current physical coordinates, and queuing status identifier, a discrete three-dimensional mesh index is constructed to directly quantify the inventory aging warning events and access control gate queuing events representing business anomalies in the structured records into the passage impedance values ​​of the bottom-level mesh nodes.

[0073] Specifically, the cloud server constructs a discrete three-dimensional mesh index in memory. This three-dimensional mesh index refers to a discrete spatial index structure composed of regularly arranged mesh nodes, obtained by dividing the storage area along the three-dimensional coordinate axes at predetermined intervals according to preset mesh side lengths. Each mesh node corresponds to a cubic sub-region with a fixed side length in the physical storage space. The three-dimensional coordinate axes include an X-axis established along the physical long side of the storage area, a Y-axis established along the physical wide side of the storage area, and a Z-axis established perpendicular to the ground height. The mesh side length refers to the straight-line distance between the centers of two adjacent mesh nodes, and its setting is based on the minimum safe turning diameter of a single intelligent mobile vehicle within the storage area, ensuring that the cubic sub-region corresponding to each mesh node can precisely accommodate a single vehicle to complete a stationary avoidance maneuver. For example, the grid side length is set to 1.2 meters. For the current physical coordinates of each structured record in the spatial event stream, the values ​​of the three-dimensional coordinates in the X, Y, and Z axes are divided by the grid side length and rounded down to obtain three integer index values. These three integer index values ​​are then concatenated sequentially to obtain the discrete grid coordinate code mapping the structured record. The cloud server assigns a dynamic feature vector to each grid node in the three-dimensional grid index. The dynamic feature vector is a numerical vector with a fixed length of three, and its three components are, in a fixed order, spatial occupancy bits, passage impedance value, and risk warning threshold. The space occupancy bit is a binary identifier representing whether the cubic sub-region corresponding to the grid node is currently occupied by a smart mobile terminal or goods. A value of zero indicates that the area is free, and a value of one indicates that it is occupied. The passage impedance value is a non-negative real number used to quantify the additional obstruction added to the grid node when a smart mobile terminal passes through it due to the presence of goods that have exceeded their storage age or gate queuing congestion. The risk warning threshold is a pre-set non-negative real constant representing that the grid node is in an impassable state. Its setting is based on the equivalent impedance limit corresponding to the upper limit of the tolerable physical probability of a smart mobile terminal causing a trajectory conflict at the grid node. For example, it is set to 8. When the passage impedance value of the grid node after being changed by an event is greater than or equal to the risk warning threshold, the cloud server directly marks the grid node as an unreachable obstacle node. Specifically, this is achieved by forcibly setting the space occupancy bit in its dynamic feature vector to one.

[0074] Spatial dimensionality reduction and discretization mapping are performed on the spatial event stream. Specifically, the cloud server traverses each structured record in the spatial event stream and updates the dynamic feature vector of the corresponding grid node according to its trigger type. Specifically: the cloud server calculates the corresponding discrete grid coordinate code based on the current physical coordinates of the current structured record, then extracts the trigger type of the structured record and performs mutually exclusive branch processing. Specifically, if the trigger type of the current structured record is a cargo scanning event, the spatial occupancy position in the dynamic feature vector of the corresponding discrete grid coordinate code grid node is set to one, and the passage impedance value remains unchanged; if the trigger type of the current structured record is a warehouse age warning event or an access control gate queuing event, the spatial occupancy position of the corresponding discrete grid coordinate code grid node is simultaneously set to one, and the passage impedance value of the grid node is incrementally superimposed. For grid nodes with trigger types of warehouse age warning events or access control gate queuing events, the calculation formula for the superimposed and updated passage impedance value is as follows: ,in, These represent the integer index values ​​of the discrete grid coordinate encoding along the X, Y, and Z axes, respectively, with value ranges of [missing values]. , , ,in, It belongs to the symbol. For a closed interval, , , The maximum value of the index range is obtained by dividing the maximum physical size of the storage area on the X-axis, Y-axis, and Z-axis by the grid side length and then rounding down. The discrete grid coordinates are encoded as follows: The passage impedance value is obtained by superimposing the impedance increments of the grid nodes; This represents the preset basic passage impedance value, which characterizes the inherent resistance constant of a grid node in the absence of any congestion or warning. Its setting is based on the normalized passage time of the physical distance corresponding to a single grid node. For example, it is set to 1.0. This indicates the current storage age of the goods. Its specific value is the time difference obtained by subtracting the goods' entry timestamp from the current system timestamp. The preset reservoir age impedance gain coefficient is a real multiplier used to linearly amplify the dimensionless reservoir age timeout ratio into the increment of the grid node's passage impedance value. It is set based on the additional obstruction applied to the grid node when the reservoir age timeout ratio reaches 100%. For example, it is set to 3. This indicates the preset warehouse age warning threshold, which is set based on the maximum physical dwell time of goods as specified in the warehouse management regulations. For example, it is set to 720 hours. This represents the preset queuing impedance gain coefficient, characterizing the strength ratio of the impedance transformation per unit number of queuing vehicles. Its setting is based on the equivalent congestion penalty value caused by a single queuing vehicle to adjacent passageways; for example, it is set to 2.5. A represents the integer quantized value extracted from the queuing status identifier field of the current structured record. After all structured records have been traversed, the cloud server constructs a four-dimensional numerical array as a dynamic grid state tensor. The first three dimensions of this four-dimensional numerical array correspond to the spatial arrangement dimensions of the three-dimensional grid index along the X, Y, and Z axes, respectively, using the discrete grid coordinate encoding i, j, and k to spatially address and locate specific grid nodes. The fourth dimension is fixed as a one-dimensional array of length 3, whose three storage bit indices are arranged in the order of features to store the spatial occupancy position, passage impedance value, and risk warning threshold of the grid node. See also... Figure 3 This is a schematic diagram illustrating the principle of dimensionality reduction discretization of warehouse space and construction of dynamic grid state tensors provided by an embodiment of the present invention. Figure 3 In the diagram, the left side constructs a three-dimensional coordinate system with the X, Y, and Z axes as spatial references. Several independent dots scattered within the system visually represent the current physical coordinates of each structured record in the spatial event flow. The middle of the diagram contains a horizontal mapping arrow pointing from left to right, representing the physical flow process of spatial dimensionality reduction and discretization mapping. The right side of the diagram shows the segmented discrete three-dimensional grid index structure. Different grayscale fills within the grid explicitly represent the dynamic feature vectors of the grid nodes numerically. Specifically: grid nodes marked "empty" indicate that their space occupancy is set to one; grid nodes marked "increased impedance" and "high impedance" indicate that they are in a state of superimposed passage impedance values ​​due to aging warning events or access control gate queuing events; grid nodes marked "obstacle node" indicate that the passage impedance value of the node has reached or exceeded the risk warning threshold and is judged as an unreachable obstacle; the remaining pure white grids represent an idle state with a passage impedance value of zero.

[0075] S10 addresses the issues of heterogeneous and disordered input states caused by the divergence of data formats and timing misalignment in multi-source heterogeneous intelligent mobile terminals, as well as the high-dimensional state space expansion and floating-point matrix operation blocking problems caused by optimization calculations in continuous spatial coordinate systems. Specifically, the spatial event stream converges isolated and disordered raw operational data from PDA devices, access control gate devices, and weighbridge nodes into a structured record sequence with a unified timestamp, constructing a strictly aligned structured data foundation. The dynamic grid state tensor converts inventory aging warning events and access control gate queuing events, representing business anomalies in the structured records, into the passage impedance values ​​of the underlying grid nodes, cutting off the high-dimensional expansion path caused by direct input of continuous three-dimensional coordinate values ​​to downstream calculations.

[0076] Step S20: Obtain the starting grid node and target grid node when the smart mobile terminal initiates the job request; call the dynamic grid state tensor to calculate the penalty coefficient for the starting grid node and target grid node to obtain the action value tensor; perform path backtracking based on the action value tensor to generate a grid path sequence.

[0077] Further, step S20 includes:

[0078] Step S21: Obtain the starting grid node and target grid node when the smart mobile terminal initiates the job request, call the dynamic grid state tensor to calculate the penalty coefficient for the starting grid node and target grid node, and obtain the action value tensor.

[0079] After obtaining the dynamic grid state tensor, in order to solve for feasible paths that avoid congested and high-frequency operation areas and avoid long computation cycles and processing blockages in continuous space, the cloud server calls the dynamic grid state tensor to perform penalty coefficient calculation. By establishing a penalty coefficient calculation mechanism based on a reinforcement learning model, the dynamic grid state tensor is mapped node by node to an action value tensor that can be directly read for path backtracking.

[0080] Specifically, the cloud server first extracts the starting grid node and target grid node when the smart mobile terminal initiates a job request. The starting grid node refers to the discrete grid coordinate code corresponding to the current physical coordinates of the smart mobile terminal after spatial dimension reduction and discretization mapping at the moment the job request is initiated; the target grid node refers to the discrete grid coordinate code corresponding to the physical coordinates of the target cargo location carried in the job request message after spatial dimension reduction and discretization mapping. The dynamic grid state tensor, the starting grid node, and the target grid node are jointly input into a pre-trained reinforcement learning model. The reinforcement learning model is a deep Q-value network model that takes the local state vector of the grid node as input and outputs the action value of the grid node in six discrete movement directions (up, down, left, right, forward, backward, and backward). The action value refers to the quantitative evaluation value of the expected cumulative benefit that the smart mobile terminal can obtain by moving along a specific discrete movement direction in the current state until it reaches the target grid node. The reinforcement learning model is modified from the standard deep Q-value network architecture in the public reinforcement learning framework and trained from scratch for this warehouse grid optimization scenario. The training data for the deep Q-value network model is collected from a warehouse digital twin simulation environment. In this environment, grid state samples containing space occupancy, passage impedance values, and risk warning thresholds are randomly generated according to the real distribution characteristics of the dynamic grid state tensor. A simulated intelligent agent—an automated computer program simulating the physical dimensions and kinematic characteristics of a real intelligent mobile terminal—repeatedly executes movement decisions within this environment, covering three typical operational scenarios: empty passage, congestion avoidance, and inventory age warning avoidance. The empty passage scenario refers to an environment with no space occupancy and a basic passage impedance value; the training objective is for the model to learn the shortest geometric path based on Manhattan distance. The congestion avoidance scenario involves randomly generating high-impedance grid nodes simulating queuing events at access control gates; the training objective is for the model to learn biased paths to avoid congested areas. The inventory age warning avoidance scenario involves randomly generating high-impedance grid nodes simulating goods exceeding their inventory age; the training objective is for the model to prioritize avoiding or responding to these warning node areas. A total of five million state transition data points were collected. Each state transition data point consists of four elements: current state, executed action, immediate reward, and next state. The current state refers to the eight-dimensional local state vector of the simulated agent at its current grid node; the executed action refers to the specific movement direction selected from six discrete movement directions; the immediate reward refers to the numerical reward / penalty signal fed back by the simulation environment according to preset reward rules after the simulated agent executes the action; and the next state refers to the eight-dimensional local state vector of the new grid node.The reward rule refers to a set of pre-set numerical mapping conditions in the simulation environment used to quantitatively evaluate the performance of the simulated agent's actions. Its construction is based on guiding the agent to safely reach the target area at the lowest cost. Specifically, the reward rule includes three mutually exclusive conditional decision branches: First, when the simulated agent performs a movement action that fails to reach the destination, a small base negative value is assigned, for example, -1, as a path length penalty to drive it to find the shortest path. Second, when the simulated agent enters a grid node with a single occupied space, or enters an obstacle node whose passage impedance exceeds the risk warning threshold, a very large negative value is assigned, for example, -100, as a violation penalty and to prematurely end the current training round, driving it to avoid obstacles and congestion. Third, when the simulated agent successfully reaches the target grid node, a very large positive value is assigned, for example, +100, as a task completion reward. The forward propagation process of the deep Q-value network model is as follows: The input layer receives a local state vector of a grid node with a dimension of eight. Its components are, in order, the current grid node's passage impedance value, the Manhattan distance from the current grid node to the target grid node, and the space occupancy of six adjacent grid nodes. In each path backtracking loop, the Manhattan distance is recalculated in real time based on the current grid node and the target grid node. Specifically, the absolute values ​​of the differences between the discrete grid coordinate codes of the current grid node and the target grid node on the X-axis, Y-axis, and Z-axis are calculated respectively, and the three absolute values ​​of the differences are algebraically added together. The input vector is multiplied by the learnable weight matrix of the first hidden layer and then superimposed with a learnable bias vector. After a nonlinear transformation using a linear rectified activation function, the first hidden feature vector is output. The first hidden feature vector is then multiplied by the learnable weight matrix of the second hidden layer and superimposed with a learnable bias vector. This is then further superimposed with a linear rectified activation function to output a second hidden feature vector. Finally, the second hidden feature vector is multiplied by the learnable weight matrix of the output layer and superimposed with a learnable bias vector to output a grid node action value vector. Its six components correspond to the action value of movement in six discrete directions: up, down, left, right, forward, and backward. The learnable weight matrices and learnable bias vectors of each layer are assigned random initial values ​​by the computer during model initialization and are iteratively updated based on the gradient of the loss function using the backpropagation algorithm during model training. The loss function of the deep Q-value network model adopts the temporal difference mean square error loss. Its loss value is calculated by the squared difference between the predicted action value in the current state and the target action value. The target action value is the expected value formed by the algebraic superposition of the maximum action value in the next state after the immediate reward is weighted by a preset decay discount factor.The attenuation discount factor is a real constant with a value range greater than zero and strictly less than one, representing the weight of the importance of long-term delayed gains relative to current immediate gains. Its setting is based on balancing the reinforcement learning model's focus on immediate local gains and future global optimal gains, while ensuring that the infinite cumulative value calculation in the Markov decision process is mathematically strictly convergent; for example, it is set to 0.95. Training uses an adaptive moment estimation optimization algorithm for gradient updates. Training is considered continuous convergence and terminated when the variance of the temporal difference mean squared error loss on the validation set is less than a preset convergence threshold for 100 consecutive training epochs. The convergence threshold is set based on the lower limit of the action value assessment error that a deep Q-value network model can tolerate under the current warehouse grid size, aiming to ensure that the model has reached a stable Nash equilibrium state and avoid premature stopping or overfitting; for example, it is set to 0.001.

[0081] The calculation of the penalty coefficient refers to the following: The cloud server traverses each grid node in the 3D mesh index where the spatial occupancy is zero. For each grid node, it extracts its local state vector and feeds it into the deep Q-value network model to obtain the original action value vector. Then, it superimposes the penalty coefficient derived from the passage impedance value onto each directional component of the original action value vector to obtain the corrected action value vector for that grid node. The formula for calculating the superimposed and corrected penalty coefficient is as follows: ,in, The discrete grid coordinates are encoded as follows: The penalty coefficient for mesh nodes is used to reduce the reachability of high-impedance mesh nodes at the action value level. The penalty coefficient for exceeding the limit represents the strong avoidance force applied after the passage impedance value exceeds the risk warning threshold. For example, it is set to 100. This represents the risk warning threshold in the dynamic feature vector of the grid node. Represents the maximum value function. The impedance gradient penalty coefficient characterizes the asymptotic suppression ratio of the passage impedance value on the motion value within the unbounded interval; for example, it is set to 0.5. The penalty coefficient is treated as an algebraic bias term and uniformly summed to the six directional components of the original motion value vector of the corresponding grid node. The resulting new six-dimensional numerical vector is defined as the corrected motion value vector of that grid node. The corrected motion value vectors obtained by calculating the penalty coefficients for all grid nodes with zero spatial occupancy bits in the three-dimensional grid index are organized and arranged according to the three-dimensional spatial arrangement order of the three-dimensional grid index, and defined as the motion value tensor. The motion value tensor refers to a four-dimensional numerical array with the first three dimensions corresponding to the grid spatial index and the fourth dimension corresponding to the six discrete movement directions.

[0082] Step S22: Based on the action value tensor, execute the path backtracking to generate a grid path sequence.

[0083] After obtaining the action value tensor, in order to convert it into an ordered path that can be directly executed by the smart mobile terminal, and to avoid forcing the computationally limited smart mobile terminal to bear the burden of path deduction locally by directly sending the action value tensor containing the global spatial state to the smart mobile terminal, the cloud server performs path backtracking based on the action value tensor. After obtaining the expected cumulative reward values ​​of each discrete movement direction contained in the action value tensor, a path backtracking mechanism is established. The purpose is to deterministically stitch together a discrete path with the lowest global passage impedance between the starting grid node and the target grid node, and output a grid path sequence composed of discrete grid coordinate codes.

[0084] Specifically, the cloud server first initializes an empty list of path nodes in memory and writes the discrete grid coordinates of the starting grid node as the first element into the list. At the same time, it sets a current grid node pointer in memory to indicate the grid node currently performing path exploration and initially points the current grid node to the starting grid node. Then, the cloud server enters a path backtracking loop. The loop object of the path backtracking loop is the pointer of the current grid node, which progresses step by step from the starting grid node to the target grid node. Within each loop, the cloud server extracts the corrected action value vector corresponding to the current grid node from the action value tensor based on the current grid node pointed to by the pointer. It then extracts the six component values ​​of the corrected action value vector corresponding to the six discrete movement directions (up, down, left, right, forward, backward) respectively. The discrete movement direction corresponding to the component with the largest cumulative gain value is selected. In the case of a tie due to equal values, the discrete movement direction that reduces the Manhattan distance between the current grid node and the target grid node the most is selected. The integer index value of the discrete grid coordinate code of the current grid node is offset by one grid unit along the discrete movement direction to obtain the discrete grid coordinate code of the next grid node. The discrete grid coordinate code of the next grid node is then appended sequentially to the path node list. Finally, the pointer of the current grid node is updated to point to the next grid node. The termination condition of this path backtracking loop adopts an exhaustive and mutually exclusive decision branch: First, if the discrete grid coordinate code of the grid node pointed to by the current grid node pointer is completely equal to the discrete grid coordinate code of the target grid node, the path backtracking is determined to be successful, and the cloud server terminates the path backtracking loop; Second, if the total number of discrete grid coordinate codes contained in the path node list reaches the preset backtracking step threshold, and the discrete grid coordinate code of the grid node pointed to by the current grid node pointer is still not equal to the discrete grid coordinate code of the target grid node, the path is determined to be unreachable, the cloud server terminates the path backtracking loop, and reports an error to end the path planning for this job request of the smart mobile terminal; Third, if the discrete grid coordinate code of the next grid node corresponding to the selected maximum component already exists in the current path node list, thus forming a spatial loop, the local deadlock is determined to be, the cloud server terminates the path backtracking loop, and reports an error to end the path planning for this job request of the smart mobile terminal. The backtracking step threshold is set based on twice the sum of the maximum physical index sizes of the 3D mesh index in the X, Y, and Z dimensions. This aims to provide a deterministic computational termination boundary for path backtracking to prevent infinite loops; for example, it is set to 512. The discrete mesh coordinates arranged in the writing order in the path node list after successful path backtracking are encoded and defined as a mesh path sequence.The grid path sequence refers to an ordered list of discrete grid coordinate codes, arranged in chronological order from the starting grid node to the target grid node. Each element is a discrete grid coordinate code composed of three integer indices, collectively representing the grid node trajectory that the intelligent mobile terminal should traverse from its starting position to the target location. See also. Figure 4 This is a schematic diagram illustrating the deductive principle of greedy path backtracking based on the action value distribution of grid nodes, as provided in this embodiment. Figure 4 In the diagram, the central square represents the current grid node in the path backtracking loop. The six arrowed value boxes radiating outwards from this current grid node represent the corrected action values ​​extracted from the grid node's action value distribution for the six discrete movement directions: up, down, left, right, forward, and backward. Specific values ​​in the diagram, such as -1.0, -5.5, and -22.1, are merely illustrative examples, intended to represent the quantitative assessment of the expected cumulative return of the current grid node moving in the corresponding direction. The "MAX" label in the diagram indicates that a rigorous comparison is made among the corrected action values ​​in the six extracted directions, and the discrete movement direction corresponding to the largest corrected action value is selected. The cloud server then offsets the discrete grid coordinates of the current grid node by one grid unit along this selected discrete movement direction to confirm it as the next grid node, thereby driving the path backtracking algorithm to complete deterministic extreme value optimization and trajectory advancement on a low-dimensional tensor base.

[0085] Step S20 addresses the long computation cycles and processing blockages caused by global traversal or heuristic search algorithms in continuous space, as well as trajectory conflicts and local deadlock anomalies caused by the inability to perceive service congestion during concurrent execution by multiple terminals, through the starting grid node, target grid node, dynamic grid state tensor, action value tensor, and grid path sequence. Specifically, the action value tensor maps the dynamic grid state tensor to the expected cumulative benefit value that incorporates the accessibility penalty, thus reducing the reachability of high-impedance grid nodes at the action value level. The grid path sequence, encoded with pure integer discrete grid coordinates, deterministically stitches together the discrete trajectory with the lowest global accessibility between the starting and target grid nodes, constructing a low-dimensional data-driven foundation for execution by intelligent mobile terminals.

[0086] Step S30: Perform lossless compression encoding on the grid path sequence to generate a lightweight instruction package, and send the lightweight instruction package to the target smart mobile terminal to execute an adaptive job loop. Detection is performed on each loop of the adaptive job loop, and obstacle detection results are output. Branch processing and position updates are performed on the obstacle detection results to obtain the local update results of the dynamic grid state tensor.

[0087] Further, step S30 includes:

[0088] Step S31: Perform lossless compression encoding on the grid path sequence to generate a lightweight instruction package, and send the lightweight instruction package to the target smart mobile terminal to execute the adaptive job cycle.

[0089] After obtaining the grid path sequence, to accurately distribute the grid path sequence with minimal communication overhead and drive the smart mobile terminal to execute it, and to avoid the cloud server directly distributing the grid path sequence as an uncompressed list of plaintext integers, which would increase the communication link load between the cloud server and the smart mobile terminal in high-concurrency scenarios, the cloud server performs lossless compression encoding on the grid path sequence. After obtaining the ordered list of discrete grid coordinate codes contained in the grid path sequence, the cloud server converts the grid path sequence into a lightweight instruction packet to further compress the transmission volume by utilizing the strong spatial correlation between the discrete grid coordinate codes of adjacent grid nodes. This lightweight instruction packet is then distributed to the target smart mobile terminal, which performs the gridded adaptive operation.

[0090] Specifically, the cloud service performs a differential pre-transformation on the grid path sequence. For each discrete grid coordinate code in the grid path sequence, arranged in the order of movement starting from the second discrete grid coordinate code, each discrete grid coordinate code is subtracted from the preceding discrete grid coordinate code component by component to obtain the coordinate increment of adjacent grid nodes. All the calculated coordinate increments in the grid path sequence are arranged in the order of movement to form a differential coordinate sequence. Since adjacent grid nodes in the grid path sequence are adjacent along one of the six discrete movement directions (up, down, left, right, forward, backward), the value of each coordinate increment component in the differential coordinate sequence is strictly constrained to the set of three values: negative one, zero, and positive one. Subsequently, the cloud server invokes the run-length encoding algorithm to perform lossless compression on the differential coordinate sequence. It replaces identical and continuously repeating coordinate increments in the differential coordinate sequence with a tuple of the increment and its number of consecutive repetitions. Non-repeating coordinate increments are retained, resulting in a compressed bitstream. Since straight-line travel segments appear as long strings of repeating increments in the differential coordinate sequence, the run-length encoding algorithm can achieve deterministic, high-ratio lossless compression of these long strings of repeating increments. The cloud server encapsulates the compressed bitstream along with the discrete grid coordinate codes of the starting grid node and the device identifier of the target smart mobile terminal to obtain a lightweight instruction packet. The lightweight instruction packet is a structured communication message with a fixed-length header and a variable-length compressed body. The header of the lightweight instruction packet encapsulates the device identifier of the target smart mobile terminal and the discrete grid coordinate codes of the starting and target grid nodes, while the compressed body encapsulates the compressed bitstream. The cloud server sends the lightweight instruction packet to the target smart mobile terminal pointed to by the device identifier through a standard communication interface. The target smart mobile terminal decodes and restores the lightweight instruction packet in its local memory. The aforementioned decoding and restoration refers to the process by which the target intelligent mobile terminal, based on the discrete grid coordinate encoding of the starting grid node in the header of the lightweight instruction packet, first performs an inverse run-length encoding transformation on the compressed bitstream in the compressed body of the lightweight instruction packet to restore the differential coordinate sequence, and then, using the discrete grid coordinate encoding of the starting grid node as the base point, performs integer accumulation on each coordinate increment in the differential coordinate sequence element by element to gradually restore the complete grid path sequence.The target intelligent mobile terminal then enters an adaptive task loop. This loop iterates through the discrete grid coordinate codes within the restored complete grid path sequence. Within each loop, the target intelligent mobile terminal reverse-calculates the current discrete grid coordinate code into the physical coordinates of the corresponding grid node center and moves to those coordinates. Upon arrival, the target intelligent mobile terminal executes a local data transmission command to report its arrival status to the cloud server via a standard communication interface. The adaptive task loop terminates when all discrete grid coordinate codes in the complete grid path sequence have been executed sequentially, meaning the target intelligent mobile terminal has reached the target grid node. See also. Figure 5 This is a schematic diagram of the timing and data structure for lossless compression and delivery of a grid path sequence and its decoding execution at the terminal, provided by an embodiment of the present invention. Figure 5 In the diagram, the upper part shows the data processing chain in the cloud. A differential pre-transformation is performed on the generated grid path sequence through backtracking, utilizing the strong spatial correlation between adjacent grids to generate a differential coordinate sequence. Subsequently, this sequence undergoes lossless compression via run-length encoding to output a compressed bitstream, which is then encapsulated into a lightweight instruction packet. This lightweight instruction packet is transmitted downwards through a standard communication interface, crossing the interaction boundary to enter the local execution area of ​​the mobile terminal below. Inside the smart mobile terminal, the local memory decodes and reconstructs the received lightweight instruction packet to rebuild the complete grid path sequence, driving the terminal into an adaptive work loop, moving node by node along the grid trajectory. During loop execution, if the terminal detects a temporary obstacle at the current grid node using an obstacle perception sensor, it immediately triggers a mutual exclusion branch, generating an abnormal state increment locally and transmitting it back incrementally via the standard communication interface. Upon receiving this abnormal increment, the cloud server accurately locates the affected node and updates its memory only for the affected node's impedance value, thus quickly generating a local update result, forming an efficient data flow and path replanning closed-loop feedback structure.

[0091] Step S32: Perform detection on each iteration of the adaptive job loop, output obstacle detection results, perform branching and setting updates on the obstacle detection results, and obtain the local update results of the dynamic mesh state tensor.

[0092] The target intelligent mobile terminal generates an abnormal state increment based on the obstacle detection results and sends the increment back. The cloud server then generates a local update result of the dynamic grid state tensor to regenerate the grid path sequence.

[0093] During the adaptive job cycle executed by the target intelligent mobile terminal, to ensure that the dynamic grid state tensor maintained by the cloud server remains consistent with the actual physical state of the storage area, and to avoid dual latency in communication and computation caused by the target intelligent mobile terminal transmitting the entire dynamic grid state tensor back to the cloud server after detecting a temporary obstacle, the target intelligent mobile terminal performs incremental backhaul of local abnormal grids. As the target intelligent mobile terminal advances along the grid path sequence, to synchronize only the information of grid nodes containing temporary obstacles to the cloud server at minimal cost, a local update of the dynamic grid state tensor is triggered instead of a global reconstruction.

[0094] Specifically, within each loop of the adaptive operation cycle, the target intelligent mobile terminal uses its onboard obstacle perception sensor to detect whether there are temporary obstacles not recorded by the dynamic grid state tensor within the cubic sub-region of the grid node corresponding to the current discrete grid coordinate encoding, and outputs the corresponding obstacle detection result. The temporary obstacle refers to a physical obstruction formed by temporarily dropped goods or temporarily stopped vehicles during warehousing operations, rendering the current grid node practically impassable. The target intelligent mobile terminal performs exhaustive and mutually exclusive branching based on the obstacle detection result. The obstacle detection result includes both no temporary obstacle detected and a temporary obstacle detected; the target intelligent mobile terminal performs exhaustive and mutually exclusive branching based on the detection result: if the obstacle detection result indicates that no temporary obstacle was detected at the current grid node, the target intelligent mobile terminal continues to advance according to the original grid path sequence without triggering any backtracking operation; if the obstacle detection result indicates that a temporary obstacle was detected at the current grid node, the target intelligent mobile terminal only generates an abnormal state increment. The abnormal state increment refers to a tuple composed of the discrete grid coordinate code of the abnormal grid node and the abnormal impedance setting value. The abnormal grid node is the current grid node that has detected a temporary obstacle. The abnormal impedance setting value is a numerical variable generated locally by the target smart mobile terminal, used to directly replace the original access impedance value of the abnormal grid node in the cloud server. The abnormal impedance setting value is set to be much greater than the risk warning threshold to ensure that the node is suppressed to extremely low accessibility in the penalty coefficient calculation; for example, it is set to 900.

[0095] The target intelligent mobile terminal transmits the abnormal state increment back to the cloud server incrementally via a standard communication interface. Upon receiving the abnormal state increment, the cloud server locates the corresponding grid node in the dynamic grid state tensor based on the discrete grid coordinate encoding carried in the increment. It only updates the pass-through impedance value in the dynamic feature vector of that grid node, directly replacing it with the high-impedance constant from the abnormal state increment. Simultaneously, it maintains all dynamic feature vectors of the remaining grid nodes in the dynamic grid state tensor unchanged, thus generating a local update result for the dynamic grid state tensor. This local update result refers to the updated dynamic grid state tensor where only the pass-through impedance value at the abnormal grid node is rewritten, while the remaining global dynamic feature vectors retain their original values. After obtaining the local update result of the dynamic grid state tensor, the cloud server, for the target intelligent mobile terminal affected by the temporary obstacle, re-extracts the starting grid node and the target grid node based on the updated dynamic grid state tensor and recalculates the corresponding action value tensor and execution path backtracking. The purpose is to regenerate a grid path sequence that avoids the abnormal grid node, thereby forming a closed-loop feedback structure for data flow.

[0096] Step S30 addresses the high communication link load caused by the cloud server directly sending an uncompressed plaintext integer list, and the dual latency issues of communication and computation caused by the target smart mobile terminal sending back the dynamic grid state tensor in full after detecting a temporary obstacle, by using lightweight instruction packets, abnormal state increments, and local update results of the dynamic grid state tensor. Specifically, the lightweight instruction packet converts the grid path sequence into a lossless compressed bitstream through differential pre-transformation and run-length encoding, significantly reducing the transmission volume sent to the target smart mobile terminal. The abnormal state increments and local update results of the dynamic grid state tensor, through a binary incremental back-transmission mechanism, directly replace and update the passage impedance value only for abnormal grid nodes, maintaining real-time consistency between the dynamic grid state tensor and the physical reality with minimal communication cost, and triggering a closed-loop feedback structure to regenerate the grid path sequence.

[0097] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.

[0098] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely preferred embodiments of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A grid-based AI path planning method for warehouse space for intelligent mobile terminals, characterized in that, The method includes: Obtain the raw job data, construct a relational mapping table, perform parsing normalization on the raw job data, generate a spatial event stream, perform spatial dimension reduction discretization mapping based on the spatial event stream, and generate a dynamic grid state tensor. Obtain the starting grid node and target grid node when the intelligent mobile terminal initiates a job request, call the dynamic grid state tensor to calculate the penalty coefficient of the starting grid node and target grid node to obtain the action value tensor, and perform path backtracking based on the action value tensor to generate a grid path sequence; Lossless compression encoding is performed on the grid path sequence to generate a lightweight instruction package, which is then sent to the target smart mobile terminal to execute an adaptive job loop. Each loop of the adaptive job loop is used for detection, and the obstacle detection results are output. The obstacle detection results are then processed by branching and setting updates to obtain the local update results of the dynamic grid state tensor.

2. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 1, characterized in that, The method for obtaining the spatial event stream includes: Parse the network packets encapsulated from the original job data, obtain the device type identifier in the packet header, and extract the local timestamp carried by the network packet; Using the device type identifier as an index, the private fields in the corresponding message payload are extracted from the relational mapping table. The private fields are then filled into the standard fields of the preset unified data pattern, and the local timestamp is filled into the standard timestamp field of the unified data pattern to generate the corresponding structured record. Multiple structured records generated for each network message are aggregated and arranged to generate a spatial event stream.

3. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 2, characterized in that, The method for obtaining the structured record includes: The unified data schema includes a standard device identifier field, a standard trigger type field, a standard coordinate field, and a standard queuing field; The private fields are mapped and filled into the standard device identifier field, standard trigger type field, standard coordinate field, and standard queue field, respectively. Convert the local timestamp to a time zone and populate it into the standard timestamp field to output a structured record.

4. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 3, characterized in that, The method for obtaining the dynamic mesh state tensor includes: Within the storage space, the grid is divided along three preset orthogonal axes according to preset grid side lengths to construct a three-dimensional grid index composed of multiple grid nodes. Extract the standard coordinate field from each structured record in the spatial event stream as the current physical coordinate. Divide the value of the current physical coordinate in the three orthogonal axis directions by the grid side length and round down to obtain three integer index values, and then concatenate them to form the discrete grid coordinate code corresponding to the current physical coordinate. Each grid node is assigned a dynamic feature vector, which includes space occupancy bits, passage impedance value, and risk warning threshold. By aggregating the dynamic feature vectors based on the coordinate codes of each discrete grid, a dynamic grid state tensor is generated.

5. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 4, characterized in that, After assigning dynamic feature vectors to each grid node, the process further includes: Parse the standard trigger type field in the structured record and extract the event category represented by the standard trigger type field. The event category includes cargo scanning event, warehouse age warning event and access control gate queuing event. If the event category is a cargo scanning event, the space occupancy bit of the corresponding discrete grid coordinate encoding associated grid node is set to 1. If the event category is a warehouse age warning event or an access control gate queuing event, the space occupancy bit of the corresponding grid node is set to 1, and incremental superposition calculation is performed on the passage impedance value of the grid node.

6. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 5, characterized in that, The method for obtaining the action value tensor includes: Traverse each grid node in the 3D grid index that has zero spatial occupancy, extract the local state vector of each grid node and input it into the pre-trained reinforcement learning model, and output the original action value vector corresponding to each grid node. The penalty coefficient of each grid node is calculated using the starting grid node, the target grid node, and the passage impedance value. The penalty coefficient is then superimposed on the corresponding original action value vector to generate the corrected action value vector for each grid node. Combine the modified action value vectors corresponding to all grid nodes to obtain the action value tensor.

7. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 6, characterized in that, The method for calculating the value vector of the corrected action includes: Calculate the difference between the passage impedance value of the corresponding grid node and the risk warning threshold, extract the maximum value between the difference and zero using the maximum value function, and multiply it by the preset over-limit penalty coefficient to obtain the over-limit penalty item; The gradient penalty term is obtained by multiplying the impedance value by a preset impedance gradient penalty coefficient. The penalty coefficient is obtained by algebraically summing the over-limit penalty term and the gradient penalty term and taking the negative value. The penalty coefficient is then added to the directional component of the corresponding original action value vector as an algebraic bias term to obtain the corrected action value vector.

8. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 6, characterized in that, The method for obtaining the grid path sequence includes: Create a path node list and write the discrete grid coordinate code corresponding to the starting grid node as the first element; Set the current grid node pointer to initially point to the starting grid node, triggering a path backtracking loop; In the path backtracking loop, extract the corrected action value vector corresponding to the grid node pointed to by the current grid node pointer in the action value tensor, and obtain the component values ​​of multiple discrete movement directions in the corrected action value vector. Select the discrete movement direction corresponding to the largest component value for coordinate offset, generate the discrete grid coordinate code of the next grid node and append it to the path node list, and update the current grid node pointer to point to the next grid node; The loop terminates when the current grid node pointer points to the target grid node, and the grid path sequence is output.

9. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 8, characterized in that, The method for obtaining the lightweight instruction package includes: For each adjacent discrete grid coordinate code in the grid path sequence, perform a component-wise integer subtraction operation with the previous discrete grid coordinate code to obtain a differential coordinate sequence composed of multiple coordinate increments; Replace consecutively repeated identical coordinate increments in the differential coordinate sequence with a tuple containing the coordinate increment and the number of consecutive repetitions to generate a compressed bitstream; Obtain the device identifier of the target smart mobile terminal; The compressed bitstream, the discrete grid coordinate encoding corresponding to the starting grid node, and the device identifier of the target smart mobile terminal are encapsulated to generate a lightweight instruction package.

10. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 4, characterized in that, The obstacle detection results include: Obtain the obstacle detection results output by the target intelligent mobile terminal in the adaptive job cycle, the obstacle detection results including the detection of temporary obstacles and the absence of temporary obstacles; If the obstacle detection result is that a temporary obstacle is detected, the abnormal state increment is received from the target smart mobile terminal. The abnormal state increment is composed of the discrete grid coordinate code of the abnormal grid node and the abnormal impedance setting value. The abnormal grid node is the current grid node that has detected a temporary obstacle, and the abnormal impedance setting value is a constant with a value greater than the risk warning threshold.

11. The warehouse space grid-based AI path planning method for intelligent mobile terminals according to claim 10, characterized in that, The method for obtaining the local update result includes: Based on the discrete grid coordinate encoding of the abnormal grid nodes, the corresponding abnormal grid nodes are located in the dynamic grid state tensor. Replace the pass-through impedance value of the abnormal grid node directly with the abnormal impedance setting value; Maintaining the dynamic eigenvectors of the remaining grid nodes in the dynamic grid state tensor unchanged, the local update result of the dynamic grid state tensor is generated; The method for obtaining the action value tensor and the method for obtaining the grid path sequence are retried using the results of local updates, so as to regenerate the grid path sequence.