An intelligent decision system for commodity circulation based on end-cloud cooperation
By constructing an edge-cloud collaborative intelligent decision-making system, and combining a lightweight image segmentation model and a time series prediction algorithm, the system addresses the shortcomings of existing systems in data acquisition and decision-making capabilities, achieving efficient and accurate intelligent warehouse management and improving the system's real-time performance and flexible scheduling capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN RUIDA INFORMATION TECH CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-05
AI Technical Summary
Existing commodity circulation management systems are insufficient to meet the real-time, accuracy, and flexible scheduling requirements of modern logistics operations in terms of data collection efficiency, on-site status perception capabilities, and intelligent prediction and optimization decision-making capabilities. In particular, they lack structured, real-time spatial perception capabilities and dynamic task generation mechanisms in scenarios involving concurrent operations in multiple regions or autonomous equipment scheduling.
We construct an intelligent decision-making system based on edge-cloud collaboration, combining a lightweight image segmentation model, time series prediction algorithm and edge computing. Through image acquisition, feature extraction, state perception, time series analysis and intelligent scheduling modules, we achieve high-precision recognition, real-time data processing and intelligent decision-making.
It significantly improves the automation level and operational efficiency of smart warehousing, possesses high-precision identification capabilities, real-time data processing, and intelligent scheduling decision-making capabilities, and can accurately predict cargo demand, inventory changes, and operational load, thereby optimizing scheduling decisions.
Smart Images

Figure CN122155600A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, computer vision, edge computing and intelligent warehouse management, and in particular to an intelligent decision-making system for commodity circulation based on edge-cloud collaboration. Background Technology
[0002] Existing commodity distribution management systems primarily rely on traditional Warehouse Management Systems (WMS) and Enterprise Resource Planning (ERP) systems for operation recording and process scheduling. These systems mostly employ rule-driven control logic and manually configured task parameters, which have significant limitations when dealing with complex and ever-changing warehouse environments. As modern logistics operations increasingly demand real-time performance, accuracy, and flexible scheduling, traditional systems are finding it increasingly difficult to meet practical needs in terms of data acquisition efficiency, on-site status awareness, and intelligent prediction and optimization decision-making capabilities.
[0003] In terms of on-site data collection, traditional systems rely on manual input, barcode scanning devices, and basic video surveillance to record operational status. They cannot accurately identify the distribution of goods within shelves, storage locations, aisles, and work areas, especially in scenarios involving concurrent operations in multiple areas or autonomous equipment scheduling, lacking structured, real-time spatial awareness capabilities. While some existing vision-based solutions incorporate object detection and image segmentation technologies, most employ static models or single-scale processing methods, making it difficult to adapt to the spatial distribution characteristics of different work areas. Furthermore, the lack of a prompting and guidance mechanism results in weak model versatility and low segmentation accuracy.
[0004] In data analysis and forecasting, traditional methods are mostly based on statistical rules or historical average models, failing to fully utilize the inherent correlations between time-series characteristics and multi-source heterogeneous data. This makes them ineffective in predicting future fluctuations in demand, inventory trends, and operational load dynamics. While some advanced systems incorporate time-series models, they often neglect the collaborative computing architecture between the edge and cloud, leading to lags in on-site data processing and limited model predictive capabilities. Furthermore, scheduling decisions still rely heavily on manually set rules, lacking dynamic task generation, path optimization, and priority control mechanisms based on forecast results, thus failing to achieve a truly intelligent warehousing decision-making closed loop.
[0005] Therefore, how to provide an intelligent decision-making system for commodity circulation based on edge-cloud collaboration is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose an intelligent decision-making system for commodity circulation based on edge-cloud collaboration. This invention fully integrates a lightweight image segmentation model, time series prediction algorithms, and edge computing and cloud computing collaborative mechanisms to construct an end-to-end processing flow consisting of modules such as image acquisition, feature extraction, prompting control, state perception, time series analysis, and intelligent scheduling. It details methods for achieving cargo state identification, inventory prediction, and scheduling decision optimization in warehousing scenarios. This system possesses advantages such as high recognition accuracy, real-time data processing, strong predictive capabilities, and a high degree of intelligent decision-making and scheduling, significantly improving the automation level and operational efficiency of intelligent warehousing.
[0007] According to an embodiment of the present invention, a smart decision-making system for commodity circulation based on edge-cloud collaboration includes: The image acquisition and preprocessing module is used to acquire image data and business data from edge computing terminals deployed at the warehouse site, preprocess the image data, and obtain image sequences of the warehouse scene. The business data acquisition module is used to collect inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, and task execution logs of AGV vehicles. The edge feature extraction module is used to input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, the encoding sub-network with EfficientViT as the image encoder is called to extract features from the image sequence and obtain multi-scale feature maps of the warehousing scene. The improved MobileSAM model includes an encoding subnetwork, a semantic preservation module, and a mask decoder; The semantic preservation module is used to perform shape normalization, position indexing, and context labeling on the multi-scale feature maps, organizing the multi-scale feature maps into semantically preserved feature maps. The cue control module is used to process point cues and box cues and associate them with region labels within the structure alignment unit, schedule the set of structure alignment cues within the region scheduling unit, generate a cue scheduling sequence, and pass the cue scheduling sequence and semantic preservation feature map to the mask decoder. The edge-side scene analysis module is used to calculate warehouse site status perception data based on the segmented mask image after the mask decoder outputs the segmented mask image; The edge-side time-series construction module is used to perform time matching and field alignment between warehouse site status perception data and business data within the edge computing terminal, generating multi-dimensional commodity time-series data for warehouse operations; The cloud-based time series forecasting module is used to receive multi-dimensional commodity time series data from warehousing operations, construct multi-dimensional time series data, input the multi-dimensional time series data into the PatchTST time series forecasting model, and generate forecast results. The cloud-based scheduling and decision-making module is used to generate intelligent warehouse scheduling decision-making result data structures based on prediction results, and then distribute the intelligent warehouse scheduling decision-making results to edge computing terminals deployed on-site in the warehouse for execution.
[0008] Optionally, modules can be integrated using the following methods: Step 1: Image data and business data are collected by edge computing terminals deployed at the warehouse site. The image data is preprocessed to obtain an image sequence of the warehouse scene. Step 2: Input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, call the encoding sub-network with EfficientViT as the image encoder to extract features from the image sequence and obtain multi-scale feature maps of the warehouse scene. Step 3: Input the multi-scale feature map into the semantic preservation module. The multi-scale feature map is structurally organized through shape normalization, position indexing, and context labeling operations to obtain the semantic preservation feature map. Step 4: The semantic preservation feature map is passed to the prompt control module. The point prompts and box prompts set for the warehouse operation area are processed and associated with the area labels according to the preset structural alignment rules. The area scheduling unit generates the corresponding prompt scheduling sequence and passes the prompt scheduling sequence and the semantic preservation feature map to the mask decoder in the preset calling order. The mask decoder outputs the segmentation mask image and calculates the warehouse site status perception data based on the segmentation mask image.
[0009] Step 5: Package the warehouse site status perception data and business data to generate multi-dimensional commodity time-series data of warehouse operations, and transmit the multi-dimensional commodity time-series data of warehouse operations to the cloud time series prediction module through the network by the edge computing terminal; Step 6: Construct a multidimensional time series in the cloud based on the multidimensional commodity time series data of warehousing operations, and input the multidimensional time series into the PatchTST time series prediction model to obtain the prediction results for each shelf, storage location and operation area within the preset time window; Step 7: Generate intelligent warehouse scheduling decision results in the cloud based on the forecast results of goods demand, inventory change and task load. These results include replenishment task allocation, AGV vehicle travel path planning, sorting operation sequence and warehouse area scheduling priority. The intelligent warehouse scheduling decision results are then distributed to the edge computing terminals deployed on the warehouse site for execution via the network.
[0010] Optionally, the image data comes from the warehouse racking area, aisle area, and work area; the business data includes inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, AGV task execution logs, and corresponding timestamps; preprocessing includes size normalization, encoding format sorting, and time order sorting. Optionally, step two specifically includes: The image sequence is divided into N image subsequences according to a preset number of frames; The images in the image subsequence are input frame by frame into the EfficientViT image encoder in the coding subnetwork of the improved MobileSAM model. Convolution operation and feature embedding processing are performed on each frame image to obtain the corresponding first-scale feature map. Within the EfficientViT image encoder, a multi-layer feature extraction structure is used to downsample and channel map the first-scale feature map to generate two intermediate feature maps with different resolutions. The intermediate feature maps are arranged in a preset scale order and channel aligned to form a multi-scale feature map representing the spatial hierarchy information of the warehousing scene.
[0011] Optionally, step three specifically includes: The multi-scale feature map of the warehousing scenario is passed as input to the input end of the semantic preservation module. The spatial size of the multi-scale feature map is unified, and the height and width of the multi-scale feature map are adjusted to the preset standard size. Multi-scale feature maps that do not meet the standard size are filled to obtain the size-standardized feature map. Based on the size-normalized feature map, a corresponding location information index is generated for each spatial location; The location information indexes are arranged according to the row and column order of the size-normalized feature map, forming a structure corresponding to the size-normalized feature map. Figure 1 A corresponding location information matrix; By associating the location information matrix with the size-normalized feature map, a feature map containing the location index sorting results is obtained; Introduce contextual tags in the channel dimension of the feature map containing the location index sorting results to indicate storage rack areas, aisle areas, and work areas; Context labels are appended to the corresponding feature channels using a preset encoding method to form a multi-scale feature map with context labels; The multi-scale feature maps with context labels are combined in the semantic preservation module according to the preset scale order and channel order, and the output is used as the semantic preservation feature map.
[0012] Optionally, step four specifically includes: The semantically preserved feature map is passed to the input of the prompt control module, which receives point prompts and box prompts set for the warehouse operation area. The prompt control module includes an input terminal, a structure alignment unit, and a region scheduling unit; The warehousing operation area includes a shelving area, a storage location area, and a work area; Within the structural alignment unit, point hints and box hints are aligned sequentially according to preset structural alignment rules. Point hints and box hints belonging to the same timestamp are grouped together to form a hint sequence that corresponds to the time sequence of the warehousing scenario. Within the structure alignment unit, spatial indices are assigned to point and box prompts in the prompt sequence based on the spatial dimensions of the semantically preserved feature map. The spatial indices are then mapped one-to-one with the prompt positions in the prompt sequence to generate a structure alignment prompt set containing spatial index information. Within the structural alignment unit, the area labels representing the storage operation area are associated with the point tips and box tips in the structural alignment tip set, respectively. An area label mark consistent with the storage operation area to which the point tip belongs is attached to each point tip and each box tip, forming a structural alignment tip set with area label marks. The set of structure alignment hints marked with area labels is input into the area scheduling unit. Within the area scheduling unit, the set of structure alignment hints is scheduled according to the label order of the storage operation area to generate a hint scheduling sequence corresponding to the area label order. The prompt scheduling sequence and semantically preserved feature map are passed together to the mask decoder in a preset calling order, and the mask decoder outputs the segmentation mask image; The segmented mask image includes the outlines of pallets, containers, storage locations, and obstacles; Based on segmented mask images, the quantity of goods, occupied area, and spatial relationship of goods corresponding to shelves, storage locations, and temporary storage areas are calculated to form warehouse site status perception data.
[0013] Optionally, step five specifically includes: Within the edge computing terminal, based on the timestamp information contained in the warehouse site status perception data and business data, the warehouse site status perception data and business data are time-matched, and warehouse site status perception data records with the same timestamp are associated with business data records to form a set of warehouse operation data records arranged in chronological order. Based on the shelf identifiers, storage location identifiers, and work area identifiers contained in the warehousing operation data record set, the warehousing operation data records are grouped together, and warehousing operation data records belonging to the same shelf, the same storage location, or the same work area are grouped together. Within each group, the quantity of goods, area occupied, and spatial location relationships in the warehouse site status perception data are aligned with the inbound and outbound records, goods identification information, and task execution logs in the business data, and arranged in time order to construct a multi-dimensional commodity time-series data record sequence corresponding to the shelves, storage locations, and work areas. The multidimensional commodity time-series data records of each group are appended with warehousing operation area labels, equipment identification information and time range identifiers to form warehousing operation multidimensional commodity time-series data, and the warehousing operation multidimensional commodity time-series data is encapsulated into a data message to be transmitted; The edge computing terminal sends the data packets to be transmitted to the receiving end of the cloud time series prediction module through a preset network communication interface.
[0014] Optionally, step six specifically includes: Based on the timestamps contained in the multidimensional commodity time-series data of warehousing operations, the multidimensional commodity time-series data of warehousing operations is arranged into a continuous data sequence in chronological order; Based on the field parsing of the data sequence, the fields of goods quantity, inventory level, occupied area, spatial location, inbound / outbound and task execution are extracted from the data sequence in sequence according to the preset field structure, and together they form a multidimensional time series. Within the cloud-based time series forecasting module, a preset time window length and forecast time span are set for each multidimensional time series. Each multidimensional time series is then divided according to the time window to obtain a time window sequence consisting of M consecutive time windows. The input feature fields within each time window are organized into a model input sequence, and the corresponding goods quantity field, inventory level field, and task execution quantity field within the target time range after the time window are organized into a prediction target sequence. Input the model input sequences corresponding to each shelf, storage location, and work area into the PatchTST time series prediction model; In the PatchTST time series forecasting model, the model input sequences are encoded and time-dependent models are modeled, and the model output sequences are output in a one-to-one correspondence with each model input sequence. The model output sequence is restored and mapped to the time axis corresponding to each shelf, storage location and work area to obtain the prediction results for each shelf, storage location and work area within the preset time window. The forecast results include forecasts of demand for goods, inventory changes, and workload.
[0015] Optionally, step seven specifically includes: The cloud server where the cloud-based time series forecasting module is located receives the forecast results and classifies them according to shelf identifiers, storage location identifiers, and work area identifiers to form a set of forecast results. Based on the inventory change forecast and the goods demand forecast, a replenishment task record list is generated; Generate a list of AGV task execution records based on the task load prediction results; A sorting sequence list is generated for the work area based on the cargo demand forecast and inbound / outbound related fields. Based on the inventory change forecast results and the task load forecast results corresponding to each shelf, storage location and operation area, generate a warehouse area scheduling priority record for each warehouse operation area; write the warehouse operation area identifier and the corresponding priority identifier into each warehouse area scheduling priority record; The warehouse scheduling priority record is linked with the replenishment task record list, the AGV vehicle execution task record list, and the sorting operation sequence list to form a data structure for intelligent warehouse scheduling decision results; The intelligent warehouse scheduling decision result data structure is encapsulated into a scheduling instruction message through a preset network communication interface, and the scheduling instruction message is sent to the edge computing terminal deployed at the warehouse site. The edge computing terminal performs local execution control based on the replenishment task records, AGV vehicle execution task records, sorting operation sequence, and warehouse scheduling priority in the scheduling instruction message.
[0016] The beneficial effects of this invention are: This invention significantly improves the efficiency and accuracy of feature extraction from warehouse image data by combining an improved MobileSAM model with the EfficientViT image encoder. It effectively adapts to the visual characteristics of various areas in a warehouse setting, such as shelves, storage locations, and aisles, ensuring that the basic data possesses good semantic discrimination and spatial hierarchy perception capabilities. Through unified processing of multi-scale feature maps and contextual labeling embedding, it achieves accurate capture of multi-target information in complex warehouse environments, providing a high-quality input foundation for subsequent prompting control and mask segmentation.
[0017] The designed prompt control module incorporates structural alignment and region scheduling mechanisms. It performs sequential alignment, spatial index mapping, and region label association between point and bounding box prompts, ensuring accurate identification and reasonable scheduling of target prompt information in different warehousing operation areas during mask decoding. Through the coordinated invocation of prompt scheduling sequences and semantically preserved feature maps, the target recognition accuracy of the segmented mask image in different regions is improved, providing a stable and reliable input basis for the state awareness module.
[0018] The warehouse site status perception data and business data fusion mechanism constructed in this invention can efficiently generate multi-dimensional commodity time-series data for warehouse operations. Through the cloud-based PatchTST time series prediction model, it achieves multi-dimensional prediction of commodity demand, inventory changes, and operational load, demonstrating excellent time-series modeling and trend capture capabilities. Simultaneously, the replenishment task allocation, AGV task path planning, and operation sequence management logic constructed based on the prediction results ensure that the intelligent scheduling solution possesses the characteristics of timely response, reasonable path routing, and clear priority.
[0019] This invention not only achieves efficient and precise scene analysis at the image recognition level, but also establishes an intelligent prediction and execution system with edge-cloud linkage in the scheduling and decision-making process. It has comprehensive advantages such as clear structure, high computational efficiency, and flexible deployment, which significantly improves the overall intelligence level and operational efficiency of commodity circulation in the intelligent warehousing environment. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the overall structure of an intelligent decision-making system for commodity circulation based on edge-cloud collaboration proposed in this invention. Figure 2 This is an overall flowchart of an intelligent decision-making method for commodity circulation based on edge-cloud collaboration proposed in this invention. Figure 3 This is a schematic diagram of the structure of an improved MobileSAM model for an intelligent decision-making system for commodity circulation based on edge-cloud collaboration, as proposed in this invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0022] refer to Figures 1-3 A smart decision-making system for commodity circulation based on edge-cloud collaboration, comprising: The image acquisition and preprocessing module is used to acquire image data and business data from edge computing terminals deployed at the warehouse site, preprocess the image data, and obtain image sequences of the warehouse scene. The business data acquisition module is used to collect inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, and task execution logs of AGV vehicles. The edge feature extraction module is used to input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, the encoding sub-network with EfficientViT as the image encoder is called to extract features from the image sequence and obtain multi-scale feature maps of the warehousing scene. The improved MobileSAM model includes an encoding subnetwork, a semantic preservation module, and a mask decoder; The semantic preservation module is used to perform shape normalization, position indexing, and context labeling on the multi-scale feature maps, organizing the multi-scale feature maps into semantically preserved feature maps. The cue control module is used to process point cues and box cues and associate them with region labels within the structure alignment unit, schedule the set of structure alignment cues within the region scheduling unit, generate a cue scheduling sequence, and pass the cue scheduling sequence and semantic preservation feature map to the mask decoder. The edge-side scene analysis module is used to calculate warehouse site status perception data based on the segmented mask image after the mask decoder outputs the segmented mask image; The edge-side time-series construction module is used to perform time matching and field alignment between warehouse site status perception data and business data within the edge computing terminal, generating multi-dimensional commodity time-series data for warehouse operations; The cloud-based time series forecasting module is used to receive multi-dimensional commodity time series data from warehousing operations, construct multi-dimensional time series data, input the multi-dimensional time series data into the PatchTST time series forecasting model, and generate forecast results. The cloud-based scheduling and decision-making module is used to generate intelligent warehouse scheduling decision-making result data structures based on prediction results, and then distribute the intelligent warehouse scheduling decision-making results to edge computing terminals deployed on-site in the warehouse for execution.
[0023] In this embodiment, the modules are interconnected using the following method: Step 1: Image data and business data are collected by edge computing terminals deployed at the warehouse site. The image data is preprocessed to obtain an image sequence of the warehouse scene. Step 2: Input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, call the encoding sub-network with EfficientViT as the image encoder to extract features from the image sequence and obtain multi-scale feature maps of the warehouse scene. The improved MobileSAM model includes an encoding subnetwork, a semantic preservation module, and a mask decoder; Step 3: Input the multi-scale feature map into the semantic preservation module. The multi-scale feature map is structurally organized through shape normalization, position indexing, and context labeling operations to obtain the semantic preservation feature map. Step 4: The semantically preserved feature map is passed to the prompt control module. According to preset structural alignment rules, the point prompts and box prompts set for the storage operation area are sequentially aligned, spatially indexed, and associated with region labels. The region scheduling unit schedules the prompt information according to the label order of the storage operation area, generating a corresponding prompt scheduling sequence. The prompt scheduling sequence and the semantically preserved feature map are then passed to the mask decoder according to a preset calling order. The mask decoder outputs a segmentation mask image and calculates the storage site status perception data based on the segmentation mask image.
[0024] Step 5: Package the warehouse site status perception data and business data to generate multi-dimensional commodity time-series data of warehouse operations, and transmit the multi-dimensional commodity time-series data of warehouse operations to the cloud time series prediction module through the network by the edge computing terminal; Step 6: Construct a multidimensional time series in the cloud based on the multidimensional commodity time series data of warehousing operations, input the multidimensional time series into the PatchTST time series prediction model, and obtain the prediction results of the demand for goods, inventory changes, and workload of each shelf, storage location, and operation area within a preset time window. Step 7: Generate intelligent warehouse scheduling decision results in the cloud based on the forecast results of goods demand, inventory change and task load. These results include replenishment task allocation, AGV vehicle travel path planning, sorting operation sequence and warehouse area scheduling priority. The intelligent warehouse scheduling decision results are then distributed to the edge computing terminals deployed on the warehouse site for execution via the network.
[0025] In this embodiment, the image data comes from the warehouse racking area, aisle area, and work area; the business data includes inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, AGV task execution logs, and corresponding timestamps; the preprocessing includes size normalization, encoding format sorting, and time order sorting. In this embodiment, step two specifically includes: The image sequence is divided into N image subsequences according to a preset number of frames, and an index relationship corresponding to a timestamp is established for each image subsequence. The images in the image subsequence are input frame by frame into the EfficientViT image encoder in the coding subnetwork of the improved MobileSAM model. Convolution operation and feature embedding processing are performed on each frame to obtain the corresponding first-scale feature map. Within the EfficientViT image encoder, a multi-layer feature extraction structure is used to downsample and channel map the first-scale feature map to generate two intermediate feature maps with different resolutions. Two intermediate feature maps with different resolutions are the first intermediate feature map and the second intermediate feature map; The spatial size of the first-scale feature map is reduced to a preset first feature scale by the downsampling unit, and the number of channels is adjusted to the target number of channels in the channel mapping unit, so that the first-scale feature map is converted into a first intermediate feature map with a spatial resolution lower than that of the input feature map.
[0026] The first intermediate feature map is input to the next layer of feature extraction structure. After a second downsampling operation, a feature map with a further reduced spatial size is generated. Under the same channel mapping rule, a second intermediate feature map is formed, wherein the height and width dimensions of the second intermediate feature map are smaller than those of the first intermediate feature map.
[0027] The intermediate feature maps are arranged in a preset scale order and channel aligned to form a multi-scale feature map representing the spatial hierarchy information of the warehousing scene.
[0028] Based on the original MobileSAM model, this invention introduces EfficientViT as an image encoder for the coding sub-network, replacing the original lightweight network structure. Through multi-layer feature extraction and resolution progressive downsampling mechanism, it constructs a multi-scale feature representation form containing a first intermediate feature map and a second intermediate feature map, which significantly improves the model's ability to perceive spatial hierarchy and semantic parsing in complex warehouse scenarios, and achieves efficient and precise image feature extraction.
[0029] In this embodiment, step three specifically includes: The multi-scale feature map of the warehousing scenario is passed as input to the input of the semantic preservation module. The spatial size of the multi-scale feature map is unified, and the height and width of the multi-scale feature map are adjusted to the preset standard size. Feature maps that do not meet the standard size are filled to obtain the size-standardized feature map. Based on the size-normalized feature map, a corresponding location information index is generated for each spatial location. The location information indexes are arranged according to the row and column order of the size-normalized feature map, forming a structure corresponding to the size-normalized feature map. Figure 1 A corresponding location information matrix is generated, and the location information matrix is associated with the size-normalized feature map to obtain a feature map containing the location index sorting results; Contextual tags for indicating warehouse racking areas, aisle areas, and work areas are introduced into the channel dimension of the feature map containing the location index sorting results. The contextual tags are then appended to the corresponding feature channels using a preset encoding method to form a multi-scale feature map with contextual tags. The multi-scale feature maps with context labels are combined in the semantic preservation module according to the preset scale order and channel order, and the output is used as the semantic preservation feature map.
[0030] This invention restructures multi-scale feature maps by designing a semantic preservation module. It innovatively introduces a mechanism for size standardization, location information indexing and organization, and context labeling fusion. After unifying the original feature maps to a standard size, a spatial location information matrix is constructed, and context labels indicating the storage rack area, aisle area, and work area are added. Finally, a semantic preservation feature map is output, which realizes the synergistic enhancement of spatial consistency and regional semantics, and significantly improves the model's ability to understand the distribution of targets and regional context in the storage scene.
[0031] In this embodiment, step four specifically includes: The semantically preserved feature map is passed to the input of the prompt control module, which receives point prompts and box prompts set for the warehouse operation area. The prompt control module includes an input terminal, a structure alignment unit, and a region scheduling unit; The point prompts and box prompts are generated by the edge computing terminal based on the spatial layout information of the warehouse site, the preset area division rules of the images captured by the camera, the RFID tag identification position, the barcode scanner acquisition position, or the area coordinates contained in the AGV vehicle navigation path. The box prompts correspond to the boundary coordinates of each area, and the point prompts are the reference point coordinates generated in each area according to the preset benchmark point rules.
[0032] The warehousing operation area includes a shelving area, a storage location area, and a work area; Within the structural alignment unit, point hints and box hints are aligned sequentially according to preset structural alignment rules. Point hints and box hints belonging to the same timestamp are grouped together to form a hint sequence that corresponds to the time sequence of the warehousing scenario. The structural alignment rules include: Arrange the point-to-point and box-shaped prompts in chronological order according to the timestamps corresponding to the prompts; The prompts are spatially aligned according to their spatial index positions in the semantically preserving feature map. The prompts are categorized and grouped according to the area labels of the shelf area, storage location area, and work area to which they belong, so that prompts with the same timestamp and the same area label form a corresponding prompt sequence in structure.
[0033] Within the structural alignment unit, spatial indices are assigned to point and box prompts in each prompt sequence based on the spatial dimensions of the semantically preserved feature map. The spatial indices are then mapped one-to-one with the prompt positions in the prompt sequence to generate a set of structurally aligned prompts containing spatial index information. Within the structure alignment unit, the area labels representing the storage operation area are associated with the point tips and box tips in the structure alignment tip set, respectively. An area label mark consistent with the area to which each tip belongs is attached to each tip, resulting in a structure alignment tip set with area label marks. The set of structure alignment hints marked with area labels is input into the area scheduling unit. Within the area scheduling unit, the set of structure alignment hints is scheduled according to the label order of the storage operation area to generate a hint scheduling sequence corresponding to the area label order. The prompt scheduling sequence and semantic preservation feature map are passed together to the input of the mask decoder in a preset calling order, and the mask decoder outputs the segmentation mask image. Based on the segmented mask image, the system calculates the quantity of goods, the area occupied, and the spatial relationship of each shelf, storage location, and temporary storage area, forming warehouse site status perception data.
[0034] The calculation of the quantity information, occupied area information, and spatial location relationship of goods corresponding to shelves, storage locations, and temporary storage areas based on segmented mask images is specifically as follows: The quantity of goods is obtained by performing connected component detection on the segmented mask image and counting the number of connected components corresponding to the pallet mask and the cargo box mask. The occupied area information is obtained by statistically analyzing the number of foreground pixels in each mask region of the segmented mask image to obtain the pixel area, and the area record is calculated based on the correspondence between image resolution and physical scale. The spatial relationship is obtained by boundary positioning and center point calculation of the foreground pixel coordinates of the mask area, and the shelf, storage location or temporary storage area to which the mask area belongs is determined according to the correspondence between the coordinate position of the outer bounding box and the coordinate interval of the storage operation area. Based on the correspondence between the spatial location of the outer bounding box and the coordinate intervals of the shelf area, storage area and temporary storage area in the warehousing scenario, the spatial ownership of the mask area is determined, and warehousing site status perception data containing information on the quantity of goods, the occupied area and the spatial location relationship is generated.
[0035] This invention innovatively proposes a prompt generation and invocation method based on structural alignment and regional scheduling by designing a collaborative mechanism between the prompt control module and the mask decoding process. First, the prompt control module receives point prompts and box prompts generated by the edge computing terminal. The prompt information sources include diverse spatial data such as camera coverage areas, RFID identification locations, barcode scanner operation points, and AGV vehicle paths, fully reflecting the dynamic operational characteristics of the warehouse site. The structural alignment unit performs multi-level synchronous processing of the prompt data, including timestamps, spatial indexes, and regional labels, ensuring that the input prompts have spatiotemporal consistency and regional semantic distinguishability. Subsequently, the prompt scheduling sequence and semantically preserved feature map are jointly input to the mask decoder in a preset order, achieving partitioned response of regional prompts and accurate segmentation of the mask image. In the state-aware computing, technologies such as connected component detection, pixel area conversion, and coordinate attribution judgment are introduced to accurately extract the quantity of goods, occupied area, and spatial location relationships, ultimately forming perceptual data reflecting the real-time state of the warehouse site. The above methods demonstrate a high degree of customization and engineering practicality in terms of prompt generation, structural alignment logic, and state data interpretation mechanism, effectively improving the accuracy of target segmentation and the detail of region perception in complex scenarios.
[0036] In this embodiment, step five specifically includes: Within the edge computing terminal, based on the timestamp information contained in the warehouse site status perception data and business data, the warehouse site status perception data and business data are time-matched, and warehouse site status perception data records with the same timestamp are associated with business data records to form a set of warehouse operation data records arranged in chronological order. Based on the shelf identifiers, storage location identifiers, and work area identifiers contained in the warehousing operation data record set, the warehousing operation data records are grouped together, and warehousing operation data records belonging to the same shelf, the same storage location, or the same work area are grouped together. Within each group, the quantity of goods, area occupied, and spatial location relationships in the warehouse site status perception data are aligned with the inbound and outbound records, goods identification information, and task execution logs in the business data, and arranged in time order to construct a multi-dimensional commodity time-series data record sequence corresponding to the shelves, storage locations, and work areas. The multidimensional commodity time-series data records of each group are appended with warehousing operation area labels, equipment identification information and time range identifiers to form warehousing operation multidimensional commodity time-series data, and the warehousing operation multidimensional commodity time-series data is encapsulated into a data message to be transmitted; The edge computing terminal sends the data packets to be transmitted to the receiving end of the cloud time series prediction module through a preset network communication interface.
[0037] In this embodiment, step six specifically includes: Based on the timestamps contained in the multidimensional commodity time-series data of warehousing operations, the multidimensional commodity time-series data of warehousing operations is arranged into a continuous data sequence in chronological order; Based on the field parsing of the data sequence, the fields of goods quantity, inventory level, occupied area, spatial location, inbound / outbound and task execution are extracted from the data sequence in sequence according to the preset field structure, and together they form a multidimensional time series. Within the cloud-based time series forecasting module, a preset time window length and forecast time span are set for each multidimensional time series. Each multidimensional time series is then divided according to the time window to obtain a time window sequence consisting of M consecutive time windows. The input feature fields within each time window are organized into a model input sequence, and the corresponding goods quantity field, inventory level field, and task execution quantity field within the target time range after the time window are organized into a prediction target sequence. Input the model input sequences corresponding to each shelf, storage location, and work area into the PatchTST time series prediction model; In the PatchTST time series forecasting model, the model input sequences are encoded and time-dependent models are modeled, and the model output sequences are output in a one-to-one correspondence with each model input sequence. The model output sequence is restored and mapped to the time axis corresponding to each shelf, storage location and work area to obtain the prediction results for each shelf, storage location and work area within the preset time window. The forecast results include forecasts of demand for goods, inventory changes, and workload.
[0038] In this embodiment, step seven specifically includes: The cloud server where the cloud time series forecasting module is located receives the forecast results and classifies them according to shelf identifiers, storage location identifiers, and work area identifiers to form a set of forecast results organized by shelf, storage location, and work area respectively. Based on inventory change forecasts and goods demand forecasts, generate replenishment task records for each shelf and each storage location; The replenishment task record obtains the inventory gap by reading the predicted demand field and the predicted inventory level field and comparing the difference between the two fields; when the inventory gap meets the preset replenishment trigger conditions, the inventory gap is written into the replenishment quantity field, and a planned replenishment time field is generated according to the time indication of the predicted demand; the replenishment task record is composed of the target shelf identifier and the corresponding goods identifier, and a replenishment task record list is formed. Write the target shelf, corresponding product identifier, replenishment quantity field and planned replenishment time field into each replenishment task record to form a replenishment task record list; Generate a list of AGV task execution records based on the task load prediction results; Extract the task requirement field from the task load prediction results and write the warehouse operation area identifier corresponding to the task requirement field into the starting position field; determine the target position field from the replenishment task record, sorting area or warehouse node according to the task type; generate the driving path identifier according to the warehouse topology and write the time record corresponding to the task requirement field into the expected execution time field to generate the AGV trolley execution task record. Based on the forecast of cargo demand and relevant fields for inbound and outbound, sorting task records are generated for each work area. The corresponding cargo identifier, target shelf, sorting quantity field and sorting time field are written into each sorting task record. The sorting task records are sorted according to the preset sorting rules to form a sorting order list. Based on the inventory change prediction results and the task load prediction results corresponding to each shelf, storage location and operation area, a warehouse area scheduling priority record is generated for each warehouse operation area. The warehouse operation area identifier and the corresponding priority identifier are written into each warehouse area scheduling priority record. The warehouse area scheduling priority record is associated with the replenishment task record list, the AGV vehicle execution task record list and the sorting operation sequence list to form an intelligent warehouse scheduling decision result data structure. The intelligent warehouse scheduling decision result data structure is encapsulated into scheduling instruction messages through a preset network communication interface, and the scheduling instruction messages are sent to the edge computing terminals deployed on the warehouse site. The edge computing terminals perform local execution control based on the replenishment task records, AGV vehicle execution task records, sorting operation sequence and warehouse area scheduling priority in the scheduling instruction messages.
[0039] Example 1: To verify the feasibility of this invention in practice, it was applied to the daily operation scenario of a comprehensive intelligent warehousing center as the experimental site. The system deployment included multiple edge computing terminals, image acquisition cameras, AGVs, barcode scanners, and RFID readers. This warehousing center has standardized shelving areas, fixed storage areas, and multiple temporary operation areas, with an average of over 600 inbound and outbound transactions per day and over 300 SKUs. During the experimental phase, an improved MobileSAM model was deployed at the edge to collect video streams from cameras within the coverage area in real time, combined with barcode scanners and RFID to collect business data, forming a multimodal input for the warehousing scenario. The system first preprocesses the image stream, dividing it into multiple image subsequences. After extraction by the EfficientViT encoder, a three-level multi-scale feature map is generated. The semantic preservation module completes shape unification, position information supplementation, and context label injection, outputting a semantically preserved feature map.
[0040] In the warehouse area division, the shelving area, storage location area, and operation area are each assigned independent area labels. Combining the camera deployment coordinates, the geographic mapping location of RFID tags, and AGV path data, point prompts and box prompts are set for different areas. The prompt information completes time alignment, spatial alignment, and area classification within the structure alignment unit. The scheduling unit generates a well-structured prompt scheduling sequence, which, together with the semantic feature map, is input into the mask decoder to finally output a high-precision mask image.
[0041] After post-processing, the masked image allows the system to automatically identify the quantity of goods (pallets / boxes), calculate the actual occupied area, generate 3D coordinate information, and construct real-time warehouse status perception data. In this scenario, the segmentation accuracy remains above 92% even under low light and complex goods stacking conditions, with a location matching error of less than 10 centimeters. Compared with traditional manual inventory methods, this system achieves a processing capacity of 180 shelves per hour without increasing manpower, which is 6.5 times that of manual methods.
[0042] The system integrates status data and business data to generate multi-dimensional product time-series data, which is then transmitted to the cloud-based PatchTST prediction model for joint prediction of goods demand, inventory changes, and task load. Using a 7-day window, the system achieved a prediction accuracy of 93.2%. Within a one-week operating cycle, a total of 186 replenishment task messages, 231 AGV path scheduling records, and 298 sorting sequence suggestions were generated. Specific experimental data is shown in Table 1. Table 1. Comparison of scheduling efficiency and accuracy during the trial of the intelligent warehousing system.
[0043] As shown in Table 1, the deployed system achieved accurate replenishment, reasonable scheduling, and efficient sorting. During the testing period, the average AGV empty-run rate decreased by 23.8%, and the time for a single picking task was reduced to 62% of the original time. Simultaneously, through the intelligent early warning function, 14 delays caused by stockouts or inventory buildup were successfully avoided.
[0044] This embodiment demonstrates that the present invention can significantly improve the data acquisition accuracy, task scheduling efficiency, and predictive control capabilities in intelligent warehousing scenarios, achieving the expected effect of improving the overall level of intelligent operation.
[0045] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart decision-making system for commodity circulation based on edge-cloud collaboration, characterized in that, include: The image acquisition and preprocessing module is used to acquire image data and business data from edge computing terminals deployed at the warehouse site, preprocess the image data, and obtain image sequences of the warehouse scene. The business data acquisition module is used to collect inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, and task execution logs of AGV vehicles. The edge feature extraction module is used to input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, the encoding sub-network with EfficientViT as the image encoder is called to extract features from the image sequence and obtain multi-scale feature maps of the warehousing scene. The improved MobileSAM model includes an encoding subnetwork, a semantic preservation module, and a mask decoder; The semantic preservation module is used to perform shape normalization, position indexing, and context labeling on the multi-scale feature maps, organizing the multi-scale feature maps into semantically preserved feature maps. The cue control module is used to process point cues and box cues and associate them with region labels within the structure alignment unit, schedule the set of structure alignment cues within the region scheduling unit, generate a cue scheduling sequence, and pass the cue scheduling sequence and semantic preservation feature map to the mask decoder. The edge-side scene analysis module is used to calculate warehouse site status perception data based on the segmented mask image after the mask decoder outputs the segmented mask image; The edge-side time-series construction module is used to perform time matching and field alignment between warehouse site status perception data and business data within the edge computing terminal, generating multi-dimensional commodity time-series data for warehouse operations; The cloud-based time series forecasting module is used to receive multi-dimensional commodity time series data from warehousing operations, construct multi-dimensional time series data, input the multi-dimensional time series data into the PatchTST time series forecasting model, and generate forecast results. The cloud-based scheduling and decision-making module is used to generate intelligent warehouse scheduling decision-making result data structures based on prediction results, and then distribute the intelligent warehouse scheduling decision-making results to edge computing terminals deployed on-site in the warehouse for execution.
2. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 1, characterized in that, The modules are connected in the following way: Step 1: Image data and business data are collected by edge computing terminals deployed at the warehouse site. The image data is preprocessed to obtain an image sequence of the warehouse scene. Step 2: Input the image sequence into the improved MobileSAM model. In the improved MobileSAM model, call the encoding sub-network with EfficientViT as the image encoder to extract features from the image sequence and obtain multi-scale feature maps of the warehouse scene. Step 3: Input the multi-scale feature map into the semantic preservation module. The multi-scale feature map is structurally organized through shape normalization, position indexing, and context labeling operations to obtain the semantic preservation feature map. Step 4: Pass the semantically preserved feature map to the prompt control module, and process and associate the point prompts and box prompts set for the warehouse operation area with the area labels according to the preset structural alignment rules; The regional scheduling unit generates a corresponding prompt scheduling sequence and passes the prompt scheduling sequence and semantic preservation feature map to the mask decoder in a preset calling order; The mask decoder outputs a segmented mask image, and the warehouse site status perception data is calculated based on the segmented mask image; Step 5: Package the warehouse site status perception data and business data to generate multi-dimensional commodity time-series data of warehouse operations, and transmit the multi-dimensional commodity time-series data of warehouse operations to the cloud time series prediction module through the network by the edge computing terminal; Step 6: Construct a multidimensional time series in the cloud based on the multidimensional commodity time series data of warehousing operations, and input the multidimensional time series into the PatchTST time series prediction model to obtain the prediction results for each shelf, storage location and operation area within the preset time window; Step 7: Generate intelligent warehouse scheduling decision results in the cloud based on the forecast results of goods demand, inventory change and task load. These results include replenishment task allocation, AGV vehicle travel path planning, sorting operation sequence and warehouse area scheduling priority. The intelligent warehouse scheduling decision results are then distributed to the edge computing terminals deployed on the warehouse site for execution via the network.
3. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, The image data comes from the warehouse racking area, aisle area, and work area; the business data includes inbound and outbound records collected by barcode scanners, cargo identification information identified by RFID tags, AGV task execution logs, and corresponding timestamps; preprocessing includes size normalization, encoding format sorting, and time order sorting.
4. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step two specifically involves: The image sequence is divided into N image subsequences according to a preset number of frames; The images in the image subsequence are input frame by frame into the EfficientViT image encoder in the coding subnetwork of the improved MobileSAM model. Convolution operation and feature embedding processing are performed on each frame image to obtain the corresponding first-scale feature map. Within the EfficientViT image encoder, a multi-layer feature extraction structure is used to downsample and channel map the first-scale feature map to generate two intermediate feature maps with different resolutions. The intermediate feature maps are arranged in a preset scale order and channel aligned to form a multi-scale feature map representing the spatial hierarchy information of the warehousing scene.
5. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step three specifically involves: The multi-scale feature map of the warehousing scenario is passed as input to the input end of the semantic preservation module. The spatial size of the multi-scale feature map is unified, and the height and width of the multi-scale feature map are adjusted to the preset standard size. Multi-scale feature maps that do not meet the standard size are filled to obtain the size-standardized feature map. Based on the size-normalized feature map, a corresponding location information index is generated for each spatial location; The location information indexes are arranged according to the row and column order of the size-normalized feature map to form a location information matrix that corresponds one-to-one with the size-normalized feature map; By associating the location information matrix with the size-normalized feature map, a feature map containing the location index sorting results is obtained; Introduce contextual tags in the channel dimension of the feature map containing the location index sorting results to indicate storage rack areas, aisle areas, and work areas; Context labels are appended to the corresponding feature channels using a preset encoding method to form a multi-scale feature map with context labels; Multi-scale feature maps with context labels are combined in the semantic preservation module according to a preset scale order and channel order, and the output is used as a semantic preservation feature map.
6. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step four specifically involves: The semantically preserved feature map is passed to the input of the prompt control module, which receives point prompts and box prompts set for the warehouse operation area. The prompt control module includes an input terminal, a structure alignment unit, and a region scheduling unit; The warehousing operation area includes a shelving area, a storage location area, and a work area; Within the structural alignment unit, point hints and box hints are aligned sequentially according to preset structural alignment rules. Point hints and box hints belonging to the same timestamp are grouped together to form a hint sequence that corresponds to the time sequence of the warehousing scenario. Within the structure alignment unit, spatial indices are assigned to point and box prompts in the prompt sequence based on the spatial dimensions of the semantically preserved feature map. The spatial indices are then mapped one-to-one with the prompt positions in the prompt sequence to generate a structure alignment prompt set containing spatial index information. Within the structural alignment unit, the area labels representing the storage operation area are associated with the point tips and box tips in the structural alignment tip set, respectively. An area label mark consistent with the storage operation area to which the point tip belongs is attached to each point tip and each box tip, forming a structural alignment tip set with area label marks. The set of structure alignment hints marked with area labels is input into the area scheduling unit. Within the area scheduling unit, the set of structure alignment hints is scheduled according to the label order of the storage operation area to generate a hint scheduling sequence corresponding to the area label order. The prompt scheduling sequence and semantically preserved feature map are passed together to the mask decoder in a preset calling order, and the mask decoder outputs the segmentation mask image; The segmented mask image includes the outlines of pallets, containers, storage locations, and obstacles; Based on segmented mask images, the quantity of goods, occupied area, and spatial relationship of goods corresponding to shelves, storage locations, and temporary storage areas are calculated to form warehouse site status perception data.
7. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step five specifically involves: Within the edge computing terminal, based on the timestamp information contained in the warehouse site status perception data and business data, the warehouse site status perception data and business data are time-matched, and warehouse site status perception data records with the same timestamp are associated with business data records to form a set of warehouse operation data records arranged in chronological order. Based on the shelf identifiers, storage location identifiers, and work area identifiers contained in the warehousing operation data record set, the warehousing operation data records are grouped together, and warehousing operation data records belonging to the same shelf, the same storage location, or the same work area are grouped together. Within each group, the quantity of goods, area occupied, and spatial location relationships in the warehouse site status perception data are aligned with the inbound and outbound records, goods identification information, and task execution logs in the business data, and arranged in time order to construct a multi-dimensional commodity time-series data record sequence corresponding to the shelves, storage locations, and work areas. The multidimensional commodity time-series data records of each group are appended with warehousing operation area labels, equipment identification information and time range identifiers to form warehousing operation multidimensional commodity time-series data, and the warehousing operation multidimensional commodity time-series data is encapsulated into a data message to be transmitted; The edge computing terminal sends the data packets to be transmitted to the receiving end of the cloud time series prediction module through a preset network communication interface.
8. The intelligent decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step six specifically involves: Based on the timestamps contained in the multidimensional commodity time-series data of warehousing operations, the multidimensional commodity time-series data of warehousing operations is arranged into a continuous data sequence in chronological order; Based on the field parsing of the data sequence, the fields of goods quantity, inventory level, occupied area, spatial location, inbound / outbound and task execution are extracted from the data sequence in sequence according to the preset field structure, and together they form a multidimensional time series. Within the cloud-based time series forecasting module, a preset time window length and forecast time span are set for each multidimensional time series. Each multidimensional time series is then divided according to the time window to obtain a time window sequence consisting of M consecutive time windows. The input feature fields within each time window are organized into a model input sequence, and the corresponding goods quantity field, inventory level field, and task execution quantity field within the target time range after the time window are organized into a prediction target sequence. Input the model input sequences corresponding to each shelf, storage location, and work area into the PatchTST time series prediction model; In the PatchTST time series forecasting model, the model input sequences are encoded and time-dependent models are modeled, and the model output sequences are output in a one-to-one correspondence with each model input sequence. The model output sequence is restored and mapped to the time axis corresponding to each shelf, storage location and work area to obtain the prediction results for each shelf, storage location and work area within the preset time window. The forecast results include forecasts of demand for goods, inventory changes, and workload.
9. A smart decision-making system for commodity circulation based on edge-cloud collaboration according to claim 2, characterized in that, Step seven specifically involves: The cloud server where the cloud-based time series forecasting module is located receives the forecast results and classifies them according to shelf identifiers, storage location identifiers, and work area identifiers to form a set of forecast results. Based on the inventory change forecast and the goods demand forecast, a replenishment task record list is generated; Generate a list of AGV task execution records based on the task load prediction results; A sorting sequence list is generated for the work area based on the cargo demand forecast and inbound / outbound related fields. Based on the inventory change forecast results and the task load forecast results corresponding to each shelf, storage location and operation area, generate a warehouse area scheduling priority record for each warehouse operation area; write the warehouse operation area identifier and the corresponding priority identifier into each warehouse area scheduling priority record; The warehouse scheduling priority record is linked with the replenishment task record list, the AGV vehicle execution task record list, and the sorting operation sequence list to form a data structure for intelligent warehouse scheduling decision results; The intelligent warehouse scheduling decision result data structure is encapsulated into a scheduling instruction message through a preset network communication interface, and the scheduling instruction message is sent to the edge computing terminal deployed at the warehouse site. The edge computing terminal performs local execution control based on the replenishment task records, AGV vehicle execution task records, sorting operation sequence, and warehouse scheduling priority in the scheduling instruction message.