Industrial Internet of Things, control methods and storage media for intelligent automated warehouses
The industrial IoT system with a five-platform structure solves the problems of complex data processing and high error rate of automated stacker cranes in intelligent warehouses, and realizes efficient and accurate data interaction and processing, thereby improving the operational efficiency and security of the warehouse.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHENGDU QINCHUAN IOT TECH CO LTD
- Filing Date
- 2023-07-10
- Publication Date
- 2026-05-26
AI Technical Summary
The data processing and management of automated stacker cranes in intelligent automated warehouses are complex, resulting in high error rates, high system costs, and easy mutual interference between multiple stacker cranes, affecting the construction of automation and safe and stable operation.
The industrial IoT system adopts a five-platform architecture, including a user platform, a service platform, a management platform, a sensor network platform, and an object platform. Through independent and front-side platform deployments, data is stored, processed, and transmitted separately to ensure accurate and efficient data interaction and processing.
It reduced data processing pressure and setup costs, improved data processing speed and accuracy, ensured the independent operation of automated stacker cranes, reduced error rates, and improved the operational efficiency and safety of automated warehouses.
Smart Images

Figure CN116902449B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of intelligent manufacturing technology, and in particular to the Industrial Internet of Things, control methods and storage media for intelligent automated warehouses. Background Technology
[0002] Automated storage and retrieval systems (AS / RS), also known as high-bay warehouses or high-bay storage facilities, generally refer to warehouses that use several, a dozen, or even dozens of layers of high shelving to store goods and use corresponding material handling equipment for inbound and outbound operations. Intelligent AS / RS, also known as automated storage and retrieval systems, are fully intelligent warehouses that combine mechanical, electrical, and information technology. They mainly consist of a goods storage system, a goods storage and retrieval and conveying system, and a control and management system, enabling fully automated, unattended operation.
[0003] In automated storage and retrieval systems (AS / RS), stacker cranes are indispensable as the main mechanical equipment for material storage and retrieval. Also known as stacker hoists, stacker cranes are specialized cranes that use forks or booms as picking devices to grab, move, and stack unitized goods in warehouses, workshops, and other locations, or to retrieve and place unitized goods from high-level racks. Their primary function is to move back and forth within the aisles of the AS / RS, storing goods located at aisle entrances into rack compartments, or retrieving goods from racks and transporting them to aisle entrances or designated locations.
[0004] In the field of intelligent manufacturing technology, intelligent automated warehouses achieve unmanned operation through automated stacker cranes. Automated stacker cranes need to pick up and place goods according to regulations. For example, goods from different workshops, of different types, or for different purposes are placed in different areas, or goods from different inbound and outbound channels are placed in zones with specific numbers. Moreover, some factory areas have large areas and multiple layers of automated warehouses, and each area needs to be equipped with multiple stacker cranes for picking up and placing goods. The number of automated stacker cranes is large, and the instantaneous data is also extremely large. When managing and controlling such a large quantity and amount of data using an automated system, a large data storage and processing system is required. Furthermore, data processing must be timely and efficient. All of this leads to the complexity and high cost of the existing system structure, and errors are prone to occur during data interaction, resulting in problems such as incorrect picking and placing of goods or incorrect placement positions by automated stacker cranes. All of these are detrimental to the automated construction and safe and stable operation of intelligent automated warehouses. Summary of the Invention
[0005] One embodiment of this specification provides an Industrial Internet of Things (IIoT) for an intelligent automated warehouse, including a user platform, a service platform, and a management platform. The management platform is configured to perform the following operations: the user platform sends a first instruction to the service platform; the service platform determines a second instruction based on the first instruction and sends it to the management platform; in response to receiving the second instruction from the service platform, it determines whether the motion force of the automated stacker crane meets the execution conditions, wherein the first instruction includes a goods retrieval instruction; in response to the automated stacker crane's motion force meeting the execution conditions, it controls the automated stacker crane to perform a pick-and-place operation; in response to the automated stacker crane completing the pick-and-place operation, it determines a target standby position and controls the automated stacker crane to move to the target standby position.
[0006] One embodiment of this specification provides a control method for an industrial Internet of Things (IIoT) for an intelligent automated warehouse. The IIoT comprises, from top to bottom, a user platform, a service platform, a management platform, a sensor network platform, and an object platform that interact sequentially. The service platform is arranged independently, while the management platform and the sensor network platform are arranged in a front-sub-platform configuration. The independent arrangement means that the service platform has multiple independent sub-platforms, each storing, processing, and / or transmitting data from different lower-level platforms. The front-sub-platform configuration means that each platform has a main platform and multiple sub-platforms, each storing and processing data of different types or receiving objects sent from lower-level platforms. The main platform aggregates, stores, processes, and transmits the data from the multiple sub-platforms to the upper-level platform. The object platform is configured as an automated stacker crane in different shelving areas of the intelligent automated warehouse. The control method includes: each sub-platform of the service platform corresponds to a different user platform; a first instruction is issued by the user platform; the sub-platform of the service platform corresponding to the user platform receives the first instruction and... The first instruction is converted into a second instruction recognizable by the management platform and sent to the main platform of the management platform. The second instruction includes at least an instruction number, shelf area, quantity of goods to be picked up or placed, and location of goods to be picked up or placed. The main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the sub-platform of the management platform corresponding to the shelf area based on the shelf area information. After receiving the second instruction, the sub-platform of the management platform extracts the instruction number, quantity of goods to be picked up or placed, and location of goods to be picked up or placed and converts them into a third instruction recognizable by the automated stacker crane. The sub-platform sends the shelf area information and the third instruction together to the main platform of the sensor network platform. After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information. The sub-platform of the sensor network platform receives the third instruction and sends it to one or more automated stacker cranes corresponding to it. The one or more automated stacker cranes perform goods picking and placing based on the third instruction and feed back the goods picking and placing result data.
[0007] One embodiment of this specification provides a computer-readable storage medium, characterized in that the storage medium stores computer instructions, and when a computer reads the computer instructions in the storage medium, the computer executes the method described in any of the above embodiments. Attached Figure Description
[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0009] Figure 1 This is an exemplary flowchart of an industrial Internet of Things (IoT) for intelligent automated warehouses, as shown in some embodiments of this specification.
[0010] Figure 2 This is an exemplary structural framework diagram of an industrial Internet of Things for intelligent automated warehouses, as shown in some embodiments of this specification.
[0011] Figure 3 This is an exemplary flowchart of an industrial Internet of Things control method for an intelligent automated warehouse, according to some embodiments of this specification;
[0012] Figure 4 This is an exemplary schematic diagram illustrating the determination of target standby positions based on reinforcement learning models according to some embodiments of this specification;
[0013] Figure 5 This is an exemplary flowchart illustrating the determination of a mobile reward value based on a first hit rate, according to some embodiments of this specification;
[0014] Figure 6 This is an exemplary flowchart illustrating the determination of mobile reward values based on percentage compliance according to some embodiments of this specification. Detailed Implementation
[0015] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0016] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0017] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0018] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0019] Industrial IoT applications for smart automated warehouses can include processing devices, networks, storage devices, and automated stacker cranes. Processing devices handle information and / or data related to the application scenario, and the management platform can be implemented within these devices. Networks enable communication between the various components within the application scenario. Storage devices store data, instructions, and / or any other information. Automated stacker cranes receive instructions from the management platform and execute tasks such as picking and placing goods within the smart automated warehouse.
[0020] Figure 1 This is an exemplary flowchart of an Industrial Internet of Things (IIoT) for an intelligent automated warehouse, according to some embodiments of this specification. In some embodiments, process 100 may be executed by a management platform. Figure 1 As shown, process 100 includes
[0021] The following steps:
[0022] Step 110: The user platform sends a first instruction to the service platform. The service platform determines a second instruction based on the first instruction and sends it to the management platform. In response to receiving the second instruction from the service platform, the management platform determines whether the motion force of the automated stacker crane meets the execution conditions.
[0023] The first instruction refers to the instruction used to control the automated stacker crane to perform a task. For example, the first instruction can be a goods retrieval instruction. A goods retrieval instruction can be an instruction used to instruct the automated stacker crane to perform a goods retrieval task. The instruction content of the first instruction may include the rack area to which the automated stacker crane is required to go, the goods retrieval location, and the quantity of goods to be retrieved or retrieved. The instruction content of the first instruction may also include movement path information instructing how the automated stacker crane should move. In some embodiments, the first instruction may be issued by the user platform and received by the service platform. For more information about user platforms and service platforms, please refer to [link to relevant documentation]. Figure 2 And its related descriptions.
[0024] In some embodiments, the first instruction may include an inbound instruction and an outbound instruction.
[0025] An inbound instruction can be a command that instructs an automated stacker crane to retrieve goods from the picking station and place them on the rack. An outbound instruction can be a command that instructs an automated stacker crane to retrieve goods from the rack and place them on the picking station.
[0026] The second instruction refers to an instruction that the management platform can recognize based on the first instruction. The second instruction has the same content as the first instruction; for example, the instruction content of the second instruction may include the racking area to which the automated stacker crane should go, the location for picking or placing goods, the quantity of goods to be picked or placed, and movement path information instructing the automated stacker crane how to move. In some embodiments, the second instruction may be obtained by converting the first instruction based on a service platform. More information on converting the first instruction based on a service platform can be found in [link to relevant documentation]. Figure 2 And its related descriptions.
[0027] A racking area can be an area consisting of one or more racks in an automated warehouse. An automated stacker crane in an automated warehouse can correspond to at least one racking area. The management platform can send a second instruction to the automated stacker crane corresponding to the racking area based on the racking area specified in the second instruction.
[0028] Goods pickup and drop-off locations can include pickup locations and drop-off locations. These locations can be represented in various ways; for example, they can be represented by two three-dimensional coordinates. For instance, a pickup location with three-dimensional coordinates (x, y, z) indicates that the pickup is located in the x-th row and y-th column of the z-th shelf, where x, y, and z are all integers greater than 0. In some embodiments, when a pickup station needs to be represented, it can be represented using three-dimensional coordinates including 0. For example, pickup station number 1 can be represented as (0, 0, 1).
[0029] The kinetic force of an automated stacker crane can be used to describe its mobility. For example, the kinetic force of an automated stacker crane can be the maximum distance it can travel in the current state. The kinetic force of an automated stacker crane can be determined based on the current remaining battery / fuel level.
[0030] In some embodiments, the management platform can determine the motion force required for the instruction by the current position of the automated stacker, as well as the shelf area, goods pick-up and drop location, and goods pick-up and drop quantity in the second instruction, and determine whether the motion force of the automated stacker meets the execution conditions. If the motion force of the automated stacker is greater than the motion force required for the instruction, then the execution conditions are met.
[0031] Step 120: In response to the automatic stacker's motion force meeting the execution conditions, control the automatic stacker to perform pick-up and place-down operations.
[0032] In some embodiments, in response to the automatic stacker's motion force meeting the execution conditions, the management platform can send a second instruction to the automatic stacker to control the automatic stacker to perform pick-up and put-down operations.
[0033] In some embodiments, in response to the automated stacker crane's motion force satisfying the execution conditions, the management platform can extract and convert the instruction number, quantity of goods picked up and placed, and location of goods picked up and placed from the second instruction into a third instruction recognizable by the automated stacker crane, and send the rack area information and the third instruction together to the main platform of the sensor network platform; after receiving the rack area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the rack area based on the rack area information; the sub-platform of the sensor network platform receives the third instruction and sends it to one or more of its corresponding automated stacker cranes to control the automated stacker cranes to perform the picking and placing operation. For more information on the third instruction and the sensor network platform, please refer to [link to relevant documentation]. Figure 2 And its related descriptions.
[0034] Step 130: In response to the automated stacker completing the pick-and-place operation, determine the target standby position and control the automated stacker to move to the target standby position.
[0035] The target standby location refers to the location that the automated stacker crane needs to go to when it is idle. The target standby location can be represented by a three-dimensional vector (x, y, z), which means that the automated stacker crane needs to go to the location where it can directly pick up or place goods in the x-th row and y-th column of the z-th rack.
[0036] Idle time refers to the period between when an automated stacker crane completes an instruction and when the next instruction arrives. Once the management platform sends a new instruction to the automated stacker crane, the idle time ends and it begins executing that instruction.
[0037] In some embodiments, the idle period can be determined based on the instruction execution status. The instruction execution status can include the start and end times of the automated stacker crane executing instructions. For example, if the instruction execution status of the automated stacker crane over a historical period is as follows: instruction A starts at 8:00, instruction A is completed at 8:05, instruction B starts at 8:15, instruction B is completed at 8:17, instruction C starts at 8:35, instruction C is completed at 8:40, and so on: then: the time period 8:05-8:15 is an idle period, with the start time at 8:05 and the end time at 8:15; the time period 8:17-8:35 is an idle period, with the start time at 8:17 and the end time at 8:35.
[0038] In some embodiments, the target standby location can be determined based on preset rules. For example, the preset rule could be that the management platform assigns a corresponding target standby location to each automated stacker, and when the automated stacker is idle, the management platform can control the automated stacker to move to its corresponding target standby location. Another example is that the preset rule could be that when the automated stacker is idle, it should remain stationary in place.
[0039] In some embodiments, the target waiting location can be determined based on a reinforcement learning model. More information about reinforcement learning models and determining the target waiting location based on reinforcement learning models can be found at [link to relevant documentation]. Figure 4 And its related descriptions.
[0040] In some embodiments of this specification, the automatic stacker cranes that receive goods retrieval instructions from the user platform through a unified management platform and allocate them to the corresponding shelf areas can better manage and control the retrieval and placement of goods. At the same time, the movement force of the automatic stacker cranes is taken into account when allocating goods retrieval instructions to avoid the automatic stacker cranes running out of energy during the execution of goods retrieval and placement, which would cause on-site operations to be blocked. This can make the intelligent automated warehouse operate more efficiently and quickly.
[0041] like Figure 2 As shown, the first embodiment of this specification aims to provide an Industrial Internet of Things (IIoT) for intelligent automated warehouses. The IIoT for intelligent automated warehouses includes, from top to bottom, an interactive user platform, a service platform, a management platform, a sensor network platform, and an object platform. The service platform adopts an independent layout, while the management platform and sensor network platform both adopt a front-sub-platform layout. The independent layout means that the service platform has multiple independent sub-platforms, each storing, processing, and / or transmitting data from different lower-level platforms. The front-sub-platform layout means that each platform has a main platform and multiple sub-platforms. The sub-platforms store and process data of different types or receiving objects sent by lower-level platforms, while the main platform aggregates, stores, processes, and transmits the data to the upper-level platform. The object platform is configured as an automated stacker crane in different shelving areas of the intelligent automated warehouse.
[0042] The service platform's sub-platforms correspond to different user platforms. The first instruction is issued by the user platform. The sub-platform of the service platform corresponding to the user platform receives the first instruction, converts it into a second instruction recognizable by the management platform, and sends the second instruction to the main platform of the management platform. The second instruction includes at least an instruction number, shelf area, quantity of goods picked up or placed, and location of goods picked up or placed. The main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the sub-platform of the management platform corresponding to the shelf area based on the shelf area information. After receiving the second instruction, the sub-platform of the management platform extracts the instruction number, quantity of goods picked up or placed, and location of goods picked up or placed in the second instruction and converts it into a third instruction recognizable by the automated stacker crane. It then sends the shelf area information and the third instruction together to the main platform of the sensor network platform. After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information.
[0043] The sub-platform of the sensor network platform receives the third instruction and sends it to one or more of its corresponding automated stacker cranes. The one or more automated stacker cranes perform cargo handling based on the third instruction and feed back the cargo handling result data.
[0044] Goods in automated warehouses must be categorized or placed according to pre-set rules or different regulations. Goods must be retrieved according to specific requirements to prevent misplacement and ensure proper storage and retrieval of subsequent goods. Currently, due to the large area, high number of floors, and large storage capacity of automated warehouses, multiple automated stacker cranes are typically installed. The sheer number of stacker cranes results in an exceptionally large volume of instantaneous data. Managing and controlling such a large quantity and amount of data requires a substantial data storage and processing system. Furthermore, data processing must be timely, efficient, and orderly. This leads to complex and costly existing system structures, and data interaction is prone to errors, causing problems such as incorrect goods retrieval, incorrect placement, and interference between multiple stacker cranes. All of these factors hinder the automated construction and safe and stable operation of intelligent automated warehouses.
[0045] The Industrial Internet of Things (IIoT) for intelligent automated warehouses described in this specification is built on a five-platform architecture. The service platforms are deployed independently, with each service platform's sub-platform corresponding to different user platforms. This ensures that each user platform has an independent service platform to perform data interaction, facilitating data management and resolving the issue of high data interaction pressure on a single service platform when multiple users are using the system. Furthermore, data can be interacted with user platforms more accurately and efficiently. Secondly, both the management platform and the sensor network platform adopt a front-side sub-platform layout. Both utilize a central platform to identify shelf areas and classify and distribute different data based on these areas. All data can be partitioned and categorized, and each data point can be associated with a corresponding sub-platform and upper-level platform. This ensures that all data can be processed and transmitted independently according to regional information. Both platforms have multiple sub-platforms, each of which operates independently and is responsible for processing and transmitting data for a unique shelf area. This enables the conversion and identification of corresponding data and allows for independent data interaction with upper and lower platforms. The data path throughout the entire processing process is unique and effective. By ensuring that the main platform and sub-platforms each perform their respective functions, the data processing requirements of both platforms are greatly reduced, the construction costs are also reduced, and the data processing speed and capabilities of each platform are further improved.
[0046] This manual describes an industrial IoT system for intelligent automated warehouses. In its use, each user platform corresponds to an independent and unique sub-platform of the service platform. Each user platform can correspond to different workshops, conveyor systems, factories, etc. This design allows for clear acquisition of information about the demand side of goods, facilitating subsequent goods statistics and traceability. The management platform, through a pre-sub-platform setup, centrally processes data from all service platform sub-platforms. Data from different service platform sub-platforms is categorized and partitioned by shelving area. Multiple management platform sub-platforms form an independent physical structure, each corresponding to its own shelving area and processing its own data independently, ensuring data interoperability between different areas. Without interference or conflict, this approach improves data processing accuracy and speed, while reducing the data processing pressure and setup costs of each sub-platform. Similarly, the sensor network platform also adopts a front-sub-platform layout, ensuring that one or more automated stacker cranes in each racking area correspond to an independent sub-platform of the sensor network platform. The main platform of the sensor network platform transmits the corresponding data independently to the corresponding sub-platforms, ensuring that the automated stacker cranes in each racking area operate relatively independently without affecting each other. This enables the automatic control of multiple automated stacker cranes in multiple racking areas, improving the efficiency and accuracy of picking and placing goods in the automated warehouse and effectively solving the problems of multiple automated stacker cranes repeatedly executing the same instructions and incorrect picking and placing.
[0047] It should be noted that the user platform in this embodiment can be a desktop computer, tablet computer, laptop computer, mobile phone, or other electronic device capable of data processing and data communication, and is not limited in many ways. In specific applications, the first server and the second server can be a single server or a server cluster, and is not limited in many ways. It should be understood that the data processing process mentioned in this embodiment can be processed by the server's processor, and the data stored on the server can be stored on the server's storage device, such as a hard disk or other storage device. In specific applications, the sensor network platform can use multiple gateway servers or multiple smart routers, and is not limited in many ways. It should be understood that the data processing process mentioned in the embodiments of this specification can be processed by the gateway server's processor, and the data stored on the gateway server can be stored on the gateway server's storage device, such as a hard disk or SSD or other storage device.
[0048] In practical applications, the corresponding service platform sub-platform receives the first instruction, converts it into a second instruction recognizable by the management platform, and sends the second instruction to the main management platform. Specifically, when the service platform sub-platform receives the first instruction, it extracts at least the instruction number, shelf area, quantity of goods picked up and placed, and location of goods picked up and placed from the first instruction; it compiles the instruction number, shelf area, quantity of goods picked up and placed, and location of goods picked up and placed into corresponding data codes in sequence, and finally forms a data code set according to the compilation rules; it converts the data code set into a second instruction recognizable by the management platform and sends the second instruction to the main management platform.
[0049] In practical applications, the main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the corresponding sub-platform of the management platform based on the shelf area information. Specifically, the main platform of the management platform has a pre-stored shelf area information association table, which includes at least the shelf area and its corresponding sub-platform of the management platform. When the main platform of the management platform receives the second instruction, it extracts the shelf area information from the second instruction and finds the sub-platform of the management platform corresponding to the shelf area as the target platform based on the shelf area information association table. The main platform of the management platform then sends the second instruction to the corresponding target platform.
[0050] In practical applications, after the main platform of the sensor network platform receives the shelf area information and the third instruction, it sends the third instruction to the corresponding sub-platform of the sensor network platform for the shelf area based on the shelf area information. Specifically, after receiving the shelf area information and the third instruction, the main platform of the sensor network platform first extracts the shelf area information and finds the corresponding sub-platform of the sensor network platform for the shelf area based on the shelf area information; the main platform of the sensor network platform then sends the third instruction to the corresponding sub-platform of the sensor network platform for the shelf area.
[0051] In practical applications, when there are multiple shelf area information extracted by the main platform of the sensor network platform, the main platform of the sensor network platform will sequentially find the sub-platforms of the sensor network platform corresponding to the multiple shelf areas, and simultaneously send the third instruction to all the found sub-platforms of the sensor network platform.
[0052] In practical applications, when the third instruction also includes execution time data, the one or more automated stacker machines, after receiving the third instruction, read the execution time data and execute the third instruction at the execution time.
[0053] In practical applications, the aforementioned Industrial Internet of Things (IIoT) for intelligent automated warehouses further includes: after one or more automated stacker cranes provide feedback on goods retrieval and placement results, the goods retrieval and placement results data is sent to all service platform sub-platforms through a sensor network platform and a management platform; all service platform sub-platforms store a goods data table, which corresponds to the quantity and storage location of all goods in all retrieval areas; when a service platform sub-platform obtains the goods retrieval and placement results data, it reads the data, retrieves the quantity and location of the goods based on the data, and updates the goods data table; before the user platform issues the first instruction, the corresponding service platform sub-platform sends the goods data table to the user platform, allowing the user platform to view the goods data and issue the first instruction based on the goods data table.
[0054] like Figure 3 As shown, the second embodiment of this specification provides a control method for an industrial Internet of Things (IoT) for an intelligent automated warehouse. The industrial IoT for the intelligent automated warehouse includes a user platform, a service platform, a management platform, a sensor network platform, and an object platform that interact sequentially from top to bottom.
[0055] The service platform adopts an independent layout, while the management platform and sensor network platform both adopt a front-sub-platform layout. The independent layout means that the service platform has multiple independent sub-platforms, each storing, processing, and / or transmitting data from different lower-level platforms. The front-sub-platform layout means that each platform has a main platform and multiple sub-platforms. The sub-platforms store and process data of different types or receiving objects sent from lower-level platforms, while the main platform aggregates, stores, processes, and transmits the data to the upper-level platform. The object platform is configured as an automated stacker crane in different shelving areas of an intelligent automated warehouse. The control method includes: each sub-platform of the service platform corresponds to a different user platform; a first instruction is issued by the user platform; the sub-platform of the service platform corresponding to the user platform receives the first instruction, converts it into a second instruction recognizable by the management platform, and issues the second instruction to the management platform. The platform consists of a main platform and a secondary platform. The secondary instruction includes at least an instruction number, a shelf area, a quantity of goods to be picked up or placed, and the location of the goods to be picked up or placed. The main platform receives the secondary instruction, extracts the shelf area information from the secondary instruction, and sends the secondary instruction to the corresponding sub-platform of the management platform for that shelf area. After receiving the secondary instruction, the sub-platform of the management platform extracts the instruction number, quantity of goods to be picked up or placed, and location of the goods to be picked up or placed, converts them into a third instruction that can be recognized by the automated stacker crane, and sends the shelf area information and the third instruction together to the main platform of the sensor network platform. After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the corresponding sub-platform of the sensor network platform for that shelf area based on the shelf area information. The sub-platform of the sensor network platform receives the third instruction and sends it to one or more automated stacker cranes corresponding to it. One or more automated stacker cranes perform goods picking and placing based on the third instruction and provide feedback on the goods picking and placing results.
[0056] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.
[0057] In the several embodiments provided in this specification, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0058] The units described as separate components may or may not be physically separate. As will be appreciated by those skilled in the art, the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.
[0059] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0060] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or grid device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0061] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this specification. It should be understood that the above descriptions are merely specific embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
[0062] Figure 4 This is an exemplary schematic diagram illustrating the determination of a target's standby position based on a reinforcement learning model, according to some embodiments of this specification. In some embodiments, process 400 may be executed by a management platform.
[0063] like Figure 4 As shown, the management platform can input environmental state information 410 into the reinforcement learning model 420, and the reinforcement learning model 420 can output the target waiting position 430 based on the input environmental state information 410. More information about the target waiting position can be found in [link to relevant documentation]. Figure 1 And its related descriptions.
[0064] Environmental status information refers to information used to describe the state of an automated stacker crane during its idle period. For example, environmental status information 410 may include the location information of the automated stacker crane, the idle duration, and the distance traveled during the idle period.
[0065] The location information of an automated stacker crane is used to represent information such as the rack, row number, and column number corresponding to the current location of the automated stacker crane. The location information of an automated stacker crane can be represented in various ways, such as a three-dimensional vector (x, y, z), which means that the automated stacker crane is currently located at a position where it can pick up and place goods in the x-th row and y-th column of the z-th rack.
[0066] In some embodiments, the management platform may obtain the location information of the automated stacker by means of position sensors deployed on the automated stacker or other methods.
[0067] Idle time refers to the time elapsed from the moment the automated stacker crane enters the idle period to the current moment. For example, if the automated stacker crane completes the pick-and-place operation at 8:00 and enters the idle period, and does not receive any new instructions between 8:00 and 8:05, then the idle time corresponding to 8:05 is 5 minutes.
[0068] In some embodiments, the management platform can determine the idle time of the automated stacker in various ways, such as by setting a clock inside the automated stacker or in the management platform. This clock can start counting from when the automated stacker completes an instruction and be reset to zero when the automated stacker receives a new instruction. The value displayed by the clock is the idle time of the automated stacker.
[0069] The travel length during the idle period refers to the path length that the automated stacker crane has traveled during the idle period. In some embodiments, the travel length during the idle period may include the path length that the stacker crane body travels on a plane (e.g., the ground) and the path length that the stacker crane loading platform travels vertically within the stacker crane body.
[0070] In some embodiments, the management platform can determine the movement length during the idle period by obtaining the location information of the automated stacker in real time or by other means.
[0071] The reinforcement learning model 420 can be used to determine the target's waiting position. The input of the reinforcement learning model 420 is the environmental state information 410, and the output is the target's waiting position 430. The reinforcement learning model 420 includes an environment module 421 and an optimal action determination module 422.
[0072] In some embodiments, when determining the target standby position 430 based on the reinforcement learning model 420, environmental state information 410 can be input into the reinforcement learning model 420. Within the model, the environmental state information 410 is input into the environment module 421, and the environment module 421 outputs a set of optional actions. Within the model, the environmental state information 410 and the set of optional actions are input into the optimal action determination module 422, and the optimal action determination module 422 outputs the optimal optional action 427. The target position corresponding to the optimal optional action 427 output by the optimal action determination module 422 is determined as the target standby position 430, and is used as the output of the reinforcement learning model 420. For example, if the optimal optional action is to remain still, then the target position corresponding to this optimal optional action is the current position information of the automated stacker crane. As another example, if the optimal optional action is to move to (x1, y1, z1), then the target position corresponding to this optimal optional action is (x1, y1, z1), and the target standby position is (x1, y1, z1).
[0073] The environment module 421 may include an optional action determination submodule 423, a state determination submodule 424, and a reward determination submodule 425. During the prediction process of the reinforcement learning model 420, the environment module 421 can determine the set of optional actions based on the environment state information 410 through the optional action submodule 423. During the training process of the reinforcement learning model 420, the state determination submodule 424 and the reward determination submodule 425 in the environment module 421 can be used to determine the environment state information and reward value at the next time step, respectively.
[0074] The optional action determination submodule 423 can determine the set of optional actions of the automated stacker crane at the current moment based on the environmental state information at the current moment.
[0075] The set of optional actions refers to the set of actions that an automated stacker crane can perform under a given environmental condition. In some embodiments, the actions that an automated stacker crane can perform may include remaining stationary and moving to a target location. The target location refers to the location that the automated stacker crane can currently move to (e.g., a row or column of a rack, or a picking station). The target locations that the automated stacker crane can move to may differ under different environmental conditions.
[0076] In some embodiments, the environment module 421 can determine the set of optional actions for the automated stacker machine at the current moment based on the location information of the automated stacker machine in the environmental state information at the current moment, namely, moving to a position that meets preset conditions with respect to the current position of the automated stacker machine and / or remaining stationary. In some embodiments, the preset conditions can be determined based on the structure of the automated warehouse, the racking structure, etc. For example, if the current position of the automated stacker machine is (1, 1, 4), and the preset condition is: the difference between the rack number at the current position and the rack number corresponding to the position to be moved does not exceed 2, then the position corresponding to remaining stationary or moving to any row and any column of racks 2-6 is determined as the set of optional actions for the automated stacker machine at the current moment.
[0077] The state determination submodule 424 can determine the environmental state information for the next moment based on the current environmental state information and the optimal optional action output by the optimal action determination module. For example, if the position information of the automated stacker in the current environmental state information is (1, 1, 1), and the optimal optional action output by the optimal action determination module is to go to the target waiting position (2, 1, 1), after the automated stacker executes the optimal optional action, the state determination submodule 424 determines the position information of the automated stacker for the next moment to be (2, 1, 1), and updates the idle time and the movement length during the idle period in the environmental state information based on the elapsed time and the path length traveled.
[0078] The reward determination submodule 425 can be used to determine the reward value. The reward value can be used to evaluate the degree to which the efficiency of the next execution instruction is improved after the automated stacker crane performs an action. For example, the reward value can be higher for actions with a high degree of improvement, and lower for actions with a low degree of improvement or negative improvement. The reward value can be represented numerically or in other ways. The degree of efficiency improvement for the next execution instruction can be determined based on the distance between the target location the automated stacker crane travels to during its idle period and the pickup location corresponding to the instruction received by the automated stacker crane at a future time. The shorter this distance, the faster the automated stacker crane travels to the pickup location corresponding to the instruction, and the greater the degree of efficiency improvement for the next execution instruction. In some embodiments, the reward determination submodule 425 can determine the reward value based on a set formula.
[0079] In some embodiments, when the automated stacker performs a stationary action or a move to a target location, the reward determination submodule 425 can determine the reward value for the corresponding action based on different methods.
[0080] In some embodiments, the reward value for an automated stacker crane performing a stationary action can be related to the idle duration and the length of movement during the idle period. For example, the longer the idle duration and the longer the length of movement during the idle period, the greater the reward value for the automated stacker crane performing a stationary action.
[0081] For example, the reward value for an automated stacker crane performing a stationary action can be calculated using the following formula (1):
[0082] r = k l t l +k2t2 (1)
[0083] Where r is the reward value of the action, t1 is the idle time, and t2 is the movement length during the idle period; k1 and k2 are the weight coefficients of the idle time and the movement length during the idle period, respectively. k1 and k2 can be determined based on experience. For example, k1 and k2 can both be 0.5.
[0084] In some embodiments, the reward value for an automated stacker crane to perform a move to a target location may be related to the expected value of the move to the target location.
[0085] The expected value of movement can be used to describe the benefit of an automated stacker crane moving to a target location. For example, if the target location of the automated stacker crane during its idle period is (1, 1, 1), and the pickup location corresponding to the instruction received by the automated stacker crane at a future time is (1, 1, 1), then the automated stacker crane can directly execute the instruction without first moving, and it can be considered that the action performed during this idle period has brought a high benefit.
[0086] In some embodiments, the expected movement value can be determined based on the expected value lookup table in the reward determination submodule 425. For example, the expected value lookup table may provide the expected movement value of the automated stacker crane to different target locations from different positions. The management platform can determine the expected movement value of the automated stacker crane to the target location when it is in its current position by querying the expected value lookup table. In some embodiments, the expected value lookup table may be stored in the management platform and updated periodically by managers based on experience.
[0087] In some embodiments, the expected value of movement may be related to a movement reward value and a movement penalty value.
[0088] Movement bonus value refers to the bonus value generated when an automated stacker moves to a target location.
[0089] In some embodiments, the mobility reward value can be determined based on the frequency with which the target location was designated as the pickup location in the instruction within a historical time period. For example, the mobility reward value can be calculated using the following formula (2):
[0090]
[0091] Where, r m p is the mobile reward value. a p represents the total number of instructions issued by the management platform within a historical time period. b The number of times the target location was designated as the pickup location in the instruction within the historical time period, and k is a preset parameter for adjusting the size of the movement reward value. k can be determined based on experience, for example, k can be 10.
[0092] In some embodiments, the mobile reward value may be related to a first hit, and the management platform may determine the mobile reward value based on the first hit. More information about the first hit and determining the mobile reward value based on the first hit can be found at [link to relevant documentation]. Figure 5 And its related descriptions.
[0093] In some embodiments, the mobile reward value can also be related to percentage compliance, and the management platform can determine the mobile reward value based on percentage compliance. More information on percentage compliance and determining mobile reward values based on percentage compliance can be found at [link to relevant documentation]. Figure 6 And its related descriptions.
[0094] The movement penalty value refers to the penalty incurred when an automated stacker crane moves to a target location. The penalty value describes the inconvenience caused by the movement of the automated stacker crane and can be related to the movement distance and energy consumption. For example, the greater the movement distance and the greater the energy consumption, the larger the penalty value.
[0095] In some embodiments, the movement penalty value can be determined based on the distance between the current position and the target position of the automated stacker crane. For example, the movement penalty value can be calculated using the following formula (3):
[0096] r p =kd (3)
[0097] Where, r p Here, d is the distance between the current position and the target position of the automated stacker, and k is a preset parameter for adjusting the size of the movement penalty value. k can be determined based on experience; for example, k can be 1.
[0098] In some embodiments, the expected value of movement can be determined based on a movement reward and a movement penalty. For example, the expected value of movement can be the difference between the movement reward and the movement penalty.
[0099] In some embodiments of this specification, the movement penalty value is determined based on the distance between the current position and the target position of the automated stacker crane, and the movement expectation value is determined based on the combination of the movement reward value and the movement penalty value. This allows the reinforcement learning model to comprehensively consider both improving efficiency and reducing energy consumption, thereby avoiding pursuing efficiency alone or reducing energy consumption alone, and enabling more efficient scheduling of automated stacker cranes in the automated warehouse.
[0100] In some embodiments, the expected movement value can be determined as the reward value for the automated stacker to perform the action of moving to the target location.
[0101] The optimal action determination module can determine the optimal optional action based on the current environmental state information 410 and the set of optional actions. The input of the optimal action determination module 422 is the environmental state information 410 and the set of optional actions. The output of the optimal action determination module 422 is the optimal optional action 427.
[0102] In some embodiments, for each optional action in the set of optional actions, the optimal action determination module can output a recommended value. The optimal action determination module can determine the optional action with the largest recommended value as the optimal optional action and output it.
[0103] In some embodiments, the optimal action determination module 422 can be a machine learning model, which can be implemented in various ways, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), etc.
[0104] In some embodiments, the optimal action determination module 422 can be trained based on reinforcement learning methods, such as Deep Q-Learning Network (DQN) or Double Deep Q-Learning Network (DDQN). Training samples can be historical environmental state information, with labels representing the optimal action available under that historical environmental state. Training samples can be obtained based on historical data. The labels of the training samples can be obtained using reinforcement learning methods.
[0105] In some embodiments, the management platform can periodically execute the reinforcement learning model 420 and output the optimal optional action based on preset trigger conditions. For example, the preset trigger condition is: the automated stacker crane has completed executing the currently output optimal optional action of the reinforcement learning model 420. Another example is: the automated stacker crane has completed executing the currently output optimal optional action of the reinforcement learning model 420, and the time since the last execution of the reinforcement learning model has exceeded a preset period value, such as 10 seconds.
[0106] In some embodiments of this specification, the target standby position of the automated stacker crane is determined by a reinforcement learning model. This allows the automated stacker crane to move in advance to the pickup location that may correspond to future instructions based on environmental information, thereby reducing the time for the automated stacker crane to execute instructions and improving the overall cargo retrieval efficiency of the automated warehouse.
[0107] Figure 5 This is an exemplary flowchart illustrating the determination of a mobile reward value based on a first hit rate, according to some embodiments of this specification. In some embodiments, process 500 may be executed by a management platform. Figure 5 As shown, process 500 includes the following steps:
[0108] Step 510: Determine the first time period.
[0109] The first time period refers to the period that starts at a historical time and ends at the current time. For example, if the current time is 9:00, the first time period can be the past hour (i.e., 8:00-9:00), the past two hours (i.e., 7:00-9:00), etc.
[0110] In some embodiments, the management platform can determine a first time period based on the historical execution of instructions by the automated stacker crane. For example, the management platform can determine the shortest historical time period in which the number of instructions executed by the automated stacker crane in the past meets a preset condition as the first time period. For instance, suppose the current time is 9:00, the preset condition is that the number of instructions executed by the automated stacker crane in the past is greater than or equal to 3, and the instructions executed by the automated stacker crane in the historical time period (the most recent few instruction executions) are: one instruction was executed at 7:05, one instruction was executed at 7:35, one instruction was executed at 8:00, one instruction was executed at 8:30, and one instruction was executed at 8:50. Then the management platform can determine the first time period as 8:00-9:00, which is the shortest time period that meets the preset condition (the number of instructions completed is 3).
[0111] Step 520: Determine at least one historical idle period based on the first time period, and determine at least one second hit rate based on the at least one historical idle period.
[0112] Historical idle periods refer to idle periods where both the start and end times fall within the first time period.
[0113] In some embodiments, the management platform can obtain the instruction execution status of the automated stacker crane within a first time period, and determine at least one historical idle period based on the instruction execution status. For example, suppose the current time is 9:00, the first time period is 8:00-9:00, and the instruction execution status of the automated stacker crane within the first time period is as follows: instruction A starts executing at 8:00, instruction A is completed at 8:05, instruction B starts executing at 8:15, instruction B is completed at 8:17, instruction C starts executing at 8:35, instruction C is completed at 8:40, and so on. Then, the time period 8:05-8:15 is a historical idle period, and the time period 8:17-8:35 is a historical idle period.
[0114] The second hit rate can be used to describe the proximity of an automated stacker crane's target standby position during a historical idle period to the pickup position in the instruction received at the end of that historical idle period. For example, if an automated stacker crane is located at the target standby position (1, 1, 1) during a certain historical idle period, and the pickup position corresponding to the instruction received by the automated stacker crane at the end of that historical idle period is (1, 1, 1), then the second hit rate for that historical idle period is 100%.
[0115] In some embodiments, the management platform can obtain the target standby position of the automated stacker crane during the historical idle period and the pickup position in the instruction received at the end of the historical idle period, based on the historical idle period, and determine the second hit rate based on the target standby position and the pickup position. For example, the second hit rate can be calculated using the following formula (4):
[0116]
[0117] Where c is the second hit rate, x a y a z a Corresponding to the target standby position coordinates (x) of the automated stacker crane during the historical idle period. a y a , z a ), x b y b z b The pickup location (x) corresponds to the instruction received at the end of the historical idle period. b y b , z b ), |x a -x b | indicates that x a -x b Take the absolute value, max(x) a ,x b ) indicates taking x a x bThe maximum values of the two, k1, k2, and k3, are weighting parameters. In some embodiments, k1, k2, and k3 can be determined based on experience; for example, k1, k2, and k3 can all be equal to 1.
[0118] Step 530: Based on at least one second hit, determine the weight corresponding to each second hit in the at least one second hit.
[0119] Weights can be used to describe how reliable the second hit is when calculating the first hit. For example, the larger the weight corresponding to the second hit, the more reliable the second hit is.
[0120] In some embodiments, the management platform can obtain the end time of the idle period corresponding to each of the at least one second hits, and determine the weight corresponding to each second hit based on the interval between the end time of the idle period and the current time. For example, the weight can be calculated using the following formula (5):
[0121]
[0122] Where m is the weight corresponding to the second hit rate, t is the interval between the end of the idle period corresponding to the second hit rate and the current time, and k is a preset parameter for adjusting the weight size. k can be determined based on experience, for example, k can be 1.
[0123] Step 540: Determine the first hit rate based on at least one second hit rate and its corresponding weight.
[0124] The first hit rate can be used to describe the average proximity of the target standby location of an automated stacker crane during at least one idle period within a first time period to the pickup location in the instruction received at the end of that idle period.
[0125] In some embodiments, the management platform can determine the first hit rate by weighted summation based on at least one second hit rate and its corresponding weight. For example, the first hit rate can be calculated using the following formula (6):
[0126]
[0127] in, Let A be the first hit score, and A be the set of second hit scores. Each second hit score in set A is numbered with a positive integer c. i Let m be the second hit score of the i-th element in this set. i For the second hit rate c i The weight.
[0128] Step 550: Determine the movement reward value based on the first hit rate.
[0129] In some embodiments, the management platform may determine the mobile reward value based on a first hit rate. For example, the mobile reward value can be calculated using the following formula (7):
[0130]
[0131] Where, r m For mobile reward value, For the first hit rate, p a p represents the total number of instructions issued by the management platform within a historical time period. b The number of times the target location was designated as the pickup location in the instruction within the historical time period, and k is a preset parameter for adjusting the size of the movement reward value. k can be determined based on experience, for example, k can be 10.
[0132] In some embodiments of this specification, by introducing a first hit rate to determine the movement reward value, the reliability of the target standby location previously visited by the automated stacker can be fully considered, thereby using historical data as a reference to guide future behavior and more accurately determining the target standby location of the automated stacker.
[0133] Figure 6 This is an exemplary flowchart illustrating the determination of mobile reward values based on percentage compliance according to some embodiments of this specification. In some embodiments, process 600 may be executed by a management platform. Figure 6 As shown, process 600 includes the following steps:
[0134] Step 610: Obtain the current percentage of goods and the percentage of empty spaces.
[0135] The cargo percentage refers to the ratio of goods stored in the automated warehouse to its total capacity. The empty space percentage refers to the percentage of empty spaces in the automated warehouse that are currently not used for storage, relative to its total capacity. For example, if there are currently 400 items stored in the automated warehouse and the total capacity is 1000 items, then the cargo percentage is 40%, and the empty space percentage is 60%. In some embodiments, the management platform can obtain the current cargo percentage and empty space percentage through sensors on the shelves, cameras, or other means.
[0136] Step 620: Obtain at least one instruction issued by the management platform during a historical time period, and determine the proportion of outbound instructions and the proportion of inbound instructions based on at least one instruction.
[0137] In some embodiments, the historical time period can be determined based on experience. For example, the historical time period can be the most recent 2 hours, the most recent 1 hour, etc. In some embodiments, the management platform can obtain the instructions it sent within the historical time period, and count the number of outbound instructions and inbound instructions. Based on the number of outbound instructions and the total number of instructions sent, the proportion of outbound instructions is determined, and based on the number of inbound instructions and the total number of instructions sent, the proportion of inbound instructions is determined.
[0138] Step 630: In response to the target location being a certain row and column of a certain shelf, determine the percentage conformity based on the percentage of goods and the percentage of outbound instructions; in response to the target location being a picking station, determine the percentage conformity based on the percentage of empty spaces and the percentage of inbound instructions.
[0139] Percentage conformity refers to the degree of conformity between the percentage of goods in the target shelf and the percentage of outbound instructions issued by the management platform during a historical period, and the percentage of empty spaces in the target shelf and the percentage of inbound instructions issued by the management platform during a historical period. For example, if the percentage of goods is 90% and the percentage of outbound instructions is 80%, then the probability of the next instruction being an outbound instruction is very high, hence the percentage conformity is high. On the other hand, if the percentage of goods is 80% and the percentage of outbound instructions is 10%, based on the percentage of goods, the next instruction is more likely to be an outbound instruction; based on the percentage of instructions, the next instruction is more likely to be an inbound instruction. In this case, it is difficult to determine whether the next instruction is more likely to be an outbound or inbound instruction, hence the percentage conformity can be lower.
[0140] In some embodiments, in response to the target location being a row and column of a shelf, the management platform can determine the percentage compliance based on the percentage of goods, the percentage of outbound instructions, the percentage of empty spaces, and the percentage of inbound instructions. For example, the percentage compliance can be calculated using the following formula (8):
[0141] e = a c b c (8)
[0142] Where e represents the percentage conformity, a c b represents the percentage of goods on the shelf corresponding to the target location. c This represents the percentage of outbound instructions issued by the management platform within a historical time period.
[0143] In some embodiments, in response to the target location being a pickup station, the management platform can determine the percentage compliance based on the percentage of empty slots and the percentage of inbound instructions. For example, the percentage compliance can be calculated using the following formula (9):
[0144] e = a d b d (9)
[0145] Where e represents the percentage conformity, a d b represents the percentage of empty spaces in the automated warehouse.d This represents the percentage of data entry instructions issued by the management platform within a historical time period.
[0146] Step 640: Determine the mobile reward value based on the percentage compliance.
[0147] In some embodiments, the management platform may determine the mobile reward value based on percentage compliance and first hit rate.
[0148] For example, the movement reward value can be calculated using the following formula (10):
[0149]
[0150] Where, r m Here, 'e' represents the mobile reward value, and 'e' represents the percentage of compliance. For the first hit rate, p a p represents the total number of instructions issued by the management platform within a historical time period. b The number of times the target location was designated as the pickup location in the instruction within the historical time period, k1 and k2 are weighting coefficients, which can be determined based on experience. For example, k1 can be 10 and k2 can be 5.
[0151] In some embodiments of this specification, the movement reward value is determined by introducing a proportion conformity degree. This not only considers the relevant information of the automated stacker crane itself, but also the information such as the proportion of goods in the automated warehouse and the proportion of outbound and inbound instructions. This allows the reinforcement learning model to learn the relevant information of the automated warehouse, thereby more accurately determining the target standby position of the automated stacker crane.
[0152] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0153] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0154] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of embodiments that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0155] Similarly, it should be noted that, in order to simplify the descriptions disclosed herein and thus aid in the understanding of one or more embodiments, the foregoing description of embodiments in this specification sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiments disclosed above.
[0156] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0157] For each patent, patent application, patent application publication, and other material, such as articles, books, specifications, publications, and documents, referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0158] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. An industrial Internet of Things (IoT) for intelligent automated warehouses, characterized in that, This includes a user platform, a service platform, and a management platform, wherein the management platform is configured to perform the following operations: The user platform sends a first instruction to the service platform, the service platform determines a second instruction based on the first instruction, and sends it to the management platform; in response to receiving the second instruction from the service platform, it determines whether the motion force of the automated stacker crane meets the execution conditions, the first instruction including a goods retrieval instruction; In response to the motion force of the automated stacker satisfying the execution condition, the automated stacker is controlled to perform a pick-and-place operation; In response to the automated stacker completing the pick-and-place operation, a target standby position is determined, and the automated stacker is controlled to move to the target standby position; in, The industrial Internet of Things also includes: a sensor network platform and an object platform; the user platform, the service platform, the management platform, the sensor network platform, and the object platform interact sequentially from top to bottom; The service platform adopts an independent layout, while the management platform and the sensor network platform both adopt a front-sub-platform layout. The independent layout means that the service platform has multiple independent sub-platforms, each storing, processing, and / or transmitting data from different lower-level platforms. The front-sub-platform layout means that each platform has a main platform and multiple sub-platforms. The sub-platforms store and process different types of data sent from lower-level platforms or data destined for different recipients. The main platform aggregates, stores, and processes the data from the multiple sub-platforms and transmits the data to the upper-level platform. The target platform is configured as an automated stacker crane in different shelving areas of an intelligent automated warehouse. The service platform's sub-platforms correspond to different user platforms. The first instruction is issued by the user platform. The sub-platform of the service platform corresponding to the user platform receives the first instruction, converts it into a second instruction that the management platform can recognize, and sends the second instruction to the management platform's main platform. The second instruction includes at least the instruction number, shelf area, quantity of goods to be picked up or placed, and location of goods to be picked up or placed. The main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the sub-platform of the management platform for the corresponding shelf area based on the shelf area information; After receiving the second instruction, the sub-platform of the management platform extracts the instruction number, the quantity of goods picked up and placed and the location of goods picked up and placed in the second instruction and converts them into a third instruction that the automatic stacker can recognize. The sub-platform then sends the shelf area information and the third instruction to the main platform of the sensor network platform. After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information. The sub-platform of the sensor network platform receives the third instruction and sends it to one or more of its corresponding automated stacker machines. The one or more automated stacker machines perform cargo handling based on the third instruction and feed back cargo handling result data. The corresponding service platform's sub-platform receives the first instruction, converts it into a second instruction recognizable by the management platform, and sends the second instruction to the management platform's main platform. Specifically: When a sub-platform of the service platform receives the first instruction, the sub-platform of the service platform shall at least extract the instruction number, the shelf area, the quantity of goods picked up and placed, and the location of goods picked up and placed from the first instruction; The instruction number, the shelf area, the quantity of goods picked up and placed, and the location of goods picked up and placed are sequentially compiled into corresponding data codes, and finally all codes are formed into a data code set according to the compilation rules. The data code set is converted into a second instruction that the management platform can recognize, and the second instruction is sent to the main platform of the management platform. The main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the sub-platform of the management platform for the corresponding shelf area based on the shelf area information, specifically: The main platform of the management platform has a pre-stored shelf area information association table, which includes at least the shelf area and its corresponding sub-platform of the management platform; When the main platform of the management platform receives the second instruction, it extracts the shelf area information from the second instruction and finds the sub-platform of the management platform corresponding to the shelf area as the target platform based on the shelf area information association table. The main platform of the management platform sends the second instruction to the corresponding target platform; After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information, specifically: After receiving the shelf area information and the third instruction, the main platform of the sensor network platform first extracts the shelf area information and finds the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information. The main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area.
2. The Industrial Internet of Things for Intelligent Automated Warehouses according to claim 1, characterized in that, When the third instruction also includes execution time data, the one or more automated stacker machines, after receiving the third instruction, read the execution time data and execute the third instruction at the execution time.
3. The Industrial Internet of Things for Intelligent Automated Warehouses according to claim 1, characterized in that, Also includes: After the one or more automated stacker cranes feed back the cargo retrieval and placement result data, the cargo retrieval and placement result data is sent to the sub-platforms of all service platforms through the sensor network platform and the management platform; All the service platforms' sub-platforms store cargo data tables, which correspond to the quantity and storage location of all cargo in all pickup areas; When the sub-platform of the service platform obtains the cargo retrieval and placement result data, it reads the cargo retrieval and placement result data, obtains the retrieval and placement quantity and location based on the cargo retrieval and placement data, and updates the cargo data table. Before the user platform issues the first instruction, the corresponding service platform's sub-platform sends the cargo data table to the user platform. The user platform can then view the cargo data and issue the first instruction based on the cargo data table.
4. The Industrial Internet of Things for Intelligent Automated Warehouses according to claim 1, characterized in that, The target standby position is determined based on a reinforcement learning model; the actions that the automated stacker crane can perform in the reinforcement learning model include: remaining stationary and moving to the target position; wherein... The reward value for the automated stacker crane to perform the stationary action is related to the idle time and the movement length during the idle period; the movement length during the idle period refers to the path length that the automated stacker crane has traveled during the idle period. The reward value for the automated stacker to perform the action of moving to the target location is related to the expected value of the movement to the target location; the expected value of the movement is used to describe the benefit brought by the automated stacker moving to the target location.
5. A control method for an industrial Internet of Things (IoT) for intelligent automated warehouses, implemented based on the industrial IoT described in any one of claims 1-4, characterized in that, The Industrial Internet of Things (IIoT) comprises, from top to bottom, a user platform, a service platform, a management platform, a sensor network platform, and an object platform that interact sequentially. The service platform adopts an independent deployment, while the management platform and the sensor network platform both adopt a front-sub-platform deployment. The independent deployment means that the service platform sets up multiple independent sub-platforms, which respectively store, process, and / or transmit data from different lower-level platforms. The front-sub-platform deployment means that each platform has a main platform and multiple sub-platforms. The multiple sub-platforms respectively store and process data of different types or different recipients sent by lower-level platforms. The main platform summarizes, stores, and processes the data from the multiple sub-platforms and transmits the data to the upper-level platform. The object platform is configured as an automated stacker crane in different shelving areas of an intelligent automated warehouse; The control method includes: The service platform's sub-platforms correspond to different user platforms. The first instruction is issued by the user platform. The sub-platform of the service platform corresponding to the user platform receives the first instruction, converts it into a second instruction that the management platform can recognize, and sends the second instruction to the management platform's main platform. The second instruction includes at least the instruction number, shelf area, quantity of goods to be picked up or placed, and location of goods to be picked up or placed. The main platform of the management platform receives the second instruction, extracts the shelf area information from the second instruction, and sends the second instruction to the sub-platform of the management platform for the corresponding shelf area based on the shelf area information; After receiving the second instruction, the sub-platform of the management platform extracts the instruction number, the quantity of goods picked up and placed and the location of goods picked up and placed in the second instruction and converts them into a third instruction that can be recognized by the automatic stacker crane. The sub-platform then sends the rack area information and the third instruction to the main platform of the sensor network platform. After receiving the shelf area information and the third instruction, the main platform of the sensor network platform sends the third instruction to the sub-platform of the sensor network platform corresponding to the shelf area based on the shelf area information. The sub-platform of the sensor network platform receives the third instruction and sends it to one or more of its corresponding automated stacker machines. The one or more automated stacker machines perform cargo handling based on the third instruction and feed back cargo handling result data.
6. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions from the storage medium, the computer runs the industrial Internet of Things control method for intelligent automated warehouses as described in claim 5.